Multi-phase cross-component prediction

By using adjacent samples of the second color component to generate samples of the first color component in the video decoding technology, the problem of low prediction accuracy and compression efficiency in the prior art is solved, and more efficient video data compression is achieved.

CN120226352APending Publication Date: 2025-06-27TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079922.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2023-10-31
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the cross-component intra prediction, it is difficult for the existing video decoding technology to effectively use adjacent samples of the second color component to predict the samples of the first color component, resulting in low prediction accuracy and compression efficiency.

Method used

By receiving syntax elements in the video bitstream, identifying samples of the first color component and the second color component, and generating samples of the first color component based on adjacent samples of the second color component, prediction is performed using linear or nonlinear functions, and the weighting factor and offset parameters are adjusted to improve prediction accuracy.

Benefits of technology

The cross-component intra prediction accuracy of video data is improved, the efficiency of video compression is enhanced, and the storage and transmission requirements of video data are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226352A_ABST
    Figure CN120226352A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems for coding video. In one aspect, a video bitstream includes a current coded block of an image frame and includes a cross-component intra prediction mode. A computing system identifies a sample of a first color component and a sample of a second color component co-located with the sample of the first color component. At least two adjacent samples of the second color component are identified, and the location of each adjacent sample is identified by horizontal delta coordinate values or vertical delta coordinate values relative to the sample of the second color component. The computing system generates samples of the first color component based on at least two adjacent samples of the second color component, and reconstructs the current coded block based at least on the generated samples of the first color component.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 538,041, filed on September 12, 2023, entitled "Multi-Phase Cross Component Prediction", and this application is a continuation of, and claims priority to, U.S. Patent Application No. 18 / 497,921, filed on October 30, 2023, entitled "Multi-Phase Cross Component Prediction". Technical Field

[0003] The disclosed embodiments generally relate to video coding, including but not limited to systems and methods for cross-component intra prediction of video data. Background Art

[0004] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital imaging devices, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive or otherwise convey digital video data across communication networks, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data.

[0005] A variety of video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Expert Group (MPEG) coding. Video coding generally uses prediction methods (e.g., inter prediction, intra prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality.

[0006] HEVC (also known as H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T (International Telecommunication Union - Telecommunication Standardization Sector, ITU-T) and ISO / IEC (International Organization for Standardization / International Electrotechnical Commission, ISO / IEC) in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC) (also known as H.266) is a video compression standard designed to succeed HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (version 1) and 2022 (version 2). AV1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 with Specification Errata 1 was released. Summary of the Invention

[0007] As mentioned above, encoding (compression) reduces the bandwidth and / or storage space requirements. As described in detail later, both lossless compression and lossy compression can be employed. Lossless compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process where the original video information is not fully preserved during encoding and cannot be fully recovered during decoding. When lossy compression is used, the reconstructed signal may not be the same as the original signal, but the distortion between the original signal and the reconstructed signal is small enough such that the reconstructed signal is useful for the intended application. The amount of allowable distortion depends on the application. For example, users of certain consumer video streaming applications may tolerate higher distortion than users of movie or television broadcast applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerable distortion generally allows an encoding algorithm that produces higher losses and higher compression ratios.

[0008] The present disclosure describes applying multiple parameters to implement cross-component intra prediction of video data in a cross-component intra prediction (CCIP) mode, where each of a plurality of samples of a first color component of a current decoding block is determined based on one or more samples of a second color component. Samples of the first color component are determined based on co-located samples and / or at least two neighboring samples of the second color component using a linear or non-linear function. In some embodiments, all parameters (e.g., weighting factors, offsets) specifying the linear or non-linear function are determined using reconstructed samples of the first and second color components in a neighboring region of the current decoding block. Alternatively, in some embodiments, at least a subset or all of the parameters specifying the linear or non-linear function are explicitly signaled via a video bitstream. In some cases, the positions of at least two neighboring samples of the second color component are asymmetric and have a center shifted in position relative to samples of the first color component. When the decoder applies neighboring samples of the second color component to predict samples of the first color component, it compensates for spatial misalignment between samples of the two color components caused by different reasons (e.g., artifacts of a camera lens, user-controlled video post-processing).

[0009] According to some embodiments, a method of video decoding is provided. The method includes: receiving a video bitstream including a current decoding block of a current image frame. The video bitstream includes a syntax element for a cross-component intra prediction (CCIP) mode that indicates whether each sample of a first color component of the current decoding block is determined based on one or more samples of a second color component. The method further includes: identifying samples of the first color component and samples of the second color component co-located with the samples of the first color component in the current decoding block. The method further includes: for each of two or more neighboring samples of the second color component, obtaining from the video bitstream at least one of (i) a horizontal increment coordinate value and (ii) a vertical increment coordinate value relative to the sample of the second color component. The method further includes: identifying each of at least two neighboring samples of the second color component in the current decoding block based on at least one of (i) the horizontal increment coordinate value and (ii) the vertical increment coordinate value. The method further includes: generating samples of the first color component based on at least two neighboring samples of the second color component, and reconstructing the current decoding block based at least on the samples of the first color component generated according to the at least two neighboring samples of the second color component.

[0010] In some embodiments, at least two adjacent samples have a geometric center that is offset from the positions of the samples of the second color component. In some embodiments, for each adjacent sample of the second color component, at least one of the horizontal incremental coordinate value and the vertical incremental coordinate value includes an integer incremental coordinate value. Alternatively, in some embodiments, at least two adjacent samples of the second color component include a first adjacent sample, and at least one of the horizontal incremental coordinate value and the vertical incremental coordinate value includes a fractional incremental coordinate value.

[0011] In some embodiments, the first color component and the second color component correspond to two different color components among a set of green, blue, and red color components. Alternatively, in some embodiments, the first color component corresponds to a chrominance component and the second color component corresponds to a luminance component.

[0012] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes control circuitry and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.

[0013] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0014] Accordingly, apparatuses and systems having methods for decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video decoding.

[0015] The features and advantages described in this specification are not necessarily all inclusive, and in particular, in view of the figures, specification, and claims provided in this disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in this specification has been principally selected for readability and guidance purposes and is not necessarily selected to depict or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] For a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and are therefore not necessarily considered restrictive, as those skilled in the art will understand that the specification may allow other valid features when reading the present disclosure.

[0017] Figure 1 is a block diagram showing an example communication system according to some embodiments.

[0018] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.

[0019] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.

[0020] Figure 3 is a block diagram showing an example server system according to some embodiments.

[0021] Figure 4 shows an example current decoding block including at least two color components according to some embodiments.

[0022] Figure 5 is an example scheme for generating samples of a first color component based on samples of a second color component 404 according to some embodiments.

[0023] Figure 6 is a flowchart showing an example method for decoding video according to some embodiments.

[0024] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, like reference numerals may be used to represent like features. Detailed Description

[0025] The present disclosure describes cross-component intra prediction of video data in a cross-component intra prediction (CCIP) mode, where each sample of a first color component of a current decoding block (e.g., each of a plurality of chrominance samples) is determined based on one or more samples of a second color component (e.g., one or more luma samples). The first color component and the second color component are different color components. Samples of the first color component are determined based on co-located samples and / or at least two neighboring samples of the second color component, using a linear or non-linear function. In some embodiments, reconstructed samples of the first color component and the second color component in an adjacent region of the current decoding block are used to derive parameters (e.g., weighting factors, offsets) that specify the linear or non-linear function. Alternatively, in some embodiments, at least a subset of the parameters that specify the linear or non-linear function are signaled explicitly. Further, in some embodiments, the first color component and the second color component are two different color components among the R, G, and B color components. In another example, the first color component is a chrominance color component and the second color component is a luma color component. The luma color component and the chrominance color component have different resolutions. Luma samples are converted to corresponding chrominance samples at either resolution of the luma samples and the chrominance samples. In some cases, at least two neighboring samples of the second color component have a center shifted from the samples of the first color component. By these means, when the decoder applies these neighboring samples of the second color component to predict samples of the first color component, it compensates for the spatial misalignment between samples of the two color components caused by different reasons (e.g., artifacts of a camera lens, user-controlled video post-processing).

[0026] Figure 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, such as for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0027] The source device 102 includes a video source 104 (e.g., a camera device component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 can be of high data volume compared to the encoded video bitstreams 108 generated by the encoder component 106. Since the encoded video bitstreams 108 are of lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth to transmit and less storage space to store. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video data to the network 110).

[0028] One or more networks 110 represent any number of networks for transmitting information between the source device 102, the server system 112 / or the electronic device 120, including for example wired (wired) communication networks and / or wireless communication networks. One or more networks 110 can exchange data in circuit - switched channels and / or packet - switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as the encoded video stream from the source device 102). The server system 112 includes a decoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the decoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the decoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the decoder component 114 is configured to decode the encoded video bitstream 108 and re - encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108.

[0030] In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim an encoded video bitstream 108 for customizing potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0031] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be presented on a display or other type of presentation device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0032] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of server systems, personal computers, portable devices (e.g., smart phones, tablet computers, or laptop computers), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0033] In an example operation of the communication system 100, the source device 102 transmits the encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may use the decoder component 114 to decode and / or encode the encoded video bitstream 108. For example, the server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more decoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to recover and optionally display the video pictures.

[0034] In some embodiments, the transmissions discussed above are unidirectional data transmissions. Unidirectional data transmissions are sometimes used in media service applications and the like. In some embodiments, the transmissions discussed above are bidirectional data transmissions. Bidirectional data transmissions are sometimes used in video conferencing applications and the like. In some embodiments, the encoded video bitstream 108 and / or the encoded video data 116 are encoded and / or decoded according to any video decoding / compression standard described herein, such as any of the video decoding / compression standards of HEVC, VVC, and / or AV1.

[0035] Figure 2A is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. A person of ordinary skill in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0036] The encoder component 106 is configured to decode and / or compress the pictures of the source video sequence into a decoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate decoding speed is a function of the controller 204. In some embodiments, the controller 204 controls the other functional units as described below and is functionally coupled to the other functional units. The parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. A person of ordinary skill in the art can easily identify other functions of the controller 204, as such functions may belong to the encoder component 106 optimized for a particular system design.

[0037] In some embodiments, the encoder component 106 is configured to operate in a decoding loop. In a simplified example, the decoding loop includes a source decoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture and reference pictures to be decoded) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a similar manner as a (remote) decoder to create sample data (in the case where the compression between the symbols and the decoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret when using prediction during decoding as reference picture samples. The principle of reference picture synchronization (and the drift that occurs in the case where synchronization cannot be maintained, e.g., due to channel errors) is known to those of ordinary skill in the art.

[0038] The operation of the decoder 210 can be the same as that of a remote decoder, such as the decoder component 122 described in detail below in conjunction with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the encoding of the symbols into the decoded video sequence by the entropy decoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding part including the buffer memory 252 and the parser 254 of the decoder component 122 may not be fully implemented in the local decoder 210.

[0039] At this point, it can be observed that any decoder technology other than the parsing / entropy decoding present in the decoder must also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described in detail. More detailed descriptions are only required and provided in certain places below.

[0040] As part of the operation of the source decoder 202, the source decoder 202 may perform motion compensated predictive decoding, which performs predictive decoding of an input frame by referring to one or more previously decoded frames designated as reference image frames from a video sequence. In this way, the decoding engine 212 decodes the difference between a pixel block of the input frame and a pixel block of the reference image frame, and the reference image frame may be selected as the prediction reference for the input frame. The controller 204 may manage the decoding operation of the source decoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0041] The decoder 210 can decode the decoded video data of a frame that can be designated as a reference image frame based on the symbols created by the source decoder 202. The operation of the decoding engine 212 can advantageously be lossy processing. When the decoded video data is decoded at a video decoder ( Figure 2A not shown), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on the reference image frame, and can cause the reconstructed reference image frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference image frame, which has the same content (in the absence of transmission errors) as the reconstructed reference image frame that will be obtained by the remote video decoder.

[0042] The predictor 206 can perform a prediction search for the decoding engine 212. That is, for a new frame to be decoded, the predictor 206 can search the reference picture memory 208 for sample data (as a candidate reference pixel block) or specific metadata such as a reference picture motion vector, block shape, etc. that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per-pixel block basis of the sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.

[0043] The outputs of all the above-mentioned functional units can undergo entropy coding in the entropy coder 214. The entropy coder 214 converts the symbols into a decoded video sequence by losslessly compressing the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0044] In some embodiments, the output of the entropy decoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the decoded video sequence created by the entropy decoder 214 in preparation for transmission via a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the decoded video data from the source decoder 202 with other data to be transmitted, such as decoded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source decoder 202 may include such data as part of the decoded video sequence. The additional data may include temporal / spatial / SNR (Signal-to-Noise Ratio) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set segments, and the like.

[0045] The controller 204 may manage the operation of the encoder components 106. During decoding, the controller 204 may assign a certain decoded picture type to each decoded picture, which may affect the decoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be decoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are familiar with those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture may be decoded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be decoded and decoded using inter prediction or intra prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, a multi-predictive picture may use more than two reference pictures and associated metadata for reconstructing a single block.

[0046] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and decoded block by block. These blocks can be predictively decoded with reference to other (already decoded) blocks, which are determined by the decoding assignments applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively decoded, or blocks of an I picture can be predictively decoded (spatial prediction or intra prediction) with reference to already decoded blocks of the same picture. Pixel blocks of a P picture can be non-predictively decoded with reference to a previously decoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be non-predictively decoded with reference to one or two previously decoded reference pictures via spatial prediction or via temporal prediction.

[0047] Video can be captured as a plurality of source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation in a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture (which is referred to as the current picture) in encoding / decoding is segmented into blocks. In a case where a block in the current picture is similar to a reference block in a previously decoded and still buffered reference picture in the video, the block in the current picture can be decoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0048] The encoder component 106 can perform decoding operations according to any predetermined video decoding technique or standard such as those described herein. In the operation of the encoder component 106, the encoder component 106 can perform various compression operations, including predictive decoding operations that utilize the temporal redundancy and spatial redundancy in the input video sequence. Thus, the decoded video data can conform to the syntax specified by the video decoding technique or standard used.

[0049] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data (e.g., via a wired connection or a wireless connection) to the display 124.

[0050] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more decoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each decoded video sequence is independent of other decoded video sequences. Each decoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as decoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the decoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the decoded video sequence. Decoder component 122 may use the additional data to more accurately decode the data and / or reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0051] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. In some embodiments, decoder component 122 is implemented at least in part in software.

[0052] The buffer memory 252 is coupled between the channel 218 and the parser 254 (e.g., to counter network jitter). In some embodiments, the buffer memory 252 is separate from the decoder component 122. In some embodiments, a separate buffer memory is provided between the output of the channel 218 and the decoder component 122. In some embodiments, in addition to the buffer memory 252 inside the decoder component 122 (e.g., which is configured to handle playback timing), a separate buffer memory is provided outside the decoder component 122 (e.g., to counter network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory 252 may not be needed, or the buffer memory 252 may be smaller. In order to use a packet network such as the Internet as much as possible, the buffer memory 252 may be needed, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not depicted) outside the decoder component 122.

[0053] The parser 254 is configured to reconstruct symbols 270 from the decoded video sequence. The symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a rendering device such as the display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set fragment (not depicted). The parser 254 parses (entropy decodes) the decoded video sequence. The decoding of the decoded video sequence may be performed according to video decoding techniques or standards, and may follow principles well known to those skilled in the art, including variable length decoding, Huffman decoding, arithmetic decoding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the decoded video sequence based on at least one parameter corresponding to a group. The subgroups may include group of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 may also extract information from the decoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0054] Depending on the type of the decoded video picture or a part thereof (e.g., inter picture and intra picture, inter block and intra block) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and the manner of involvement can be controlled by the parser 254 through subgroup control information parsed from the decoded video sequence. For the sake of brevity, this subgroup control information flow between the parser 254 and the multiple units below is not depicted.

[0055] In addition to the functional blocks already mentioned, the decoder component 122 can conceptually be subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it remains conceptually subdivided into the following functional units.

[0056] The scaler / inverse transform unit 258 receives the quantized transform coefficients as symbol 270 and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254. The scaler / inverse transform unit 258 can output a block including sample values, and the sample values can be input into the aggregator 268.

[0057] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra decoded block; that is, a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed part of the current picture. Such prediction information can be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 can generate a block of the same size and shape as the block being reconstructed using the surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 can add the prediction information already generated by the intra picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.

[0058] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion compensating the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 obtains the prediction samples can be controlled by a motion vector. The motion vector can be available to the motion compensation prediction unit 260 in the form of symbols 270, which can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of sample values such as those obtained from the reference picture memory 266 when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0059] The output samples of the aggregator 268 can undergo various loop filtering techniques in the loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the decoded video bitstream and available to the loop filter unit 256 as symbols 270 from the parser 254, but video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the decoded picture or decoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0060] The output of the loop filter unit 256 can be a sample stream that can be output to a rendering device such as the display 124, and stored in the reference picture memory 266 for use in future inter-picture prediction.

[0061] Once a decoded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once the decoded picture is fully reconstructed and the decoded picture (by, for example, the parser 254) has been identified as a reference picture, the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent decoded pictures.

[0062] The decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in any standard such as the standards described herein. As specified in a video compression technique document or standard and particularly in the profiles therein, in the sense that the decoded video sequence follows the syntax of the video compression technique or standard, the decoded video sequence can conform to the syntax specified by the video compression technique or standard used. Additionally, to conform to some video compression techniques or standards, the complexity of the decoded video sequence can be within the range defined by the levels of the video compression technique or standard. In some cases, the levels limit the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, megasamples per second), the maximum reference picture size, and the like. In some cases, the limits set by the levels can be further defined by a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the decoded video sequence.

[0063] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and / or a Data Processing Unit (DPU)). In some embodiments, the control circuitry includes one or more Field-Programmable Gate Arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application-specific integrated circuits).

[0064] The network interface 304 can be configured to interface with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication networks can be local, wide area, metropolitan area, vehicle, and industrial, real-time, delay-tolerant, etc. Examples of communication networks include: local area networks such as Ethernet, wireless LAN (Local Area Network, LAN); cellular networks including GSM (Global System for Mobile Communications, GSM), 3G (the Third Generation, 3G), 4G (the Fourth Generation, 4G), 5G (the Fifth Generation, 5G), LTE (Long Term Evolution, LTE), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicle and industrial networks including CANBus (Controller Area Network - BUS, CANBus), etc. Such communication can be only one-way receiving (e.g., broadcast TV), only one-way sending (e.g., CAN bus to certain CAN bus devices), or two-way (e.g., to other computer systems using local digital networks or wide area digital networks). Such communication can include communication to one or more cloud computing networks.

[0065] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 can include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera device, etc. The output device 308 can include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., displays or monitors), etc.

[0066] Memory 314 may include high-speed random access memory (e.g., DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more disk storage devices, optical storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). Memory 314 optionally includes one or more storage devices remote from control circuitry 302. Memory 314, or alternatively, the non-volatile solid-state memory device within memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, memory 314 or the non-transitory computer-readable storage medium of memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof:

[0067] · Operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks;

[0068] · Network communication module 318, which is used to connect server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection);

[0069] · Decoding module 320, which is used to perform various functions regarding encoding and / or decoding data such as video data. In some embodiments, decoding module 320 is an instance of decoder component 114. Decoding module 320 includes, but is not limited to, one or more of the following:

[0070] ο Decoding module 322, which is used to perform various functions regarding decoding encoded data, such as those functions previously described regarding decoder component 122; and

[0071] ο Encoding module 340, which is used to perform various functions regarding encoding data, such as those functions previously described regarding encoder component 106; and

[0072] · Picture memory 352, such as for storing pictures and picture data for use with decoding module 320. In some embodiments, picture memory 352 includes one or more of the following: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.

[0073] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described with respect to the scalar / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).

[0074] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to the source decoder 202 and / or the decoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0075] Each of the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., the instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the decoding module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as an audio processing module.

[0076] In some embodiments, server system 112 includes: a web or Hypertext Transfer Protocol (HTTP) server; a File Transfer Protocol (FTP) server; and web pages and applications implemented using Common Gateway Interface (CGI) scripts, PHP Hypertext Preprocessor (PHP), Active Server Page (ASP), Hyper Text Markup Language (HTML), Extensible Markup Language (XML), Java, JavaScript, Asynchronous JavaScript and XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), etc.

[0077] Although Figure 3 server system 112 is shown in accordance with some embodiments, Figure 3 it is intended more as a functional description of the various features that may exist in one or more server systems rather than a structural diagram of the embodiments described herein. In practice, and as would be recognized by one of ordinary skill in the art, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement server system 112 and how the features are distributed among the servers will vary depending on the implementation, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods.

[0078] Figure 4 An example current decoding block 400 including at least two color components 402 and 404 is shown in accordance with some embodiments. The GOP includes a sequence of image frames. The sequence of image frames includes a current image frame, and the current image frame further includes the current decoding block 400. The decoder 122 ( Figure 1)Receive a video bitstream 116 including a current decoded block 400 with a current image frame. The video bitstream 116 includes a syntax element for a cross-component intra prediction (CCIP) mode that indicates whether each of a plurality of samples (e.g., 406C) of a first color component 402 of the current decoded block 400 is determined based on one or more samples (e.g., 408C) of a second color component 404. The decoder 122 identifies the samples 406C of the first color component 402 in the current decoded block 406 and the samples 408C of the second color component 404 that are co-located with the samples 406C of the first color component 402. The decoder 122 also identifies at least two neighboring samples 408X of the second color component 404 in the current decoded block 400. The decoder 122 generates the samples 406C of the first color component 402 based on the at least two neighboring samples 408X of the second color component 404 and reconstructs the current decoded block 400 including the samples 406C of the first color component 402 and the samples 408C of the second color component 404.

[0079] In some embodiments, at least two neighboring samples 408X applied to determine the samples 406C of the first color component 402 have a geometric center offset from the positions of the samples of the second color component. For example, the at least two neighboring samples 408X include one or more of the neighboring samples 408-1, 408-2, and 408-3, and the offset of the geometric center is to the right of the samples 406C of the first color component 402. In another example, the at least two neighboring samples 408X include the neighboring samples 408-4 and 408-5, and the offset of the geometric center is to the left of the samples 406C of the first color component 402 and on the same row as the samples 406C of the first color component 402.

[0080] In some embodiments, the samples 406C of the first color component 402 are generated by combining the samples 408C of the second color component 404 with the at least two neighboring samples 408X according to one of the following equations:

[0081] predChromaVal = c0C + c1A1 + c2A2 +... + c N AN + c P P + c B B (1.4)

[0082] Wherein, PredChromaVal is the predicted value of sample 406C of the first color component 402; C is the value of sample 408C of the second color component 404 at the same position as sample 406C of the first color component 402; A1, A2, ……, and AN are the values of adjacent samples 408X of the second color component 404, where N is a positive integer; P is a non - linear term, for example, equal to (C * C + median luminance value) >> bit depth; B is an offset; and c0, c1, c2, ……, c N 、c P and c B are weighting factors. In some embodiments (e.g., in Equation (1.1)), the non - linear term P and the offset B are not applied to predict sample 406C of the first color component 402. Alternatively, in some embodiments (e.g., in Equation (1.2) or (1.3)), only one of the non - linear term P and the offset B is applied to predict sample 406C of the first color component 402. Alternatively, in some embodiments (e.g., in Equation (1.4)), both the non - linear term P and the offset B are applied to predict sample 406C of the first color component 402. In some embodiments, B is the median luminance value or the average luminance value of the samples of the first color component 402 in the current decoding block 400. In some embodiments, the weighting factors c0, c1, c2, ……, c N 、c P and c B are received via the video bitstream 116 for combining at least two adjacent samples 408X of the second color component 404 to generate a sample of the first color component 402.

[0083] Alternatively, in some embodiments, sample 406C of the first color component 402 is generated by combining at least two adjacent samples 408X of the second color component 404 according to one of the following equations:

[0084] predChromaVal = c1A1 + c2A2+... + c N AN + c P P + c B B (2.4) where N is a positive integer, and the sample 408C of the second color component 404 at the same position as sample 406C is not used to determine sample 406C.

[0085] In some embodiments, the first color component 402 and the second color component 404 correspond to two different color components in a set of green, blue, and red (RGB) color components. Alternatively, in some embodiments, the first color component 402 corresponds to a chrominance component (e.g., Cr, Cb), and the second color component 404 corresponds to a luminance component (e.g., Y). In some embodiments, the second color component 404 (e.g., luminance samples) has a higher resolution than the first color component 402 (e.g., chrominance samples), and at least two adjacent samples 408X of the second color component 404 are identified based on the resolution of the second color component 404.

[0086] The position of each adjacent sample 408X is identified by a displacement (e.g., 410-1, 410-2, 410-3, 410-4, or 410-5) relative to a sample 408C of the second color component 404, and the displacement is represented by at least one of a horizontal increment coordinate value and a vertical increment coordinate value. In some embodiments, at least one of the horizontal increment coordinate value and the vertical increment coordinate value of each adjacent sample 408X of the sample 408C of the second color component 404 is encoded in the video bitstream 116 and provided to the decoder 122 via the video bitstream 116. The displacement including at least one of the horizontal increment coordinate value and the vertical increment coordinate value represents a phase. Each adjacent sample 408X corresponds to a respective phase.

[0087] In some embodiments, for each adjacent sample 408X of the second color component 404, at least one of the horizontal increment coordinate value and the vertical increment coordinate value includes an integer increment coordinate value. For example, at least two adjacent samples 408X include a first adjacent sample 408-1, which is identified relative to the sample 408C by a horizontal increment coordinate value equal to 2 horizontal samples and a vertical increment coordinate value equal to 0. The first adjacent sample 408-1 is on the same line as the sample 408C. In another example, at least two adjacent samples 408X include a second adjacent sample 408-2, which is identified relative to the sample 408C by a horizontal increment coordinate value equal to 3 horizontal samples and a vertical increment coordinate value equal to 2 vertical samples. In another example, at least two adjacent samples 408X include a third adjacent sample 408-2, which is identified relative to the sample 408C by a horizontal increment coordinate value equal to 1 horizontal sample and a vertical increment coordinate value equal to -2 vertical samples. Each of the adjacent samples 408-2 and 408-3 is not on the same line or the same column as the sample of the second color component, and the position of the corresponding adjacent sample 408-2 or 408-3 is identified by both the horizontal increment coordinate value and the vertical increment coordinate value.

[0088] Alternatively, in some embodiments, at least two adjacent samples of the second color component include adjacent sample 408-4, and adjacent sample 408-4 is identified relative to sample 408C by a horizontal increment coordinate value equal to -1 2 / 3 horizontal samples and a vertical increment coordinate value equal to 2 vertical samples. In some embodiments, at least two adjacent samples of the second color component include adjacent sample 408-5, and adjacent sample 408-5 is identified relative to sample 408C by a horizontal increment coordinate value equal to -1.5 horizontal samples and a vertical increment coordinate value equal to -1.5 vertical samples. For each of adjacent samples 408-4 and 408-5, the horizontal increment coordinate value includes a fractional increment coordinate value. In some embodiments, at least one of the horizontal increment coordinate value and the vertical increment coordinate value has a given fractional precision 116A (e.g., 1 / 3, 1 / 10), and the given fractional precision 116A is extracted from the video bitstream 116 Figure 1 received from the encoder 106. Further, in some embodiments, the given fractional precision 116A is signaled as a high-level syntax selected from: sequence level parameter, GOP level parameter, picture level parameter, sub-picture level parameter, slice level parameter, tile level parameter, and maximum decoded block row level parameter.

[0089] In some embodiments, for adjacent sample 408-5 having two fractional increment coordinate values, decoder 122 identifies a set of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D of adjacent sample 408-5 of the second color component 404 based on the fractional increment coordinate values. Adjacent sample 408-4 or 408-5 is interpolated by the set of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D, and adjacent sample 408-4 or 408-5 is equal to the average of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D. The interpolated adjacent sample 408-5 is applied to generate sample 408C of the first color component 404. More specifically, in some embodiments, a set of weights is generated for the set of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D based on the distances of the set of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D from adjacent sample 408-5, and the distances are determined based on the two fractional increment coordinate values. Adjacent sample 408-5 is a weighted combination of the set of adjacent samples 408-5A, 408-5B, 408-5C, and 408-5D. Further, in some embodiments, a set of adjacent samples 408-5A through 408-5D is identified and adjacent sample 408-5 is interpolated based on a predefined filter. The predefined filter is also applied to determine samples at fractional positions within the current image frame in at least one of motion compensation processing and directional intra prediction processing.

[0090] In some embodiments, at least two adjacent samples 408X correspond to predefined candidate phases (also referred to as candidate displacements, offsets, incremental coordinate values) and are selected from a plurality of candidate phases 412. For example, the plurality of candidate phases 412 includes three candidate phases (e.g., three candidate incremental coordinate value options). The first candidate phase includes adjacent samples 408-1. The second candidate phase includes adjacent samples 408-1, 408-2, and 408-3. The third candidate phase includes adjacent samples 408-3, 408-4, and 408-5. Additionally, in some embodiments, the video bitstream 116 includes a candidate index 116B that selects one of the plurality of candidate phases 412 for identifying at least one of the horizontal incremental coordinate value and the vertical incremental coordinate value for each adjacent sample 408X of at least two adjacent samples 404X and samples 408C of the second color component 404. The decoder 122 extracts the candidate index 116B from the video bitstream 116. In some embodiments, the candidate index 116B is signaled in the video bitstream 116 at one of a superblock level, a coded block level, a prediction block level, a transform block level, and a fixed block size level. Alternatively, in some embodiments, the candidate index 116B is signaled as a high-level syntax selected from: sequence level parameters, GOP level parameters, picture level parameters, sub-picture level parameters, slice level parameters, tile level parameters, and maximum coded block row level parameters.

[0091] In some embodiments, a set of candidate phases 412A and a candidate index 116B are extracted from the video bitstream 116. The candidate index 116B selects one of the set of candidate phases 116B for identifying at least one of the horizontal incremental coordinate value and the vertical incremental coordinate value for each adjacent sample 408X of at least two adjacent samples and samples 408C of the second color component 404 of the current coded block 400. The set of candidate phases 412A is, for example, selected from the plurality of predefined candidate phases 412 during an encoding process and has a lower prediction error than the remaining candidate phases in the plurality of predefined candidate phases 412.

[0092] In some embodiments, the decoder 122 extracts a high-level flag or syntax 116C from the video bitstream 116 that indicates whether phase selection is applied. Based on the determination of applying phase selection based on the high-level flag or index 116C, at least two adjacent samples 408X of the second color component 404 are identified in the current coded block 400 for determining samples 406C of the first color component 402.

[0093] Figure 5Example scheme 500 for generating samples 406C of the first color component 402 based on samples 408C and 408X of the second color component 404 according to some embodiments. The samples 406C of the first color component 402 are generated based on the CCIP mode, in which, based on one or more samples of the second color component 404, each sample of the first color component 402 of the current decoding block 400 is determined. In the current decoding block 400, the samples 406C of the first color component 402 are co-located with the samples 408C of the second color component 404. At least two adjacent samples 408X of the second color component 404 are identified in the current decoding block 400. The position of each adjacent sample 408X is identified by at least one of a horizontal increment coordinate value and a vertical increment coordinate value relative to the sample 408C of the second color component 404. The decoder 122 generates the samples 406C of the first color component 402 based on at least two adjacent samples 408X of the second color component 404. The current decoding block 400 is reconstructed and includes the samples 406C of the first color component 402 and the samples 408C of the second color component 404.

[0094] In some embodiments, the decoder 122 identifies a reference region 502 of the current decoding block 400 in the current image frame. The reference region 502 includes one or more decoding blocks that are adjacent to and decoded before the current decoding block 400. Based on the reference region 502, the decoder 122 determines a candidate index 116B( Figure 4 ), which selects one of a plurality of candidate phases 412 for identifying at least two adjacent samples 408X and at least one of a horizontal increment coordinate value and a vertical increment coordinate value for each adjacent sample 408X of the current decoding block 400. In some embodiments, the reference region 502 includes a plurality of reference samples 406R of the first color component 402 and a plurality of reference samples 408R of the second color component 404. The decoder 122 determines a plurality of prediction errors corresponding to the plurality of candidate phases 412 based on the reference samples 406R of the first color component 402 and the reference samples 408R of the second color component 404. The candidate index 116B is identified as the selected one phase among the plurality of candidate phases 412 that provides the minimum prediction error among the plurality of prediction errors.

[0095] In some embodiments, the reference region 502 of the current decoding block 400 includes one or more of the following: upper left reference region 502TL, top reference region 502T, upper right reference region 502TR, lower left reference region 502BL, and left reference region 502L. In the example, the reference region 502 includes the top reference region 502T and the left reference region 502L. Each of the reference regions includes one or more decoding blocks.

[0096] In some embodiments, for example, based on equations (1.1) to (1.4), one or more weighting factors c0, c1, c2, ……, c N , c P and c B are used to combine at least two adjacent samples 408X of the second color component 404 to generate a sample 406C of the first color component 402. Additionally, in some embodiments, the decoder 122 identifies a reference region 502 of a current decoded block 400 in a current image frame, and the reference region 502 includes one or more decoded blocks adjacent to and decoded before the current decoded block 400. One or more weighting factors c0, c1, c2, ……, c N , c P and c B are determined based on a plurality of reference samples 406R of the first color component 402 in the reference region 502 and a plurality of reference samples 408R of the second color component 404 that are co-located with the plurality of reference samples 406R of the first color component 402. Additionally, in some embodiments, the decoder 122 determines one or more weighting factors c0, c1, c2, ……, c N , c P and c B by determining a least mean square (LMS) value based on the plurality of reference samples 406R of the first color component 402 and the plurality of reference samples 408R of the second color component 404. The one or more weighting factors c0, c1, c2, ……, c N , c P and c B are iteratively adjusted to reduce the LMS value until the LMS value meets a predefined criterion (e.g., minimizing the LMS value or being below an LMS threshold).

[0097] In some embodiments, a set of candidate phases 412A and candidate indices 116B are extracted from the video bitstream 116. The candidate index 116B selects one of a set of candidate phases 116B for identifying at least one of a horizontal increment coordinate value and a vertical increment coordinate value for each adjacent sample 408X of at least two adjacent samples 408X and a sample 408C of the second color component 404 for the current decoded block 400. The set of candidate phases 412A is selected, for example, based on the reference region of the current decoded block 400 and from a plurality of predefined candidate phases 412 during an encoding process, and has a lower prediction error than the remaining candidate phases among the plurality of predefined candidate phases 412.

[0098] In addition, in some embodiments, for the current decoding block 400, the reference region 502 of the current decoding block 400 includes one or more of the following: the top-left reference region 502TL, the top reference region 502T, the top-right reference region 502TR, the bottom-left reference region 502BL, and the left reference region 502L of the current decoding block 400. Each of the reference regions includes one or more decoding blocks.

[0099] Figure 6 FIG. 600 is a flowchart illustrating an example method 600 for decoding video according to some embodiments. The method 600 may be performed at a computing system (e.g., the server system 112, the source device 102, or the electronic device 120), which has control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, the method 600 is applied in conjunction with one or more video codecs, including but not limited to H.264, H.265 / HEVC, H.266 / VVC, AV1, and AVS / AVS2 / AVS3. In some embodiments, the method 600 is performed by executing instructions stored in the memory of the computing system (e.g., the decoding module 320 of the memory 314). The decoder 122 applies one or more phases (or displacements or offsets) of the samples 408C of the second color component 404 as inputs to perform cross-component prediction. The phase (or displacement, or offset) refers to the relative position with respect to the coordinates of the co-located samples 408C of the second color component 404, which is used to obtain another neighboring sample 408X of the second color component 404. The neighboring sample 408X is used as an input (instead of the collocated sample) to perform cross-component prediction. Using the samples 408C of the second color component 404 and / or at least two neighboring samples 408X, the sample 406C of the first color component 402 co-located with the sample 408C is determined based on a linear or non-linear function (e.g., any one of equations (1.1) to (1.4)). In some embodiments, the reference samples 406R of the first color component 402 and the reference samples 408R of the second color component 404 in the reference region 502 of the current decoding block 400 are used to derive the parameters (e.g., weighting factors, offsets) of the specified linear or non-linear function. Alternatively, in some embodiments, at least a subset of the parameters of the specified linear or non-linear function is explicitly signaled via the video bitstream 116.

[0100] In some embodiments, the plurality of phases includes a plurality of candidate phases 412( Figure 4 ) and are defined as one or more integer increment coordinate values to identify the neighboring sample 408X with reference to the co-located sample 408C of the second color component 404 (e.g., Figure 4(i) the horizontal coordinate value, or (ii) the vertical coordinate value, or (iii) both the horizontal coordinate value and the vertical coordinate value of 408-1 to 408-3 therein.

[0101] In some embodiments, the plurality of phases includes a plurality of candidate phases 412 ( Figure 4 ) and is defined as one or more fractional increment coordinate values to identify adjacent samples 408X with reference to the co-located samples 408C of the second color component 404 (e.g., Figure 4 (i) the horizontal coordinate value, or (ii) the vertical coordinate value, or (iii) both the horizontal coordinate value and the vertical coordinate value of 408-4 and 408-5 therein. Additionally, in some embodiments, an interpolation filter is applied to obtain a sample 408X at a coordinate with a fractional increment coordinate value located above the juxtaposed sample 408C of the second color component 404.

[0102] In some embodiments, the interpolation filter can be the same set of filters used to obtain sample values at fractional positions used in different image / video decoding processes (e.g., motion compensation processing or directional intra prediction processing).

[0103] In some embodiments, the plurality of phases is applied to samples 408C and / or 408X at the original resolution of the second color component 404 (i.e., without downsampling the second color component 404).

[0104] In some embodiments, the selection of a phase (e.g., the positions of at least two adjacent samples 408X) from the plurality of candidate phases 412 for encoding / decoding the current decoding block 400 is signaled in the video bitstream 116 (e.g., candidate index 116B). In other words, the video bitstream 116 includes a candidate index 116B that selects one of the plurality of candidate phases 412 for generating the sample 406C of the first color sample 402. In some embodiments, the candidate index 116B is signaled at a block level, which includes but is not limited to a superblock level, a decoding block level, a prediction block level, a transform block level, and / or a fixed block size level. In some embodiments, the candidate index 116B is signaled as a high-level syntax (HLS), and the high-level syntax includes but is not limited to sequence-level parameters, GOP-level parameters, picture-level parameters, sub-picture-level parameters, slice-level parameters, tile-level parameters, and / or maximum decoding block row-level parameters. Conversely, in some embodiments, the selection of a phase (e.g., the positions of at least two adjacent samples 408X) from the plurality of candidate phases 412 for encoding / decoding the current decoding block 400 is implicitly derived without explicit signaling in the video bitstream 116.

[0105] In some embodiments, the reference samples 406R of the first color component and the reference samples 408R of the second color component in the reference region 502 ( Figure 5 ) of the current decoding block 400 are used to derive the selected phase (e.g., the positions of at least two adjacent samples 408X) of the sample 408C of the current decoding block 400. In the example, the reference region 502 includes a left reference region 502L and a top reference region 502T.

[0106] In some embodiments, the reference samples 406R of the first color component and the reference samples 408R of the second color component in the reference region 502 ( Figure 5 ) of the current decoding block 400 are used to evaluate a plurality of candidate phases 412. The evaluation is completed by predicting the samples of the first color component 402 of the reference region 502 using the reference samples 408R of the second color component 404 of the reference region 502. The predicted samples are compared with the reference samples 406R of the first color component 402 of the reference region 502, and for each sample of the first color component 402, the prediction error is determined to be equal to the difference between the predicted sample and the reference sample 406R. The selected candidate phase (e.g., the positions of at least two adjacent samples 408X) provides the minimum prediction error (measured by a given error metric such as SAD or SSE) and is selected for encoding and decoding the current decoding block 400 including the samples 406C and 408C.

[0107] In some embodiments, based on the same function used for applying cross-component prediction, the adjacent reconstructed co-located samples 408X of the second color component 404 are used to predict the adjacent reconstructed samples of the first color component 402 of the current decoding block 400. However, the parameters (c0, c1, c2, ……, c N , c P or c B ) used in the function are derived using the adjacent reference samples 406R of the current decoding block 400 of the first color component 402 and the adjacent reference samples 408R of the co-located block of the second color component 404 (e.g., using the least mean square approximation).

[0108] In some embodiments, multiple candidate phases 412 are evaluated using neighboring reconstructed samples of a co-located block of the second color component 404. The evaluation is done by predicting neighboring reconstructed samples (e.g., top neighboring reconstructed sample and left neighboring reconstructed sample) of the current decoded block 400 of the first color component 402 using neighboring reconstructed samples of a co-located block of the second color component 404 (e.g., top neighboring reconstructed sample and left neighboring reconstructed sample), and a subset of candidate phases 412A that provide less prediction error (measured by a given error metric such as SAD or SSE) than other candidate phases is marked as the most likely phase for encoding and decoding the current decoded block 400, and the index of the selected phase 116B in this subset is signaled into the bitstream 116 and parsed by the decoder 122.

[0109] In some embodiments, the selected phase (e.g., the positions of at least two neighboring samples 408X) is explicitly signaled. In some embodiments, multiple candidate phase values are predefined and locally stored by the decoder 122, and the index of the selected phase (i.e., candidate index 116B) is signaled.

[0110] In some embodiments, the horizontal and / or vertical components of the phase value are signaled with a given fractional precision 116A. In some embodiments, the fractional precision 116A is signaled with a high-level syntax, which includes but is not limited to sequence-level parameters, GOP-level parameters, picture level, sub-picture level, slice level, tile level, or maximum decoded block row level.

[0111] In some embodiments, the selected phase (e.g., the positions of at least two neighboring samples 408X) is signaled at a block level, which includes but is not limited to super-block level, decoded block level, prediction block level, transform block level, or fixed block size level.

[0112] In some embodiments, a high-level flag or high-level syntax (HLS) is introduced, e.g., at the frame level and using a phase selection flag or syntax 116C, to indicate whether phase selection is enabled. If the HLS indicates that the function is not enabled, a filter without phase shift is applied.

[0113] Now turning to some example embodiments.

[0114] (A1)In some implementations, a method 600 for decoding video data is implemented. The method 600 includes: receiving (602) a video bitstream including a current decoded block of a current image frame, where the video bitstream includes (604) a syntax element for a cross-component intra prediction (CCIP) mode, and the syntax element indicates whether each sample of a first color component of the current decoded block is determined based on one or more samples of a second color component. The method 600 further includes: identifying (606) samples of the first color component and samples of the second color component at the same positions as the samples of the first color component in the current decoded block. The method 600 further includes: for each of two or more adjacent samples of the second color component, obtaining (608) at least one of (i) a horizontal increment coordinate value and (ii) a vertical increment coordinate value with respect to the sample of the second color component from the video bitstream. The method further includes: identifying (610) each of at least two adjacent samples of the second color component in the current decoded block based on at least one of (i) the horizontal increment coordinate value and (ii) the vertical increment coordinate value. The method 600 further includes: generating (612) samples of the first color component based on at least two adjacent samples of the second color component, and reconstructing (614) the current decoded block based at least on the samples of the first color component generated according to at least two adjacent samples of the second color component.

[0115] (A2)In some embodiments of A1, wherein the at least two adjacent samples have a geometric center offset from the position of the sample of the second color component.

[0116] (A3)In some embodiments of A1 or A2, wherein for each adjacent sample of the second color component, at least one of the horizontal increment coordinate value and the vertical increment coordinate value includes an integer increment coordinate value.

[0117] (A4)In some embodiments of any one of A1 to A3, wherein the at least two adjacent samples of the second color component include first adjacent samples located in the same row or the same column as the sample of the second color component, and the position of the first adjacent sample is identified by one of the horizontal increment coordinate value and the vertical increment coordinate value.

[0118] (A5)In some embodiments of any one of A1 to A3, wherein the at least two adjacent samples of the second color component include first adjacent samples not located in the same row or the same column as the sample of the second color component, and the position of the first adjacent sample is identified by both the horizontal increment coordinate value and the vertical increment coordinate value.

[0119] (A6) In some embodiments of A1 or A2, wherein at least two adjacent samples of the second color component include a first adjacent sample, and at least one of the horizontal increment coordinate value and the vertical increment coordinate value includes a fractional increment coordinate value.

[0120] (A7) In some embodiments of A6, method 600 further includes, for the first adjacent sample and based on the fractional increment coordinate value: identifying a set of adjacent samples of the first adjacent sample of the second color component; and interpolating the first adjacent sample through the set of adjacent samples, and the interpolated first adjacent sample is applied to generate a sample of the first color component.

[0121] (A8) In some embodiments of A7, wherein identifying the set of adjacent samples, and interpolating the first adjacent sample based on a predefined filter, and method 600 further includes: applying the predefined filter in at least one of motion compensation processing and directional intra prediction processing to determine a sample located at a fractional position of the current image frame.

[0122] (A9) In some embodiments of any one of A1 to A8, wherein the second color component has a higher resolution than the first color component, and at least two adjacent samples of the second color component are identified based on the resolution of the second color component.

[0123] (A10) In some embodiments of any one of A1 to A9, method 600 further includes: extracting a candidate index from the video bitstream, the candidate index selecting one of a plurality of candidate phases for identifying at least two adjacent samples and at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the samples of the second color component.

[0124] (A11) In some embodiments of A10, wherein for the current decoding block, the candidate index is signaled in the video bitstream at one of a superblock level, a decoding block level, a prediction block level, a transform block level, and a fixed block size level.

[0125] (A12) In some embodiments of A10, wherein the candidate index is signaled as a high-level syntax selected from: sequence level parameter, GOP level parameter, picture level parameter, sub-picture level parameter, slice level parameter, tile level parameter, and maximum decoding block row level parameter.

[0126] (A13)In some embodiments of any one of A1 to A12, method 600 further includes: identifying a reference region of the current decoded block in the current image frame, wherein the reference region includes one or more decoded blocks adjacent to and decoded before the current decoded block; and based on the reference region, determining a candidate index that selects one of a plurality of candidate phases for identifying at least one of the horizontal increment coordinate value and the vertical increment coordinate value of each adjacent sample of the at least two adjacent samples and the samples of the second color component for the current decoded block.

[0127] (A14)In some embodiments of A13, wherein the reference region includes a plurality of reference samples of the first color component and a plurality of reference samples of the second color component, and determining the candidate index further includes: determining a plurality of prediction errors corresponding to the plurality of candidate phases based on the reference samples of the first color component and the reference samples of the second color component; and identifying the candidate index that represents the selected one of the plurality of candidate phases that provides the minimum prediction error among the plurality of prediction errors.

[0128] (A15)In some embodiments of any one of A1 to A14, wherein one or more weighting factors are used to combine at least two adjacent samples of the second color component to generate a sample of the first color component.

[0129] (A16)In some embodiments of A15, method 600 further includes: identifying a reference region of the current decoded block in the current image frame, wherein the reference region includes one or more decoded blocks adjacent to and decoded before the current decoded block; and determining the one or more weighting factors based on the plurality of reference samples of the first color component and the plurality of reference samples of the second color component at the same positions as the plurality of reference samples of the first color component in the reference region.

[0130] (A17)In some embodiments of A16, wherein determining the one or more weighting factors further includes: determining a least mean square (LMS) value based on the plurality of reference samples of the first color component and the plurality of reference samples of the second color component; and iteratively adjusting the one or more weighting factors to reduce the LMS value until the LMS value meets a predefined criterion.

[0131] (A18)In an implementation of any one of A13, A14, and A16, wherein, for the current decoding block, the reference region of the current decoding block includes one or more of the following: the upper left reference region, the top reference region, the upper right reference region, the lower left reference region, and the left reference region of the current decoding block.

[0132] (A19)In some implementations of any one of A1 to A18, method 600 further includes: extracting a set of candidate phases and candidate indices from the video bitstream, the candidate indices selecting one of the set of candidate phases for identifying at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the at least two adjacent samples and the samples of the second color component for the current decoding block; wherein the set of candidate phases is selected from a plurality of predefined candidate phases and has a lower prediction error than the remaining candidate phases in the plurality of predefined candidate phases.

[0133] (A20)In some implementations of any one of A1 to A19, wherein at least one of the horizontal increment coordinate value and the vertical increment coordinate value has a given fractional precision, method 600 further includes: extracting the given fractional precision from the video bitstream.

[0134] (A21)In some implementations of A20, wherein the given fractional precision is signaled as a high-level syntax selected from the following: sequence-level parameter, GOP-level parameter, picture-level parameter, sub-picture-level parameter, slice-level parameter, tile-level parameter, and maximum decoding block row-level parameter.

[0135] (A22)In some implementations of any one of A1 to A21, method 600 further includes: extracting a high-level flag or syntax indicating whether to apply phase selection from the video bitstream; wherein, based on a determination of applying the phase selection based on the high-level flag or index, at least two adjacent samples of the second color component are identified in the current decoding block for determining the samples of the first color component.

[0136] (A23)In some implementations of any one of A1 to A22, wherein the video bitstream further includes one or more weighting factors for combining at least two adjacent samples of the second color component to generate the samples of the first color component.

[0137] (A24)In some implementations of any one of A1 to A23, wherein the video bitstream further includes at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the current decoding block.

[0138] (A25)In some embodiments of any one of A1 to A24, wherein the first color component and the second color component correspond to two different color components among a set of green, blue, and red color components.

[0139] (A26)In some embodiments of any one of A1 to A24, wherein the first color component corresponds to a chrominance component and the second color component corresponds to a luminance component.

[0140] (A27)In some embodiments of any one of A1 to A26, wherein generating a sample of the first color component based on at least two adjacent samples of the second color component further includes: combining one or more adjacent samples of the second color component with at least one of the following based on a plurality of weighting factors: (1) a sample of the second color component, (2) a non - linear term of a subset of one or more adjacent samples of the second color component and a sample of the second color component, and (3) a bias term.

[0141] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 to A27 above).

[0142] In yet another aspect, some embodiments include a non - transitory computer - readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A27 above).

[0143] The proposed methods can be used alone or in any combination. Additionally, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non - transitory computer - readable medium. Hereinafter, the term block can be interpreted as a prediction block, a decoding block, or a decoding unit, i.e., a CU.

[0144] It will be understood that although terms such as "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0145] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. Unless the context clearly indicates otherwise, as used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0146] As used herein, the term "if" can be interpreted, depending on the context, as meaning "when the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrases "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" can be interpreted, depending on the context, as meaning "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true".

[0147] For purposes of illustration, the foregoing description has been made with reference to particular embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Given the above teachings, many modifications and variations are possible. The embodiments were chosen and described in order to best illustrate the operating principles and practical applications, so as to enable others skilled in the art to implement them.

Claims

1. A method for decoding video data, comprising: Receiving a video bitstream including a current decoded block of a current picture frame, wherein the video bitstream includes a syntax element for a cross-component intra prediction (CCIP) mode, the syntax element indicating whether each sample of a first color component of the current decoded block is determined based on one or more samples of a second color component; Identifying samples of the first color component and samples of the second color component at the same positions as the samples of the first color component in the current decoded block; For each of two or more adjacent samples of the second color component, obtaining at least one of (i) a horizontal increment coordinate value and (ii) a vertical increment coordinate value relative to the sample of the second color component from the video bitstream; Identifying each of at least two adjacent samples of the second color component in the current decoded block based on at least one of (i) the horizontal increment coordinate value and (ii) the vertical increment coordinate value; Generating samples of the first color component based on at least two adjacent samples of the second color component; and Reconstructing the current decoded block based at least on the samples of the first color component generated according to at least two adjacent samples of the second color component.

2. The method according to claim 1, wherein The at least two adjacent samples have a geometric center offset from the position of the sample of the second color component.

3. The method according to claim 1, wherein For each adjacent sample of the second color component, at least one of the horizontal increment coordinate value and the vertical increment coordinate value includes an integer increment coordinate value.

4. The method according to claim 1, wherein The at least two adjacent samples of the second color component include first adjacent samples located in the same row or the same column as the sample of the second color component, and the position of the first adjacent sample is identified by one of the horizontal increment coordinate value and the vertical increment coordinate value.

5. The method according to claim 1, wherein The at least two adjacent samples of the second color component include first adjacent samples not located in the same row or the same column as the sample of the second color component, and the position of the first adjacent sample is identified by both the horizontal increment coordinate value and the vertical increment coordinate value.

6. The method according to claim 1, wherein, The at least two adjacent samples of the second color component include first adjacent samples, and at least one of the horizontal increment coordinate value and the vertical increment coordinate value includes a fractional increment coordinate value.

7. The method according to claim 6, further comprising: For the first adjacent sample and based on the fractional increment coordinate value: Identifying a set of adjacent samples of the first adjacent sample of the second color component; And Interpolating the first adjacent sample through the set of adjacent samples, and the interpolated first adjacent sample is applied to generate samples of the first color component.

8. The method according to claim 7, wherein Identifying the set of adjacent samples and interpolating the first adjacent sample based on a predefined filter, the method further comprising: Applying the predefined filter in at least one of motion compensation processing and directional intra prediction processing to determine samples located at fractional positions in the current picture frame.

9. The method according to claim 1, wherein The second color component has a higher resolution than the first color component, and at least two adjacent samples of the second color component are identified based on the resolution of the second color component.

10. The method according to claim 1, further comprising: extracting a candidate index from the video bitstream, the candidate index selecting one of a plurality of candidate phases for identifying at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the at least two adjacent samples and the samples of the second color component.

11. The method according to claim 10, wherein, For the current decoding block, signaling the candidate index in the video bitstream at one of a superblock level, a decoding block level, a prediction block level, a transform block level, and a fixed block size level.

12. The method according to claim 10, wherein, Signaling the candidate index as a high-level syntax selected from: sequence level parameters, GOP level parameters, picture level parameters, sub-picture level parameters, slice level parameters, tile level parameters, and maximum decoding block row level parameters.

13. A computing system, comprising: control circuitry; and a memory storing one or more programs configured to be executed by the control circuitry, the one or more programs further comprising instructions for: receiving a video bitstream including a current decoding block of a current image frame, wherein the video bitstream includes syntax elements for cross-component intra prediction (CCIP) mode, the syntax elements indicating whether each sample of the first color component of the current decoding block is determined based on one or more samples of a second color component; identifying samples of the first color component and samples of the second color component co-located with the samples of the first color component in the current decoding block; for each of two or more adjacent samples of the second color component, obtaining at least one of (i) a horizontal increment coordinate value and (ii) a vertical increment coordinate value relative to the sample of the second color component from the video bitstream; identifying each of at least two adjacent samples of the second color component in the current decoding block based on at least one of (i) the horizontal increment coordinate value and (ii) the vertical increment coordinate value; generating samples of the first color component based on at least two adjacent samples of the second color component; and reconstructing the current decoding block based at least on the samples of the first color component generated according to at least two adjacent samples of the second color component.

14. The computing system according to claim 13, the one or more programs further comprising instructions for: Identify a reference region of the current decoding block in the current image frame, where The reference region includes one or more decoding blocks adjacent to and decoded before the current decoding block; and determining a candidate index based on the reference region, the candidate index selecting one of a plurality of candidate phases for identifying at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the two or more adjacent samples and the samples of the second color component of the current decoding block.

15. The computing system according to claim 14, wherein, The reference region includes a plurality of reference samples of the first color component and a plurality of reference samples of the second color component. Determining the candidate index further includes: Determining a plurality of prediction errors corresponding to the plurality of candidate phases based on the reference samples of the first color component and the reference samples of the second color component; and Identifying the candidate index, where the candidate index represents a selected one of the plurality of candidate phases that provides the minimum prediction error among the plurality of prediction errors.

16. The computing system according to claim 13, wherein, Using one or more weighting factors to combine two or more adjacent samples of the second color component to generate a sample of the first color component.

17. The computing system according to claim 16, wherein the one or more programs further include instructions for performing the following operations: Identify the reference region of the current decoding block in the current image frame, wherein, The reference region includes one or more decoded blocks that are adjacent to and prior to the current decoded block. And Determining the one or more weighting factors based on the plurality of reference samples of the first color component in the reference region and the plurality of reference samples of the second color component that are co-located with the plurality of reference samples of the first color component.

18. The computing system according to claim 17, wherein, Determining the one or more weighting factors further includes: Determining a least mean square (LMS) value based on the plurality of reference samples of the first color component and the plurality of reference samples of the second color component; and Iteratively adjusting the one or more weighting factors to reduce the LMS value until the LMS value meets a predefined criterion.

19. The computing system according to claim 13, wherein the one or more programs further include instructions for performing the following operations: Extracting from the video bitstream a set of candidate phases and candidate indices, where the candidate index selects one of the set of candidate phases for identifying at least one of the horizontal increment coordinate value and the vertical increment coordinate value for each adjacent sample of the two or more adjacent samples and the sample of the second color component of the current decoded block; Among them, The set of candidate phases is selected from a plurality of predefined candidate phases and has a lower prediction error than the remaining candidate phases among the plurality of predefined candidate phases.

20. The computing system according to claim 19, wherein, For the current decoded block, the reference region of the current decoded block includes one or more of the following: the upper left reference region, the top reference region, the upper right reference region, the lower left reference region, and the left reference region of the current decoded block.

21. A non-transitory computer-readable storage medium storing one or more programs for execution by a control circuitry of a computing system, the one or more programs including instructions for performing the following operations: Receiving a video bitstream including a current decoded block of a current picture frame, wherein, The video bitstream includes syntax elements for cross-component intra prediction (CCIP) mode that indicate whether each sample of the first color component of the current decoded block is determined based on one or more samples of the second color component; Identifying samples of the first color component and samples of the second color component that are co-located with the samples of the first color component in the current decoded block; For each of two or more adjacent samples of the second color component, obtain at least one of (i) a horizontal incremental coordinate value and (ii) a vertical incremental coordinate value with respect to the sample of the second color component from the video bitstream; Identify each of at least two adjacent samples of the second color component in the current decoding block based on at least one of (i) the horizontal incremental coordinate value and (ii) the vertical incremental coordinate value; Generate samples of the first color component based on at least two adjacent samples of the second color component; and Reconstruct the current decoding block based at least on the samples of the first color component generated according to at least two adjacent samples of the second color component.

22. The non-transitory computer-readable storage medium according to claim 21, wherein, At least one of the horizontal incremental coordinate value and the vertical incremental coordinate value has a given fractional precision, and the one or more programs further include instructions for: Extracting the given fractional precision from the video bitstream.

23. The non-transitory computer-readable storage medium according to claim 22, wherein, Signaling the given fractional precision as a high-level syntax selected from: sequence-level parameter, GOP-level parameter, picture-level parameter, sub-picture-level parameter, slice-level parameter, tile-level parameter, and maximum decoding block row-level parameter.

24. The non-transitory computer-readable storage medium according to claim 21, wherein the one or more programs further include instructions for performing the following operations: Extract a high-level flag or syntax indicating whether phase selection is applied from the video bitstream; Among them, Identify two or more adjacent samples of the second color component in the current decoding block for determining samples of the first color component according to a determination of applying the phase selection based on the high-level flag or index.

25. The non-transitory computer-readable storage medium according to claim 21, wherein, The video bitstream further includes one or more weighting factors for combining two or more adjacent samples of the second color component to generate samples of the first color component.

26. The non-transitory computer-readable storage medium according to claim 21, wherein, The video bitstream further includes at least one of the horizontal incremental coordinate value and the vertical incremental coordinate value for each adjacent sample of the current decoding block.

27. The non-transitory computer-readable storage medium according to claim 21, wherein, The first color component and the second color component correspond to two different color components in a set of green, blue, and red color components.

28. The non-transitory computer-readable storage medium according to claim 21, wherein, The first color component corresponds to a chrominance component, and the second color component corresponds to a luminance component.

29. The non-transitory computer-readable storage medium according to any one of claims 21, wherein, Generating samples of the first color component based on two or more adjacent samples of the second color component further includes: Combining two or more adjacent samples of the second color component with at least one of the following based on a plurality of weighting factors: (1) samples of the second color component, (2) non-linear terms of a subset of two or more adjacent samples and samples of the second color component of the second color component, and (3) a bias term.