Video encoding and decoding method, medium and system

By employing subsampling technology with cross-component prediction modes in video coding, hardware complexity and memory requirements are reduced, encoding and decoding efficiency is improved, and the problems of high hardware complexity and large memory requirements in existing technologies are solved, achieving a more efficient encoding and decoding process.

CN122069356APending Publication Date: 2026-05-19TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2025-11-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video coding technologies suffer from high hardware complexity and large memory requirements when processing video data, especially when using cross-component prediction modes, where the computational complexity and memory requirements of reference samples are high, affecting encoding and decoding efficiency.

Method used

The cross-component prediction mode (CCP) is adopted, which reduces hardware complexity and CCP parameter derivation complexity by subsampling the reference region. It uses a multi-tap model and subsampling techniques to selectively reduce the number of reference samples, and applies irregular sampling methods, pooling and low-pass filtering to optimize reference sample selection.

Benefits of technology

It reduces the computational complexity and hardware requirements of the encoding and decoding process, improves encoding and decoding efficiency, shortens encoding and decoding time, enhances the scalability of resource-limited devices, and significantly reduces complexity while having a minimal impact on video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069356A_ABST
    Figure CN122069356A_ABST
Patent Text Reader

Abstract

The invention provides a video encoding and decoding method, a medium and a computing system. A video decoding method includes receiving a video bitstream including a plurality of blocks, the plurality of blocks including a current block. The method also includes determining that a multi-hypothesis cross-component prediction (MHCCP) mode is enabled for the current block, and identifying a reference region for the MHCCP mode. The method further includes subsampling the reference region, and applying an MHCCP mode to the current block using the subsampled reference region.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority and related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 722,547, entitled “Simplification of Model Parameter Derivation,” filed November 19, 2024, and U.S. Patent Application No. 19 / 366,481, filed October 22, 2025, which are incorporated herein by reference in their entirety. Technical Field

[0003] The disclosed embodiments generally relate to video encoding and decoding, including but not limited to video encoding and decoding methods, media, and systems. Background Technology

[0004] Digital video is supported by a variety of electronic devices such as digital televisions, laptops or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing equipment, and video streaming devices. Electronic devices transmit and receive digital video data via communication networks or otherwise transmit digital video data, and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video encoding can be used to compress video data according to at least one video encoding standard before transmission or storage. Video encoding can be performed by hardware and / or software on electronic / client devices or servers providing cloud services.

[0005] Video coding typically uses prediction methods that leverage the inherent redundancy in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form using a lower bitrate while avoiding or minimizing degradation in video quality. Several video codec standards have been developed. High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (Revision 1), 2014 (Revision 2), 2015 (Revision 3), and 2016 (Revision 4). Universal Video Coding (VVC / H.266) is a video compression standard designed as a successor to HEVC. The ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (Revision 1) and 2022 (Revision 2). The Open Media Alliance Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, a confirmed version 1.0.0 with errata table 1 was released. Summary of the Invention

[0006] As mentioned above, encoding (compression) reduces bandwidth and / or storage space requirements. Both lossless and lossy compression can be used, as detailed later. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal via a decoding process. Lossy compression refers to an encoding / decoding process where the original video information is not fully preserved during encoding and is not fully recovered during decoding. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signals is small enough to make the reconstructed signal useful for the intended application. The amount of distortion tolerated depends on the application. For example, users of some consumer video streaming applications may tolerate higher distortion than users of film or television broadcasting applications. The compression ratio achievable by a particular encoding algorithm can be selected or adjusted to reflect various distortion tolerances: higher tolerance for distortion generally allows encoding algorithms that produce higher loss and higher compression ratios.

[0007] Among other things, this disclosure describes prediction of video data using a cross-component prediction (CCP) mode, wherein each of a plurality of samples of a second color component of a current coded block is determined based on at least one associated sample of a first color component of a reference block. This CCP mode can use a multi-tap model that includes multiple taps. As an example, subsampling of the reference samples can be used for the CCP mode (e.g., a multi-hypothesis cross-component prediction (MHCCP) mode). Subsampling of the reference region can reduce hardware complexity (e.g., smaller buffer size) and can reduce the complexity of CCP parameter derivation.

[0008] According to some embodiments, a video decoding method includes: (i) receiving a video bitstream comprising a plurality of blocks, the plurality of blocks including a current block; (ii) determining that an MHCCP mode is enabled for the current block; (iii) identifying a reference region for the MHCCP mode; (iv) subsampling the reference region; and (v) applying the MHCCP mode to the current block using the subsampled reference region.

[0009] According to some embodiments, a video coding method includes: (i) receiving video data comprising a plurality of blocks, the plurality of blocks including a current block; (ii) determining that an MHCCP mode is enabled for the current block; (iii) identifying a reference region for the MHCCP mode; (iv) subsampling the reference region; and (v) applying the MHCCP mode to the current block using the subsampled reference region.

[0010] According to some embodiments, a computing system, such as a streaming system, server system, personal computer system, or other electronic device, is provided. The computing system includes control circuitry and a memory storing one or more sets of instructions. The set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes encoder components and decoder components (e.g., a transcoder). A computing system is also provided, including: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code to perform a method according to an embodiment of this application. A method for storing a video stream, generating a video stream according to a video encoding method, and storing the video stream is also provided.

[0011] According to some embodiments, a non-volatile computer-readable storage medium is provided. This non-volatile computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.

[0012] Therefore, apparatus and systems having methods for encoding and decoding video are disclosed. Such methods, apparatus, and systems can complement or replace conventional methods, devices, and systems for encoding and decoding video.

[0013] The features and advantages described in the specification are not necessarily exhaustive, and in particular, some additional features and advantages will be apparent to those skilled in the art in light of the accompanying drawings, specification, and claims provided in this disclosure. Furthermore, it should be noted that the language used in the specification has been chosen primarily for readability and instruction purposes and is not necessarily intended to depict or limit the subject matter described herein. Attached Figure Description

[0014] To gain a more detailed understanding of this disclosure, reference can be made to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings only illustrate relevant features of this disclosure and are therefore not necessarily to be considered limiting, as those skilled in the art will understand upon reading this disclosure that other valid features may be permitted.

[0015] Figure 1 This is a block diagram illustrating an example communication system according to some embodiments.

[0016] Figure 2A This is a block diagram illustrating example elements of an encoder component according to some embodiments.

[0017] Figure 2B This is a block diagram illustrating example elements of a decoder component according to some embodiments.

[0018] Figure 3 This is a block diagram illustrating an example server system according to some embodiments.

[0019] Figure 4 An example scheme for generating a first chromaticity sample from at least one luminance sample in CCP mode is illustrated according to some embodiments.

[0020] Figure 5A This is a diagram of an example image frame of the current block located at the top boundary of the superblock, according to some embodiments.

[0021] Figure 5B This is a diagram of an example image frame of the current block located at the left boundary of the superblock, according to some embodiments.

[0022] Figure 5C An example reference area for encoding blocks is illustrated according to some embodiments.

[0023] Figures 5D to 5I The illustration shows an example subsampling of an example reference area according to some embodiments.

[0024] Figure 6A This is a diagram showing the application of multiple filter shapes to the current block according to some embodiments.

[0025] Figure 6B The illustration shows example syntax when the MHCCP flag is enabled, according to some embodiments.

[0026] Figure 7A The illustration shows an example video decoding method according to some embodiments.

[0027] Figure 7B The illustration shows an example video encoding method according to some embodiments.

[0028] According to conventional practice, the various features illustrated in the accompanying drawings need not be drawn to scale, and similar reference numerals may be used to indicate similar features throughout the specification and the accompanying drawings. Detailed Implementation

[0029] This disclosure describes a video compression method using intra-frame prediction and inter-frame prediction. Samples of the current coding block can be reconstructed from samples of a reference coding block based on a model with multiple model parameters. For example, when it is determined that CCP mode is enabled for the current block, a reference region can be identified for that CCP mode. This reference region can be subsampled, and the CCP mode can be applied to the current block using the subsampled reference region. In this way, the complexity and memory requirements of using reference samples can be reduced.

[0030] Among other things, this disclosure describes a set of methods for simplifying model parameter derivation in video coding, particularly within MHCCP modes. In some embodiments, subsampling techniques are applied to a reference region used for model parameter calculation. This is achieved by selectively reducing the number of reference samples (e.g., by employing...). Subsampling rates, cloverleaf downsampling, or changing the subsampling rate across different rows, columns, or regions can simplify the computational complexity and hardware requirements of both the encoding and decoding processes. For example, more aggressive subsampling can be applied to regions farther from the current block, or the subsampling ratio can be adjusted based on block shape and size, ensuring efficient use of memory and processing resources. Furthermore, irregular sampling methods, pooling, and / or low-pass filtering can be applied to further optimize reference sample selection. These techniques collectively result in smaller buffer sizes, reduced memory bandwidth, and faster parameter derivation, leading to more efficient hardware implementations and reduced power consumption. These benefits manifest in improved encoding and decoding efficiency, reduced encoding and decoding times, and enhanced scalability for resource-constrained devices, as simulation results (e.g., Table 1 below) demonstrate, achieving a significant reduction in complexity with minimal impact on video quality.

[0031] Figure 1 This is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) that are communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is, for example, a streaming system used with video-enabled applications such as video conferencing applications, digital television applications, and media storage devices and / or distribution applications.

[0032] Source device 102 includes a video source 104 (e.g., a camera component or media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates at least one encoded video stream from the video stream. The video stream from video source 104 can have a high data volume compared to the encoded video stream 108 generated by encoder component 106. Because the encoded video stream 108 has a lower data volume (less data) compared to the video stream from video source 104, it requires less bandwidth for transmission and less storage space for storage. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to at least one network 110).

[0033] At least one network 110 represents any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired (wired) and / or wireless communication networks. At least one network 110 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. At least one network 110 includes server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content such as encoded video streams from source device 102). Server system 112 includes encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, encoder component 114 includes encoder components and / or decoder components. In various embodiments, encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, encoder component 114 is configured to decode the encoded video stream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings from the encoded video stream 108.

[0034] In some embodiments, server system 112 serves as a media-aware network element (MANE). For example, server system 112 may be configured to trim encoded video streams 108 to cut potentially different streams for at least one of the electronic devices 120. In some embodiments, the MANE is provided separately from server system 112.

[0035] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, at least one of the electronic devices 120 does not include a display component (e.g., communicatively coupled to an external display device and / or includes a media storage device). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0036] The source device and / or multiple electronic devices 120 are sometimes referred to as “terminal devices” or “user devices”. In some embodiments, at least one of the source device 102 and / or electronic devices 120 is an example of a server system, a personal computer, a portable device (e.g., a smartphone, tablet, or laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0037] In an example operation of communication system 100, source device 102 transmits an encoded video stream 108 to server system 112. For example, source device 102 may encode a stream of images captured by the source device. Server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using encoder component 114. For example, server system 112 may apply encoding to the video data that is optimized for network transmission and / or storage. Server system 112 may transmit encoded video data 116 (e.g., at least one encoded video stream) to at least one of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display video images.

[0038] Figure 2A This is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence as a digital video sample stream, which may have any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 YCrCb or RGB), and any suitable sampling structure (e.g., YCrCb 4:2:0 or YCrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed sequentially, produce motion. An image itself can be organized as a spatial array of pixels, where each pixel includes at least one sample, depending on the sampling structure, color space, etc. Those skilled in the art can readily understand the relationship between pixels and samples.

[0039] Encoder component 106 is configured to encode and / or compress images of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. Implementing an appropriate encoding rate is a function of controller 204. In some embodiments, encoder component 106 is configured to perform a conversion between the source video sequence and the bitstream of visual media data (e.g., a video bitstream). In some embodiments, controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by controller 204 may include parameters related to rate control (e.g., image skipping, quantizer and / or rate distortion optimization techniques with λ values), image size, group of images (GOP) layout, maximum motion vector search range, etc. Other functions of controller 204 can be readily identified by those skilled in the art, as they may be related to encoder component 106 optimized for a particular system design.

[0040] In some embodiments, encoder component 106 is configured to operate within an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder 210. Decoder 210 reconstructs the symbols in a manner similar to that of the (remote) decoder to create sampled data (when compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to reference image memory 208. Since decoding of the symbol stream results in bit-accurate results independent of the decoder location (local or remote), the contents of reference image memory 208 are also bit-accurate between the local encoder and the remote encoder. Thus, the prediction portion of the encoder interprets the same sample values ​​as those interpreted by the decoder during prediction as reference image samples. Those skilled in the art will understand that this reference image synchronization principle (which leads to drift if synchronization cannot be maintained due to channel errors, etc.) is common knowledge.

[0041] The operation of decoder 210 can be the same as that of a remote decoder, such as decoder component 122, which will be discussed below. Figure 2B Detailed description. However, a brief reference is provided. Figure 2B Since symbols are available, and the encoding / decoding of symbols for the encoded video sequence by the entropy encoder 214 and the parser 254 can be lossless, the entropy decoding portion of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0042] Aside from parsing / entropy decoding, the decoder techniques described in this paper can exist in the corresponding encoders with essentially the same functional form. Therefore, the subject matter focuses on decoder operations. Descriptions of encoder techniques may be abbreviated, as they can be the inverses of decoder techniques.

[0043] As part of its operation, the source encoder 202 can perform motion-compensated predictive coding, which uses at least one previously encoded frame in the video sequence designated as a reference frame to predictively encode the input frame. In this way, the encoding engine 212 encodes the differences between pixel blocks of the input frame and pixel modules of the reference frame, which can be selected as a predictive reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0044] Decoder 210 decodes encoded video data based on symbol pairs created by source encoder 202, which can be designated as reference frames. The operation of encoding engine 212 can advantageously be a lossy process. When in the video decoder ( Figure 2A When decoding the encoded video data at a location (not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process that can be performed on the reference frame by a remote video decoder, and can store the reconstructed reference frame in reference picture memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame with common content as a reconstructed reference frame (without transmission errors) to be obtained by the remote video decoder.

[0045] Predictor 206 can perform a prediction search on encoding engine 212. That is, for a new frame to be encoded, predictor 206 can search the reference image memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., which can be used as appropriate prediction references for the new image. Predictor 206 can operate on sample blocks pixel by pixel to find appropriate prediction references. As determined by the search results obtained by predictor 206, the input image can have prediction references drawn from multiple reference images stored in reference image memory 208.

[0046] The outputs of all the above-described functional units can be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by lossless compression of the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable-length coding, and / or arithmetic coding).

[0047] In some embodiments, the output of entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequences created by entropy encoder 214 in preparation for transmission via communication channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The transmitter can be configured to combine encoded video data from source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter can send additional data along with the encoded video. Source encoder 202 can include such data as part of the encoded video sequence. Additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, fragments of visual usability information (VUI) parameter sets, etc.

[0048] Controller 204 can manage the operation of encoder component 106. During encoding, controller 204 can assign a specific encoding picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be designated as an intra-frame picture (I-picture), a prediction picture (P-picture), or a bidirectional prediction picture (B-picture). Intra-frame pictures can be encoded and decoded without using any other frames in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. These variations of I-pictures and their respective applications and characteristics are known to those skilled in the art, and therefore will not be repeated here. Predictive pictures can be encoded and decoded using intra-frame prediction or inter-frame prediction, using at most one motion vector and reference index to predict sample values ​​for each block. Bidirectional prediction pictures can be encoded and decoded using intra-frame prediction or inter-frame prediction, using at most two motion vectors and reference indices to predict sample values ​​for each block. Similarly, multiple predictive pictures can be used to reconstruct a single block using two or more reference pictures and associated metadata.

[0049] The source image can typically be spatially subdivided into multiple sample blocks (e.g., each sample block is 4×4, 8×8, 4×8, or 16×16) and encoded block by block. Blocks can be predicted by referencing other (already encoded) blocks determined by the encoding assignment applied to the corresponding image. For example, blocks of image I can be unpredictably encoded or predictively encoded by referencing already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of images P can be unpredictably encoded by referencing a previously encoded reference image through spatial or temporal prediction. Blocks of image B can be unpredictably encoded by referencing one or two previously encoded reference images through spatial or temporal prediction.

[0050] Video can be captured as multiple source images (video frames) in a time series. Intra-frame prediction (often abbreviated as intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame prediction utilizes (temporal or other) correlations between images. In one example, a specific image in the encoding / decoding process (called the current image) is divided into blocks. When a block in the current image is similar to a reference block in a previously encoded and still buffered reference image in the video, that block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image and, in the case of using multiple reference images, can have a third dimension identifying the reference images.

[0051] Figure 2B This is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 is coupled to channel 218 and display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to loop filter 256 and configured to transmit data to display 124 (e.g., via a wired or wireless connection).

[0052] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive at least one encoded video sequence to be decoded by decoder component 122. In some embodiments, decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive encoded video data with other data (e.g., encoded audio data and / or auxiliary data streams), which may be forwarded to their respective user entities (not depicted). The receiver may separate the encoded video sequences from other data. In some embodiments, the receiver receives additional (redundant) data accompanying the encoded video. The additional data may be included as part of at least one encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0053] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also called an entropy decoder), a scaler / inverse transform unit 258, an intra-frame prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference image memory 266, and a current image memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The decoder component 122 may be implemented at least partially in software.

[0054] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, a separate buffer memory is provided outside decoder component 122 (e.g., to combat network jitter) in addition to buffer memory 252 inside decoder component 122 (e.g., configured to handle playback timing). Buffer memory 252 may not be necessary when receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, or buffer memory 252 may be very small. Buffer memory 252 may be necessary for use on best-effort packet networks such as the Internet; buffer memory 252 may be relatively large and / or have an adaptive size, and may be implemented at least partially outside decoder component 122 in an operating system or similar component.

[0055] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. Symbols may include, for example, information for managing the operation of decoder component 122, and / or information for controlling rendering devices such as display 124. Control information for at least one rendering device may be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Availability Information (VUI) parameter set fragments (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. Parser 254 may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to that group. Subgroups may include picture groups (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. Parser 254 may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc., from the encoded video sequence.

[0056] The reconstruction of symbol 270 can involve multiple different units, depending on the type of encoded video picture or its portions (such as inter- and intra-frame pictures, inter- and intra-frame blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed from the encoded video sequence by parser 254. For clarity, the flow of this subgroup control information between parser 254 and the multiple units is not depicted below.

[0057] The decoder component 122 can be conceptually subdivided into a plurality of functional units as described below. In some embodiments, many of these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, this document maintains a conceptual subdivision of the functional units.

[0058] The scaler / inverse transform unit 258 receives quantized transform coefficients and control information (such as the transform to be used, block size, quantization factor, and / or quantization scaling matrix) as at least one symbol 270 from the parser 254. The scaler / inverse transform unit 258 can output a block comprising sample values ​​that can be input into the aggregator 268.

[0059] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use prediction information from previously reconstructed images, but can use prediction information from previously reconstructed portions of the current image. This prediction information can be provided by the intra-prediction unit 262. The intra-image prediction unit 262 can use surrounding reconstructed information obtained from the current (partially reconstructed) image from the current image memory 264 to generate blocks with the same size and shape as the reconstructed blocks. The aggregator 268 can add the prediction information already generated by the intra-prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0060] In other cases, the output samples of the scaler / inverse transform unit 258 belong to inter-frame encoded blocks that may have undergone motion compensation. In this case, the motion compensation prediction unit 260 can access the reference image memory 266 to obtain samples for prediction. After motion compensation is performed on the obtained samples according to the symbols 270 belonging to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address at which the motion compensation prediction unit 260 obtains the predicted samples from the reference image memory 266 can be controlled by motion vectors. Motion vectors can be used by the motion compensation prediction unit 260 in the form of symbols 270, which can have, for example, X, Y, and reference image components. Motion compensation can also include interpolation of sample values ​​obtained from the reference image memory 266 when using subsampled precise motion vectors, motion vector prediction mechanisms, etc.

[0061] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbol 270 from parser 254, but may also be in response to metadata obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to sample values ​​obtained from previous reconstruction and loop filtering. The output of loop filter unit 256 can be a sample stream, which can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-frame prediction.

[0062] Once certain encoded images are reconstructed, they can be used as reference images for future predictions. Once an encoded image has been reconstructed and has been identified as a reference image (e.g., by parser 254), the current reference image can become part of the reference image memory 266, and a new current image memory can be reallocated before reconstructing subsequent encoded images begins.

[0063] Figure 3This is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, at least one network interface 304, memory 314, a user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes at least one processor (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes at least one field-programmable gate array (FPGA), a hardware accelerator, and / or at least one integrated circuit (e.g., an application-specific integrated circuit). The at least one network interface 304 may be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical network).

[0064] User interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 may include at least one of the following: keyboard, mouse, touchpad, touchscreen, microphone, scanner, camera, etc. The at least one output device 308 may include at least one of the following: audio output device (e.g., speaker), visual output device (e.g., display or viewer), etc.

[0065] Memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash memory device and / or other non-volatile solid-state storage devices). Memory 314 may optionally include at least one storage device remote from control circuitry 302. Alternatively, memory 314 or at least one non-volatile solid-state storage device within memory 314 may include a non-volatile computer-readable storage medium. In some embodiments, memory 314 or the non-volatile computer-readable storage medium of memory 314 stores programs, modules, instructions, and data structures, or subsets or supersets thereof:

[0066] ● Operating system 316, which includes programs for handling various basic system services and for performing hardware-related tasks;

[0067] ● Network communication module 318, which is used to connect server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connection);

[0068] ● Encoding module 320, which performs various functions related to encoding and / or decoding data (such as video data). In some embodiments, encoding module 320 is an example of encoder component 114. Encoding module 320 includes, but is not limited to, at least one of the following:

[0069] Decoding module 322, which performs various functions related to decoding encoded data, such as those previously described with respect to decoder component 122; and

[0070] The encoding module 340 performs various functions related to encoding data, such as those previously described with respect to encoder component 106; and

[0071] ● Image memory 352 is used to store images and image data, for example, for use by encoding module 320. In some embodiments, image memory 352 includes at least one of the following: reference image memory 208, buffer memory 252, current image memory 264, and reference image memory 266.

[0072] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transformation unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra-frame prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter unit 256).

[0073] In some embodiments, the encoding module 340 includes an encoding module 342 (e.g., configured to perform various functions previously described with respect to source encoder 202 and / or encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes... Figure 3 The modules shown are subsets of those modules. For example, both decoding module 322 and encoding module 340 use a shared prediction module. Each of the modules identified above, stored in memory 314, corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., a set of instructions) do not need to be implemented as separate software programs, procedures, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments.

[0074] although Figure 3 The illustration depicts a server system 112 according to some embodiments, but... Figure 3 This is intended more as a functional description of various features that can exist in at least one server system, rather than a structural schematic diagram of the embodiments described herein. In practice, as those skilled in the art will know, items shown individually can be combined and some items can be separated. For example, Figure 3Some items shown individually can be implemented on a single server, and a single item can be implemented by at least one server. The actual number of servers used to implement server system 112, and how features are allocated among them, will vary depending on the implementation method and, optionally, in part, on the amount of data traffic processed by the server system during peak usage periods and during average usage periods.

[0075] The encoding / decoding processes and techniques described below can be performed at the devices and systems mentioned above (e.g., source device 102, server system 112, and / or electronic device 120). According to some embodiments, methods for signaling, parsing, and using cross-component prediction modes are described below.

[0076] The method described in this paper can be applied to intra- or inter-frame prediction modes using a least-mean-square optimization-based model, where model parameters are derived from neighboring reconstructed samples of the current block and a reference block. For a first example, the intra-frame prediction mode can be a cross-component prediction mode that uses reconstructed samples of the second color component to derive predicted samples of the first color component, where the current block can be a chroma block and the reference block can be a co-positional luma block. For a second example, the intra-frame prediction mode can be an intra-block duplication or intra-template matching mode that uses a block vector that can be signaled (e.g., intra-block duplication) or implicitly derived (e.g., using template matching) to derive predicted samples of the current block, where the reference block is the block identified by that block vector. For a third example, the inter-frame prediction mode can be an illumination compensation mode that uses neighboring reconstructed samples of the current block and neighboring reconstructed samples of a reference block in a reference frame, based on least-mean-square optimization, to derive predicted samples.

[0077] Figure 4 An example scheme 400 is illustrated according to some embodiments for generating a first chroma sample 402A from at least one luminance sample 404 (e.g., 404A and 404X) in CCP mode (e.g., multiple hypothetical CCP (MHCCP) mode). In some embodiments, the video stream 116 includes a current coded block 406C of the current image frame 408 and a syntax element 420 for CCP mode. The syntax element 420 indicates whether the first chroma sample 402A of the current coded block 406C is reconstructed based on a set of at least one luminance sample 404 of a reference coded block, according to a plurality of model parameters 410. Reference Figure 4 In the example, the reference coding block is the current coding block 406C itself. In some embodiments, for the current coding block 406C, the syntax element 420 is signaled in the video stream 116 at one of the block level, superblock level, image frame level, stripe level, tile level, and image sequence level.

[0078] In some embodiments ( Figure 4 In the CCP mode, there is a Cross-Component Intra-Prediction (CCIP) mode, and the current coding block 406C of the current image frame 408 is encoded in this CCIP mode. In CCIP mode, the current coding block 406C includes a chroma block and corresponds to a reference coding block that includes a co-occurrence luma block. Decoder 122 ( Figure 2B The first chroma sample 402 of the current coding block 406C is determined based on at least one luminance sample 404 of a reconstructed reference coding block. In some cases, the CCIP mode includes a cross-component linear model (CCLM) mode, in which the first chroma sample 402A is derived from a reconstructed luminance sample 404A that is co-located with the chroma sample 402A based on a linear model. Alternatively, in some cases, the CCIP mode includes a convolutional cross-component mode (CCCM), in which the first chroma sample 402A is directly predicted based on a filter shape based on a plurality of reconstructed luminance samples 404X located adjacent to the first luminance sample 404A. Alternatively and additionally, in some cases, the CCIP mode includes an MHCCP mode, in which the first chroma sample 402A is generated by combining at least one first luminance sample 404A that is co-located with the first chroma sample 402A and a plurality of hypotheses using a plurality of weighting factors. A plurality of coefficients are used to combine a plurality of adjacent luminance samples 404X of the first luminance sample 404A to generate a plurality of hypotheses. In other words, in MHCCP mode, the first luminance sample 404A is combined with multiple adjacent luminance samples 404X using multiple model parameters 410 (which are associated with weighting factors and coefficients) to generate the first chromaticity sample 402A. The first chromaticity sample 402A is either a blue difference chromaticity (Cb) sample or a red difference chromaticity (Cr) component.

[0079] In some embodiments, the video stream 116 includes syntax elements 420 for the MHCCP mode. A first chroma sample 402A of the current coded block 406C is configured to be generated by combining at least a first luminance sample 404A co-located with the first chroma sample 402A and at least one adjacent luminance sample 404X of the first luminance sample 404A using multiple model parameters (e.g., ci, cP, cB). Based on the determination that the MHCCP mode is applied, the first chroma sample 402A is predicted according to the following model:

[0080]

[0081] Example chromaticity prediction of Equation 1-MHCCP

[0082] Where predChromaVal is the predicted chromaticity value of the first chromaticity sample 402A; Num is the total number of adjacent luminance samples 404X; S i It is the luminance value of the first luminance sample 404A (where i equals 0) or the luminance value of the adjacent luminance sample 404X (where i is greater than 0), with i as the index; P is the non-linear term; B is the offset term; and c i c P c B These are model parameters. In the example, the nonlinear term P equals (C × C + B) >> bit. `depth`, where `C` is the sample value of the first luminance sample 404A, and `bit_depth` is the number of bits required to represent the luminance sample of the current image frame 408 during encoding and decoding. In some embodiments, `B` is the median luminance value, intermediate luminance value, or average luminance value of the luminance sample 404 of the current coding block 406C. In another example, `B` equals 1 << (bit_depth - 1). In MHCCP mode, it is not necessary to transmit the chroma sample 402 of the current coding block 406C in the video bitstream 116, thereby saving communication bandwidth for the video codec.

[0083] In some embodiments, each of at least one neighboring luminance sample 404X of the first luminance sample 404A is adjacent to the first luminance sample 404A and shares at least one corresponding edge or vertex with the first luminance sample 404A. In some embodiments, the at least one neighboring luminance sample 404X includes a subset or all of the following: the north neighboring luminance sample (also known as the top luminance sample) 404N, the south neighboring luminance sample (also known as the bottom luminance sample) 404S, the west neighboring luminance sample (also known as the left luminance sample) 404W, the east neighboring luminance sample (also known as the right luminance sample) 404E, the northwest neighboring luminance sample (also known as the top left luminance sample) 404NW, the southeast neighboring luminance sample (also known as the bottom right luminance sample) 404SE, the southwest neighboring luminance sample (also known as the bottom left luminance sample) 404SW, and the northeast neighboring luminance sample (also known as the top right luminance sample) 404NE.

[0084] In some embodiments, Equation (1) comprises five terms and represents a five-tap model for determining the first chroma sample 402A of the current coding block 406C based on three linear terms (e.g., associated with the first luminance sample 404A and adjacent luminance samples 404W and 404E), a nonlinear term P, and an offset term B in MHCCP mode. Alternatively, in some embodiments, Equation (1) comprises seven terms and represents a seven-tap model for determining the first chroma sample 402A of the current coding block 406C based on three linear terms (e.g., associated with luminance samples 404A, 404W, 404E, 404N, and 404S), a nonlinear term P, and an offset term B in MHCCP mode.

[0085] In some embodiments, the luminance sample 404 and chrominance sample 402 of the current coding block have different resolutions corresponding to the chrominance subsampling scheme (e.g., 4:2:2 or 4:2:0).

[0086] In some embodiments, multiple model parameters ci, cP, and cB are determined based on a set of at least one set of reference luminance samples 404R and a set of at least one set of co-located reference chrominance samples 402R within a reference region 412 of the current coding block 406C. The reference region 412 is located in the current image frame 408. Furthermore, in some embodiments, the reference luminance samples 404R of the reference region 412 are combined based on equation (1) to regenerate at least one chrominance sample 402A. In some embodiments, the set of at least one co-located reference chrominance sample 402R is compared with at least one regenerated chrominance sample to generate a minimum mean square (LMS) value. The multiple model parameters ci, cP, and cB are iteratively adjusted to reduce the LMS value until the LMS value meets a predefined criterion (e.g., where the LMS value is below a threshold LMS value or is minimized).

[0087] In some embodiments, the plurality of model parameters ci, cP, or cB are derived at least in part based on chroma and luminance samples within a reference region 412 of the current coding block 406C, and the reference region 412 includes at least one coding block that was decoded prior to the current coding block 406C (e.g., Figure 4 (4 coded blocks in the current coding block 406C). In some embodiments, a subset of the at least one coded block is adjacent to the current coding block 406C. In some embodiments, the subset of the at least one coded block is separated from the current coding block 406C by at least one coded block. In some embodiments, the reference region 412 includes at least a portion of at least one row above the current coding block 406C and / or a portion of at least one column to the left of the current coding block 406C. For example, the reference region 412 includes at least a portion of at least one row above the current coding block 406C and / or a portion of at least one column to the left of the current coding block 406C. Figure 4Reference region 412 includes seven rows of luminance samples 404R above the current coding block 406C and nine columns of luminance reference samples 404R to the left of the current coding block 406C. Reference region 412 may include padding rows and padding columns (e.g., in...). Figure 4 (The middle is shadowed).

[0088] Additionally, in some embodiments, the reference region 412 of the current coding block 406C includes at least one of the following: an upper-left reference region 412TL, a top reference region 412T, an upper-right reference region 412TR, a lower-left reference region 412BL, and a left-side reference region 412L. In an example, the reference region 412 includes the top reference region 412T and the left-side reference region 412L. Each of these reference regions includes at least one coding block. In other words, in some embodiments, the reference region 412 includes at least a portion of a plurality of rows above the current coding block 406 and / or a portion of a plurality of columns to the left of the current coding block 406. For example, the reference... Figure 4 Reference region 412 includes a first portion of 6 rows of chroma samples above the current coding block 406C and a second portion of 8 columns of chroma samples to the left of the current coding block 406C. The number of columns in the first portion is determined by the number of columns in the current coding block 406C, and the number of rows in the second portion is determined by the number of rows in the current coding block 406C. In some embodiments, reference region 412 extends to the right of the right boundary of the current coding block 406 by one coding block width and below the bottom boundary of the current coding block 406 by one coding block height. In some embodiments, reference region 412 is adjusted to include only available samples. The extended portion 412E of reference region 412 is padded in unavailable areas to provide side samples for the filter.

[0089] In some embodiments, the reconstructed luminance sample 404R and reconstructed chrominance sample 402R of the reference region 412 are used to generate model parameters in CCP mode. The reference region 412 may be L-shaped, including a lower left reference region, a left side reference region, an upper left reference region, an upper top reference region, and an upper right reference region. For example, the reference region 412 has a first integer K (e.g., 6) reference lines above the current coding block 406C and a second integer L (e.g., 8) columns to the left of the current coding block 406C. The extended portion 412E of the reference region 412 includes padding pixels for the reference samples.

[0090] Figure 5A This is a diagram of an example image frame 500, including the current coded block 406C located at the top boundary 504 of superblock 502, according to some embodiments. Figure 5BThis is a diagram of an example image frame 520, including the current coding block 406C located at the left boundary 510 of superblock 502, according to some embodiments. In MHCCP mode, the current coding block 406 is reconstructed based on model parameters 410 determined according to reference samples 402R and 404R of reference region 412. Reference region 412 includes a first number (N1) rows of chroma reference samples 402R and a second number (N2) rows of luma reference samples 404R above the current coding block 406C. In this example, reference region 412 includes three rows of chroma reference samples 402R and eight rows of luma reference samples 404R. Buffers (e.g., Figure 2B The buffer memory 252 in the buffer stores the first set of reference samples 506. For example, a second set of reference samples 508 is generated from the first set of reference samples 506 by a padding scheme. Based on the first set of reference samples 506 and the second set of reference samples 508, multiple model parameters 410 for use in MHCCP mode are determined for the first chroma sample 402A of the current coding block 406C. Using the multiple model parameters 410, at least one set of luminance samples 404 of the current coding block 406C (e.g., ...) are processed. Figure 4 Samples 404A and 404X in the set are combined to generate the first chroma sample 402A of the current coded block 406C. Image frame 500 or 520 is reconstructed based on the first chroma sample 402A generated from at least one luminance sample 404 in this set.

[0091] refer to Figure 5A In some embodiments, the current coding block 406C is located at the top superblock boundary 504. The topmost row of luma or chroma samples of the current coding block 406C is defined by and located immediately adjacent to the top superblock boundary 504. A first set of reference samples 506 stored in the buffer is located at the top superblock boundary 504 and includes a row of luma reference samples 404R, a row of chroma reference samples 402R, or both. The row of luma samples 404R immediately adjacent to the top superblock boundary 504 can be applied (e.g., copied) to generate each of the remaining seven rows of luma reference samples 404R in the reference region 412. The row of chroma samples 402R immediately adjacent to the top superblock boundary 504 can be applied (e.g., copied) to generate each of the remaining two rows of chroma reference samples 402R in the reference region 412, thereby constructing the reference region 412 of the current coding block 406C located at the top superblock boundary 504.

[0092] In some embodiments, the first set of reference samples 506 may include an entire row of chroma samples 402R or luma samples 404R. Alternatively, in some embodiments, the first set of reference samples 506 is a portion of that row of chroma samples 402R or luma samples 404R (e.g., in reference areas 412TL, 412T, and 412TR). In some embodiments, the left reference area 412L and the lower left reference area 412BL are located within the current coding block 406C. In some embodiments, the left reference area 412L has a first number of lines (e.g., columns), and the top reference area 412T is located outside the current coding block 406C and has a second number of lines (e.g., rows). The first number is equal to or less than the second number.

[0093] refer to Figure 5B In some embodiments, the current coding block 406C is located at the left superblock boundary 510. The leftmost column of luma or chroma samples of the current coding block 406C is defined by and located immediately adjacent to the left superblock boundary 510. A first set of reference samples 506 stored in the buffer is located at the left superblock boundary 510 and includes a column of luma reference samples 404R, a column of chroma reference samples 402R, or both. The column of luma samples 404R immediately adjacent to the left superblock boundary 510 can be applied (e.g., copied) to generate each of the remaining seven columns of luma reference samples 404R of the reference region 412. The column of chroma samples 402R immediately adjacent to the left superblock boundary 510 can be applied (e.g., copied) to generate each of the remaining two columns of chroma reference samples 402R of the reference region 412, thereby constructing the reference region 412 of the current coding block 406C located at the left superblock boundary 510.

[0094] In some embodiments, the first set of reference samples 506 may include an entire column of chroma samples 402R or luma samples 404R. Alternatively, in some embodiments, the first set of reference samples 506 is a portion of that column of chroma samples 402R or luma samples 404R (e.g., within reference areas 412TL, 412L, and 412BL). In some embodiments, the top reference area 412T and the upper right reference area 412TR are located within the current coding block 406C.

[0095] In some embodiments not shown, the current coding block 406C is located at the top left corner of superblock 502. The top sample of the leftmost column of luminance or chrominance samples of the current coding block 406C is defined by both the top superblock boundary 504 and the left superblock boundary 510, and is positioned immediately adjacent to both the top superblock boundary 504 and the left superblock boundary 510. The first set of reference samples 506 includes multiple rows of reference samples 402R and 404R (… Figure 5A ), multi-line reference samples 402R and 404R ( Figure 5B(or both), as is the case with the second set of reference samples 508 generated from reference sample 506.

[0096] Figure 6 is a diagram of an example current coding block 404C applying multiple filter shapes 800 according to some embodiments. In some embodiments, the MHCCP mode has two different filter shapes 800, including a vertical filter shape 800V and a horizontal filter shape 800H. In some embodiments, the second syntax element 440 ( Figure 4 The signal is sent to the video stream 116 to indicate the selected filter shape between the vertical filter shape 800V and the horizontal filter shape 800H. In the example, the second syntax element 440 includes a flag. The selected filter shape has a number of model parameters (e.g., 5 parameters) for prediction, and these model parameters are derived for both of the two different filter shapes 800. For the horizontal filter shape 800H, the predicted chromaticity value (predChromaVal) of the first chromaticity sample 402A is determined by applying the model parameters c0, c1, c2, c3, and c4 to the first luminance sample 404A (C), the left luminance sample 404W (L), the right luminance sample 404E (R), the nonlinear term (E), and the offset term (F) in the following manner:

[0097]

[0098] Equation 2

[0099] For a vertical filter shape of 800V, the predicted chromaticity value (predChromaVal) of the first chromaticity sample 402A is determined by applying the model parameters c0, c1, c2, c3, and c4 to the first luminance sample 404A (C), the top luminance sample 404N (T), the bottom luminance sample 404S (B), the nonlinear term (E), and the offset term (F) as follows:

[0100]

[0101] Equation 3

[0102] Furthermore, in some embodiments, for each of the horizontal filter shape 800H and the vertical filter shape 800V, the corresponding five model parameters are derived by Gaussian elimination based on LMS optimization.

[0103] In some embodiments, in MHCCP mode, the right luminance sample 404E(R) and the bottom luminance sample 404S(B) are not applied in the prediction of the first chromaticity sample 402A. For the horizontal filter shape 800H and the vertical filter shape 800V, Equations 2 and 3 used for predicting the first chromaticity sample 402A are updated as follows:

[0104]

[0105] Equation 4

[0106]

[0107] Equation 5

[0108] In some embodiments, the number of upper and left reference rows is 3 each, and there is 1 fill row in the chroma channel. For a 4:2:0 format sequence, the luma channel may require 6 rows and 2 fill rows as the corresponding reference area. In MHCCP mode, vertical prediction mode and horizontal prediction mode are implemented to apply vertical filter shape 800V and horizontal filter shape 800H respectively. For vertical prediction, the first luma sample 404A(C), the top luma sample 404N(T), and the bottom luma sample 404S(B) are combined to generate the first chroma sample 402A (predChromaVal) as follows:

[0109]

[0110] Equation 6

[0111] in is the first luminance sample; L and R are the left luminance sample 404W and the right luminance sample 404E, respectively; and m is the offset value, which represents the median value of the pixel intensity (e.g., the median of the range of luminance sample values). For horizontal prediction, the first luminance sample 404A (C), the left luminance sample 404W (L), and the right luminance sample 404E (R) are combined to generate the first chroma sample 402A (predChromaVal) as follows:

[0112]

[0113] Equation 7

[0114] Where T and B are the top luminance sample 404N and the bottom luminance sample 404S, respectively. When the current coding block 406C is located at the boundary of a superblock or coding tree unit (e.g., top boundary 504, left boundary 510), a row of reference samples 506 located above the boundary of that superblock or coding tree unit ( Figure 5A (This can be used in hardware implementation within a buffer.)

[0115] In some embodiments, the current coding block 406C includes a first prediction block 802 located at the top superblock boundary 504 and a second prediction block 804 separated from the top superblock boundary 504 by the first prediction block 802. A horizontal filter shape 800H is applied to combine a set of luminance samples 404 to generate a chrominance sample 402 of the first prediction block 802. A vertical filter shape 800V is applied to combine a set of luminance samples 404 to generate a chrominance sample 402 of the second prediction block 804. In other words, in some embodiments, the current coding block 406C has a block-level filter shape (e.g., a vertical filter shape 800V) and includes a first prediction block 802 positioned immediately adjacent to the boundary 504 and a second prediction block 804 separated from the boundary 504 by at least one sample of the first prediction block 802. In MHCCP mode, each chroma sample 402 of the first prediction block 802 is reconstructed using a first filter shape (e.g., horizontal filter shape 800H) based on a set of corresponding first luminance samples 404, and each chroma sample of the second prediction block 804 is reconstructed using a block-level filter shape (e.g., vertical filter shape 800V) different from the first filter shape (e.g., horizontal filter shape 800H) based on a set of corresponding second luminance samples.

[0116] In some embodiments, the current coding block 406C has a predefined block-level filter shape corresponding to a plurality of model parameters 410. The current coding block 406C is located in superblock 502 ( Figure 5A At the top boundary 504 of the predefined block-level filter shape, the predefined block-level filter shape is a horizontal filter shape 800H, and at least one luminance sample of the group applied to reconstruct the first chroma sample 402A of the current coding block 406C is located on the same row as the first luminance sample 404A, which is in the same position as the first chroma sample 402A.

[0117] In some embodiments not shown, the current coding block has a block-level filter shape and includes a superblock boundary 510 immediately adjacent to the left. Figure 5B The first prediction block 802 is located by a first filter shape (e.g., a horizontal filter shape 800H) and the second prediction block is separated from the boundary by at least one sample. In MHCCP mode, each chroma sample (e.g., a column of chroma samples) of the first prediction block 802 is reconstructed based on a set of corresponding first luminance samples using a first filter shape (e.g., a horizontal filter shape 800H), and each chroma sample of the second prediction block is reconstructed based on a set of corresponding second luminance samples using a block-level filter shape different from the first filter shape (e.g., a vertical filter shape 800V).

[0118] In some instances, intra-frame prediction mode is a cross-component prediction mode using a linear model (e.g., CfL mode, Chrominance from Luminance, chroma prediction based on luminance) that uses reconstructed samples of the second color component to derive predicted samples of the first color component, where the current block can be a chroma block and the reference block can be a co-occurring luminance block. In some systems, CfL mode has two modes: explicit CfL and implicit CfL. Explicit CfL means explicitly signaling the weight factor or the index of the set of weight factors, while implicit CfL means implicitly deriving the weight factor at the encoder and decoder sides. For a second example, intra-frame prediction mode could be a cross-component prediction mode using a non-linear model (e.g., MHCCP) that uses reconstructed samples of the second color component to derive predicted samples of the first color component, where the current block can be a chroma block and the reference block can be a co-occurring luminance block. In some systems, there are three MHCCP prediction modes, which use in-situ downsampled luminance samples, the downsampled left adjacent sample of the in-situ luminance sample, and the downsampled upper adjacent sample of the in-situ luminance sample to perform prediction.

[0119] In some embodiments, MHCCP is used to generate a chromaticity prediction block by combining several linearly or non-linearly weighted luminance samples. Weighting factors are derived based on adjacent luminance reconstruction samples and adjacent chromaticity reconstruction samples. The reference region is... Figure 5A and Figure 5B The reference sample is illustrated in the figure. It can be L-shaped, including a lower left reference sample, a left reference sample, a top left reference sample, a top reference sample, and a top right reference sample. There can be three reference rows and reference columns. In some embodiments, the third row (e.g., the one furthest from the block) is filled.

[0120] In some embodiments, the MHCCP pattern has three different shapes: vertical, horizontal, and center-shaped. Each pattern requires only three parameters for derivation. For the center-prediction pattern, the prediction pattern is shown in Equation 8.

[0121]

[0122] Equation 8

[0123] For the left-hand (horizontal) prediction pattern, the prediction pattern is shown in Equation 9.

[0124]

[0125] Equation 9

[0126] For the top (vertical) prediction pattern, the prediction pattern is shown in Equation 10.

[0127]

[0128] Equation 10

[0129] Where E and F represent the nonlinear components (F) of the center pixel, respectively. 2 + median) and bit-depth offset (median).

[0130] As described above, reference regions in the first color component and / or the second / third color component are used to generate a model for cross-component prediction (such as MHCCP). The reference sample is L-shaped (e.g., as shown in the image). Figure 5C (As indicated), including the lower left reference sample, left reference sample, upper left reference sample, upper top reference sample, and upper right reference sample of the coded block. There can be K reference rows above the coded block 550, and L columns to the left of the coded block 550. Dark pixel 554 is a padding pixel used for the reference pixels. Figure 5C The diagram illustrates an example reference region. In this example, both the upper and left reference rows have three rows. The techniques and methods for reducing the number of samples used in the reference region to derive model parameters are described below.

[0131] In some embodiments, the sequence-level flag "enable-mhccp" is used to control the MHCCP mode. At the code block level, the 3-symbol CfL_mode flag can be replaced with two 2-symbol flags. A new syntax element, cfl_mhccp_switch_flag, is added to indicate whether the chroma prediction mode from luma is CfL or MHCCP mode. Additionally, the cfl_mode syntax can be reduced from 3 symbols to 2 symbols and is now only used to represent CfL prediction mode. When "enable-mhccp" is enabled, the signaling logic can proceed as follows: Figure 6B As shown. In some embodiments, when “enable-mhccp” is off, and the prediction mode is CfL, the CfL mode flag can be signaled to indicate whether to use the CfL explicit mode or the CfL implicit mode.

[0132] In some embodiments, for MHCCP, two reference rows are used in the chroma channel and four rows plus two fill rows are used in the luma channel. In some embodiments, non-adjacent rows in the chroma channel are not used for MHCCP, and only one reference row from the top and left in the chroma channel is used. In some embodiments, two rows plus two fill rows from the top and left in the luma channel are required to derive the model parameters for MHCCP. In some embodiments, three luma rows and two fill rows are used.

[0133] In some embodiments, for luma padding, considering that the Multiple Reference Line (MRL) mode uses four adjacent lines above and to the left of the coded block, the MHCCP mode also utilizes four adjacent lines and one padding line as a reference area. In some embodiments, the luma channel reference samples are downsampled. For example, downsampling is performed using weights {1, 2, 1; 1, 2, 1}.

[0134] Using the example techniques above, coding gain can be improved. Table 1 below illustrates the improvements in signal-to-noise ratio and encoding / decoding time based on simulations performed using the current design (e.g., AVM research-v9) with various video data (e.g., representing AOM test conditions). Results are reported for intra-frame, random access, and low-latency configurations.

[0135]

[0136] Table 1 - Simulation Results

[0137] Figure 7A This is a flowchart illustrating a method 700 for decoding video according to some embodiments. The method 700 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 700 is executed by executing instructions stored in memory (e.g., memory 314) of the computing system.

[0138] The system receives (702) a video stream comprising multiple blocks, including the current block. The system determines (704) that an MHCCP mode is enabled for the current block. The system identifies (706) a reference region for the MHCCP mode. The system subsamples (708) the reference region. The system applies (710) the MHCCP mode to the current block using the subsampled reference region. In this way, the subsampling method can be applied to the reference data collection process to reduce computational complexity.

[0139] In some embodiments, when collecting data in a reference region, the subsampling rate is This means that for each row or each column, or for each of the two dimensions of rows and columns... One sample is selected from each sample. It is a positive integer. In one example, a quincunx downsampling method is used to reduce the number of samples used in the reference region for the model derivation process. In another example, the subsampling rate is 1 / 2 for both the row and column dimensions. Figure 5E An example of such subsampling is illustrated, where diagonally patterned pixels 558 (e.g., 558-1 and 558-2) indicate subsampled pixels.

[0140] In some embodiments, when collecting data in a reference region, the subsampling rate varies for different rows / columns. It is changing, among which It is a row or column index. When When it is 0, the first Zero samples are selected for a row or column. In some embodiments, the subsampling rate is smaller when the reference sample for the row or column is closer to the current coding block.

[0141] For example, the subsampling rate might be 1 for the nearest neighboring reference row / column, 1 / 2 for the second nearest neighboring reference row / column, and 0 for the third reference row / column. Figure 5F An example of such subsampling is illustrated, where the diagonally patterned pixel 558 (e.g., pixel 558-3) indicates the subsampled pixel.

[0142] For example, , and This means that 0 samples are selected in the first row and the first column, and one sample is selected for every two pixels in the second row and the second column, as well as in the third row and the third column. Figure 5G An example of such subsampling is illustrated, where the diagonally patterned pixel 558 (e.g., 558-4) indicates the subsampled pixel.

[0143] In another example, , and This means that 0 samples are selected in the first row and first column, all samples are selected in the second row and second column, and one sample is selected for every two pixels in the third row and third column. Figure 5H An example of such subsampling is illustrated, where the diagonally patterned pixel 558 (e.g., 558-5) indicates the subsampled pixel.

[0144] In another example, , and This means that one sample is selected for every two pixels in the first row and first column, and for the third row and third column, and all samples are selected for the second row and second column. Figure 5I An example of such subsampling is illustrated, where the diagonally patterned pixel 558 (e.g., 558-6) indicates the subsampled pixel.

[0145] In some embodiments, the subsampling ratio varies for different rows and / or columns. In some embodiments, the subsampling ratio depends on the block shape; for example, the larger side may have a more aggressive subsampling ratio. For example, when the ratio of block width to block height is 2:1, the subsampling ratio for the row dimension may be 1 / 4, while the subsampling ratio for the column dimension may be 1 / 2.

[0146] In some embodiments, different regions in the reference region (e.g., such as...) Figure 5D The regions shown (556) have different subsampling ratios. For example, the lower left and upper right regions can be skipped, and the above techniques can be applied only to the upper left, top, and left regions.

[0147] In some embodiments, other techniques such as pooling methods, downsampling, low-pass filtering, compression, or any form of irregular sampling are used for subsampling. In some embodiments, the subsampling rate of the reference samples in the row and column dimensions depends on the block width, block height, and sample availability. In some embodiments, when one side is more than twice as large as the other, reference samples are selected only from the larger side. For example, when the block width to block height ratio is 4:1 (or 8:1), the "left" and "bottom left" regions may not have samples to be selected. In another example, when the block width to block height ratio is 1:4 (or 1:8), the "top" and "top right" regions may not have samples to be selected.

[0148] In some embodiments, when one side is no more than twice the size of the other side, a certain number of reference samples are evenly distributed on both sides. For example, the "top left" area, "top" area, "top right" area, "left side" area, and "bottom left" area can have the same number of reference samples to be selected.

[0149] In some embodiments, the total number of samples is represented as N squared, where N is a positive integer. In some embodiments, the subsampling technique needs to satisfy the requirement that the total number of samples equals the square of N. Some samples may be discarded to meet this condition.

[0150] In some embodiments, the total number of reference samples depends on the block size. For example, larger blocks may have more reference samples. In the example, for an 8×8 block, the total number of reference samples is 32; while for a 16×16 block, the total number of reference samples is 64×64.

[0151] Figure 7BThis is a flowchart illustrating a method 750 for encoding video according to some embodiments. The method 750 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 750 is executed by executing instructions stored in memory (e.g., memory 314) of the computing system. In some embodiments, method 750 is executed by the same system as method 700 described above.

[0152] The system receives (752) video data comprising multiple blocks, including the current block. The system determines (754) that an MHCCP mode is enabled for the current block. The system identifies (756) a reference region for the MHCCP mode. The system subsamples the reference region (758). The system applies the MHCCP mode (760) to the current block using the subsampled reference region. As previously mentioned, the encoding method can be a mirror image of the decoding method described herein (e.g., the application of cross-component prediction modes). For brevity, these details will not be repeated here.

[0153] although Figure 7A and Figure 7B Multiple logical stages are illustrated in a specific order, but stages that are not in order can be reordered, and other stages can be combined or split. Any reordering or other grouping not specifically mentioned will be obvious to those skilled in the art, and therefore the ordering and grouping presented herein are not exhaustive. Furthermore, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0154] Now let’s turn to some example implementations.

[0155] (A1) In one aspect, some embodiments include a video decoding method (e.g., method 700). In some embodiments, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). In some embodiments, the method is performed at a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) comprising a plurality of blocks (e.g., corresponding to a set of pictures), the plurality of blocks including a current block; (ii) determining that an MHCCP mode is enabled for the current block; (iii) identifying a reference region for the MHCCP mode; (iv) subsampling the reference region; and (v) applying the MHCCP mode to the current block using the subsampled reference region. In this way, the subsampling method can be applied to the reference data collection process to reduce computational complexity.

[0156] (A2) In some embodiments of A1, the reference region is the same reference region used for intra-frame prediction multiple reference line selection (MRLS) for the current block.

[0157] (A3) In some embodiments of A1 or A2, subsampling the reference region involves applying a subsampling rate of 1 / N, where N is the number of samples. For example, when collecting data within the reference region, the subsampling rate could be... This means that for each row or each column, or for each of the two dimensions of rows and columns... One sample is selected from each sample. is a positive integer.

[0158] (A4) In some embodiments of A3, N equals 2. For example, the subsampling rate is 1 / 2 for both the row and column dimensions. In some embodiments, the total number of samples is the square of M, where M is a positive integer. In some embodiments, the total number of samples needs to be equal to the square of N, for example, some samples may need to be discarded to satisfy this condition. As an example, the total number of reference samples may depend on the block size; for example, a larger block may have more reference samples. For example, for an 8×8 block, the total number of reference samples is 32; while for a 16×16 block, the total number of reference samples is 64.

[0159] (A5) In some embodiments of A3 or A4, the subsampling rate is different for different portions of the reference region. For example, when collecting data in the reference region, the subsampling rate is different for different rows / columns. It can be variable, among which It is a row or column index. When When it is 0, the first Zero samples are selected for each row or column. As an example, the subsampling ratio can vary per row and / or per column. For instance, different regions can have different subsampling ratios, such as... Figure 5D As illustrated. As an example, sub-sampling can be skipped for the lower left and / or upper right regions.

[0160] (A6) In some embodiments of A5, the subsampling rate is lower for reference samples closer to the current block and higher for reference samples farther from the current block. For example, the subsampling rate is smaller when the row or column reference sample is closer to the current coding block. As an example, the subsampling rate is 1 for the nearest neighboring reference row / column, 1 / 2 for the second neighboring reference row / column, and 0 for the third reference row / column. In another example, , and This means that 0 samples are selected for the first row and first column, and one sample is selected for every two pixels in the second row and second column, and the third row and third column. In another example, , and In another example, , and This means that the first row and first column select 0 samples, the second row and second column select all samples, and the third row and third column select one sample every two pixels. In another example, , and This means that one sample is selected for every two pixels in the first row and first column, and for the third row and third column, and all samples are selected for the second row and second column.

[0161] (A7) In some embodiments of A5 or A6, the subsampling rate is different for different rows of the reference region. For example, the subsampling rate of the reference samples in the row and column dimensions depends on the block width, block height, and sample availability. As an example, when one side is more than twice as large as the other, reference samples are selected only from the larger side. In one example, when the block width to block height ratio is 4:1 (or 8:1), the "left" and "bottom left" regions have no samples to be selected. In another example, when the block width to block height ratio is 1:4 (or 1:8), the "top" and "top right" regions have no samples to be selected. In some embodiments, when one side is not more than twice as large as the other, a certain number of reference samples are evenly distributed on both sides. For example, the "top left," "top," "top right," "left," and "bottom left" regions may have the same number of reference samples to be selected.

[0162] (A8) In some embodiments of any of A1 to A7, subsampling the reference region includes applying quincunx downsampling to the reference region. For example, a quincunx downsampling method can be used to reduce the number of samples in the reference region used for the model derivation process.

[0163] (A9) In some embodiments of any of A1 through A8, subsampling the reference region includes applying a sampling rate based on the block size or block shape of the current block. For example, the subsampling rate may depend on the block shape; for instance, the larger side may have a more aggressive subsampling rate. As an example, when the block width to block height ratio is 2:1, the subsampling rate for the row dimension is 1 / 4, while the subsampling rate for the column dimension is 1 / 2.

[0164] (A10) In some embodiments of any of A1 to A9, subsampling the reference region includes applying irregular sampling to the reference region. Other methods include, for example, pooling methods, downsampling, low-pass filtering, compression, or any form of irregular sampling.

[0165] (B1) In another aspect, some embodiments include a video encoding method (e.g., method 750). In some embodiments, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed at an encoding / decoding module (e.g., encoding / decoding module 320). The method includes: (i) receiving video data (e.g., a source video sequence) comprising a plurality of blocks (e.g., corresponding to a set of pictures), the plurality of blocks including a current block; (ii) determining that an MHCCP mode is enabled for the current block; (iii) identifying a reference region for the MHCCP mode; (iv) subsampling the reference region; and (v) applying the MHCCP mode to the current block using the subsampled reference region. In some embodiments, the method further includes signaling information about the encoding of the current block.

[0166] (B2) In some embodiments of B1, the reference region is the same reference region used for intra-frame prediction multi-reference line selection (MRLS) for the current block.

[0167] (B3) In some embodiments of B1 or B2, subsampling the reference region includes applying a subsampling rate of 1 / N, where N is the number of samples.

[0168] (B4) In some embodiments of B3, the subsampling rate is different for different parts of the reference region.

[0169] (B5) In some embodiments of B3 or B4, the subsampling rate is different for different rows of the reference region.

[0170] (B6) In some embodiments of any of B1 to B5, subsampling the reference region includes applying irregular sampling to the reference region.

[0171] On the other hand, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A10 above and B1 to B6 above).

[0172] In another aspect, some embodiments include a non-volatile computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A10 above and B1 to B6 above). In some embodiments, the memory or non-transitory computer-readable storage medium stores a video stream containing any features disclosed herein (e.g., syntax elements and encoding information).

[0173] Unless otherwise stated, any syntax element described herein can be a High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS can correspond to sequence level, frame level, slice level, or tile level. As another example, HLS elements can be signaled in Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, picture header, tile header, and / or CTU header.

[0174] It will be understood that although the terms “first,” “second,” etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of embodiments and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term “and / or,” as used herein, refers to and covers any and all possible combinations of at least one of the associated listed items. It will be further understood that, when used in this specification, the terms “comprises” and “comprising” indicate the presence of the stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of at least one other feature, integral, step, operation, element, component, and / or group thereof.

[0175] As used herein, depending on the context, the term "if" can be interpreted as meaning "when" or "at," or "in response to determination" or "according to determination" or "in response to detection," the stated prerequisite is true. Similarly, depending on the context, the phrases "if determination [the stated prerequisite is true]" or "if [the stated prerequisite is true]" or "when [the stated prerequisite is true]" can be interpreted as meaning "at determination" or "in response to determination" or "according to determination" or "at detection" or "in response to detection," the stated prerequisite is true.

[0176] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to understand them.

Claims

1. A video decoding method, characterized in that, The method includes: Receive a video stream comprising multiple blocks, including the current block; Determine whether to enable the Multi-Hypothesis Cross-Component Prediction (MHCCP) mode for the current block; Identify the reference region for the MHCCP mode; Subsampling is performed on the reference region; and The MHCCP mode is applied to the current block using a reference region that has been subsampled.

2. The method according to claim 1, characterized in that, The reference region is the same reference region used by Multi-Reference Line Selection (MRLS) for intra-frame prediction of the current block.

3. The method according to claim 1 or 2, characterized in that, Subsampling the reference region involves applying a subsampling rate of 1 / N, where N is the number of samples.

4. The method according to claim 3, characterized in that, N equals 2.

5. The method according to claim 3, characterized in that, The subsampling rate is different for different parts of the reference region.

6. The method according to claim 5, characterized in that, The subsampling rate is lower for reference samples closer to the current block and higher for reference samples farther from the current block.

7. The method according to claim 5, characterized in that, The subsampling rate is different for different rows of the reference region.

8. The method according to any one of claims 1-7, characterized in that, Subsampling the reference region includes applying quincunx downsampling to the reference region.

9. The method according to any one of claims 1-7, characterized in that, Subsampling the reference region includes applying a sampling rate based on the block size or block shape of the current block.

10. The method according to any one of claims 1-7, characterized in that, Subsampling the reference region includes applying irregular sampling to the reference region.

11. A video encoding method, characterized in that, The method includes: Receive video data comprising multiple blocks, including the current block; Determine that MHCCP mode is enabled for the current block; Identify the reference region for the MHCCP mode; Subsampling is performed on the reference region; and The MHCCP mode is applied to the current block using a reference region that has been subsampled.

12. The method according to claim 11, characterized in that, The reference region is the same reference region used by Multi-Reference Line Selection (MRLS) for intra-frame prediction of the current block.

13. The method according to claim 11 or 12, characterized in that, Subsampling the reference region involves applying a subsampling rate of 1 / N, where N is the number of samples.

14. The method according to claim 13, characterized in that, The subsampling rate is different for different parts of the reference region.

15. The method according to claim 13, characterized in that, The subsampling rate is different for different rows of the reference region.

16. The method according to any one of claims 11-15, characterized in that, Subsampling the reference region includes applying irregular sampling to the reference region.

17. A method for storing video streams, characterized in that, Generate a video stream by performing the method according to any one of claims 11-16; and store the video stream.

18. A non-volatile computer-readable storage medium, characterized in that, The video stream generated by the video encoding method is stored, the video stream comprising: Encoded information for multiple blocks of video data, including the current block; and The video encoding method mentioned above includes: Determine that MHCCP mode is enabled for the current block; Identify the reference region for the MHCCP mode; Subsampling is performed on the reference region; and The MHCCP mode is applied to the current block using a reference region that has been subsampled.

19. A computing system, characterized in that, include: At least one memory is configured to store program code; as well as At least one processor is configured to read the program code and operate according to the instructions of the program code to perform the method as described in any one of claims 1-16.