Method and apparatus for inheriting temporal cross-component models with buffer constraints for video coding
Patent Information
- Application Number
- CN202480027913.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-24
- Filing Date
- 2024-04-22
- Publication Date
- 2025-11-21
AI Technical Summary
Existing video encoding and decoding systems suffer from cross-component redundancy when processing color video, resulting in low encoding and decoding efficiency. This is especially true when limited by the size of the buffer, making it difficult to effectively utilize reference data for efficient prediction.
By inheriting the parameters of temporally and spatially adjacent models, the prediction of cross-component models is optimized. Multiple models (such as CCLM, MMLM, CCCM, and GLM) are used for color image encoding and decoding. Combined with buffer constraints, the prediction accuracy and encoding/decoding performance are improved.
It improves the efficiency and quality of color video encoding and decoding, reduces cross-component redundancy, optimizes buffer utilization, and enhances the overall performance of the video encoding and decoding system.
Smart Images

Figure CN121002872A_ABST
Abstract
Description
[0001] Related citations This application is a non-provisional application, claiming priority to U.S. Provisional Patent Application No. 63 / 497,759, filed April 24, 2023. The contents of the aforementioned U.S. Provisional Patent Application are incorporated herein by reference. Technical Field
[0002] This invention relates to video encoding and decoding systems. Specifically, it relates to inheriting temporal adjacency model parameters for cross-component prediction of correlation modes in a video encoding and decoding system with a limited buffer size for storing reference data. Background Technology
[0003] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Video Coding Experts Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Joint Video Experts Team (JVET) of the Moving Picture Experts Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC). This standard was published as an ISO standard in February 2021: ISO / IEC 23090-3:2021, Information technology – Codec representation of immersive media – Part 3: Versatile Video Coding. VVC was developed based on its predecessor, High Efficiency Video Coding (HEVC). It improves encoding and decoding efficiency by adding more encoding and decoding tools, and can handle various types of video sources, including three-dimensional (3D) video signals.
[0004] Figure 1AAn exemplary adaptive inter-frame / intra-frame video coding system incorporating loop processing is illustrated. For intra-frame prediction, prediction data is derived from previously encoded / decoded picture video data in the current frame. For inter-frame prediction 112, motion estimation (ME) is performed at the encoder, and motion compensation (MC) is performed based on the results of ME to provide prediction data derived from other pictures and motion data. Switch 114 selects either intra-frame prediction 110 or inter-frame prediction 112 and provides the selected prediction data to adder 116 to form a prediction error, also known as a residual. The prediction error is then processed by transform (T) 118, followed by quantization (Q) 120. The residuals from the transform and quantization are then encoded by entropy encoder 122 to be included in the video bitstream corresponding to the compressed video data. The bitstream associated with the transform coefficients is then packaged with auxiliary information (such as motion and encoding / decoding modes associated with intra-frame and inter-frame prediction) and other information such as parameters associated with loop filters applied to the underlying image regions. Figure 1A As shown, auxiliary information related to intra-frame prediction 110, inter-frame prediction 112, and loop filter 130 is provided to entropy encoder 122. When inter-frame prediction mode is used, one or more reference images must also be reconstructed at the encoder. Therefore, the residuals from transform and quantization are processed by inverse quantization (IQ) 124 and inverse transformation (IT) 126 to recover the residuals. The residuals are then added back to the prediction data 136, reconstructing the video data at reconstruction (REC) 128. The reconstructed video data can be stored in reference image buffer 134 and used for prediction of other frames.
[0005] like Figure 1A As shown, the incoming video data undergoes a series of processes in the encoding system. The reconstructed video data from REC 128 may be subject to various degradations due to these processes. Therefore, before storing the reconstructed video data in the reference picture buffer 134, a loop filter 130 is typically applied to the reconstructed video data to improve video quality. For example, a deblocking filter (DF), sample adaptive offset (SAO), and adaptive loop filter (ALF) may be used. It may be necessary to incorporate loop filter information into the bitstream so that the decoder can correctly recover the required information. Therefore, loop filter information is also provided to the entropy encoder 122 for incorporation into the bitstream. Figure 1A In the process, the loop filter 130 is applied to the reconstructed video, and then the reconstructed samples are stored in the reference image buffer 134. Figure 1A The system described herein is intended to illustrate an exemplary architecture of a typical video encoder. It may correspond to a High Efficiency Video Coding (HEVC) system, VP8, VP9, H.264, or VVC.
[0006] like Figure 1B As shown, apart from transform 118 and quantization 120, the decoder can use the same or partially the same function blocks as the encoder, since the decoder only needs inverse quantization 124 and inverse transform 126. The decoder uses entropy decoder 140 instead of entropy encoder 122 to decode the video bitstream into quantized transform coefficients and the required encoding / decoding information (e.g., ILPF information, intra-frame prediction information, and inter-frame prediction information). Intra-frame prediction 150 at the decoder end does not require mode search. Instead, the decoder only needs to generate intra-frame predictions based on the intra-frame prediction information received from entropy decoder 140. Furthermore, for inter-frame prediction, the decoder only needs to perform motion compensation (MC 152) based on the inter-frame prediction information received from entropy decoder 140, without motion estimation.
[0007] According to VVC, the input image is divided into non-overlapping square block regions called Coding Tree Units (CTUs), similar to HEVC. Each CTU can be further divided into one or more smaller coding units (CUs). The resulting CU segments can be squares or rectangles. Furthermore, VVC divides the CTUs into prediction units (PUs), which serve as units for applying prediction processing (such as inter-frame prediction, intra-frame prediction, etc.).
[0008] The VVC standard incorporates various new encoding and decoding tools to further improve encoding and decoding efficiency relative to the HEVC standard. Some of the new tools related to this invention are described below.
[0009] Cross-Component Linear Model (CCLM) Prediction To reduce cross-component redundancy, VVC uses a cross-component linear model (CCLM) prediction mode, where a linear model is used to predict chrominance samples based on reconstructed luminance samples from the same CU, as shown below: (1) in This represents the predicted chromaticity samples in the CU. This represents downsampled reconstructed luminance samples from the same CU.
[0010] CCLM parameters ( and This is derived from a maximum of four adjacent chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set as follows: When the LM_LA mode is applied, W' = W, H' = H; When the LM_A mode is applied, W' = W + H; When the LM_L mode is applied, H' = H + W.
[0011] The adjacent positions above are denoted as S[0, -1]…S[W' -1, -1], and the adjacent positions to the left are denoted as S[-1,0]…S[-1, H' -1]. Then, the four samples are selected as follows: When the LM_LA mode is applied and both the upper and left adjacent samples are available, S[W' / 4, -1], S[ 3 W' / 4, -1 ], S[-1, H' / 4 ], S[-1, 3 H' / 4 ]; When LM_A mode is applied or only the upper adjacent sample is available, S[W' / 8, -1], S[3] W' / 8,-1 ], S[ 5 W' / 8, -1 ], S[ 7 W' / 8, -1 ]; When the LM_L mode is applied or only the left adjacent sample is available, S[-1, H' / 8], S[-1, 3] H' / 8 ], S[ -1, 5 H' / 8 ], S[ -1, 7 H' / 8 ].
[0012] Four adjacent luminance samples at a selected location are downsampled and compared four times to find two larger values: x0A and x1A, and two smaller values: x0B and x1B. Their corresponding chromaticity sample values are represented as y0A, y1A, y0B, and y1B. Then xA, xB, yA, and yB are derived as: ; ; ; (2) Finally, the linear model parameters and Obtained according to the following procedure.
[0013] (3) (4) Figure 2 This shows examples of the left and top samples involved in the LM_LA mode, as well as the sample positions of the current block. Figure 2 Display color block 210 The corresponding brightness block 220 And the relative sample positions of their neighboring samples (shown as filled circles).
[0014] In addition to the upper and left templates being used together to calculate linear model coefficients, they can also be used alternately in the other two LM modes, known as LM_A and LM_L modes.
[0015] In LM_A mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H+W) samples.
[0016] In LM_LA mode, the left and top templates are used to calculate the coefficients of the linear model.
[0017] The terms {LM_LA, LM_L, LM_A} and {CCLM_LT, CCLM_L, CCLM_T} can be used interchangeably in this document.
[0018] To match the chroma sample positions of a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve downsampling ratios of 2 to 1 in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the SPS level flag. These two downsampling filters correspond to "Type-0" and "Type-2" content, respectively.
[0019] ; (5) ; (6) Please note that when the upper reference line is at the CTU boundary, only one luminance line (the general line buffer in intra-frame prediction) is used to create the downsampled luminance sample.
[0020] This parameter calculation is performed as part of the decoding process, not just as part of the encoder's search operation. Therefore, no syntax is used to convey the α and β values to the decoder.
[0021] Multi-model CCLM (MMLM) In JEM (J. Chen, E. Alshina, GJ Sullivan, J.-R. Ohm, and J. Boyce, Algorithm Description of Joint Exploration Test Model 7, document JVET-G1001, ITU-T / ISO / IEC Joint Video Exploration Group (JVET), Jul. 2017), a multi-model CCLM mode (MMLM) was proposed to predict chroma samples of the entire CU from luminance samples using two models. In MMLM, neighboring luminance and chroma samples of the current block are classified into two groups, each group serving as a training set to derive a linear model (i.e., deriving specific α and β for a specific group). Furthermore, samples of the current luminance block are also classified according to the classification rules of neighboring luminance samples. Three MMLM model modes (MMLM_LA, MMLM_T, and MMLM_L) are allowed to select neighboring samples from the left and top, top only, and left only, respectively.
[0022] In this article, the terms MMLM_LA and MMLM_LT can be used interchangeably.
[0023] Figure 3 This shows an example of classifying neighboring samples into two groups. The threshold is calculated as the average of the neighboring reconstructed brightness samples. Neighboring samples with Rec'L [x,y] <= the threshold are classified into group 1; while neighboring samples with Rec'L [x,y] > the threshold are classified into group 2.
[0024] (7) Therefore, MMLM uses two models based on the sample level of adjacent samples.
[0025] Convolutional Cross-Component Model (CCCM) - Single Model and Multiple Models In CCCM, convolutional models are used to improve chromaticity prediction performance. The convolutional model has a 7-tap filter, which consists of a 5-tap plus shapespace component, a nonlinear term, and a bias term. The input to the filter's 5-tap shapespace component includes the center (C) luminance sample co-located with the chromaticity sample to be predicted, and its above / north (N), below / south (S), left / west (W), and right / east (E) neighbors, such as... Figure 4 As shown.
[0026] The non-linear term (denoted as P) is expressed as a power of two of the center brightness sample C, scaled to the range of sample values for the content: , For example, for 10 bits of content, the nonlinear term is calculated as follows: , The offset term (labeled B) represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10 bits).
[0027] The output of the filter is calculated as the filter coefficients c. i The convolution between the input and output values, and the cropping to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B Filter coefficients c i It is calculated by minimizing the mean squared error (MSE) between the predicted and reconstructed chromaticity samples in the reference region. Figure 5 This example shows a reference region consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and downwards by one PU height beyond the PU boundary. The region is adjusted to include only available samples. This expansion of the region (labeled "padding") is to support... Figure 4 The plus sign shape space filter's "side sample" is used to fill in unavailable areas.
[0028] MSE minimization is performed by calculating the autocorrelation matrix of the luminance input and the cross-correlation vector between the luminance input and chrominance output. The autocorrelation matrix is decomposed using LDL, and the final filter coefficients are calculated using back-substitution. This process roughly follows the calculation of ALF filter coefficients in ECM, but LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.
[0029] Similarly, like CCLM, CCCM offers the option to use a single model or a multi-model variant. The multi-model variant uses two models: one derived from samples above the average luminance reference value, and the other derived from the remaining samples (following the spirit of CCLM's design). The multi-model CCCM mode can be selected when at least 128 reference samples are available for the PU.
[0030] Gradient Linear Model (GLM) For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from the gradient of luminance samples. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0031] Compared to CCLM, GLM does not use downsampled luminance values, but instead uses the gradient of luminance samples to derive a linear model. Specifically, when GLM is applied, CCLM processes the input, i.e., downsampled luminance samples... The gradient of the brightness sample Replacement. Other parts of CCLM (e.g., parameter derivation, linear transformation of predicted samples) remain unchanged.
[0032] , In the three-parameter GLM, chromaticity samples can be predicted based on different parameters of the luminance sample gradient and downsampled luminance values. The model parameters of the three-parameter GLM are derived from 6 row and column adjacent samples using the MSE minimization method based on LDL decomposition, similar to the method used in CCCM.
[0033] , For transmission, when the current CU has CCLM mode enabled, a flag is sent to indicate whether GLM is enabled for the Cb and Cr components; if GLM is enabled, another flag is sent to indicate which of the two GLM modes is selected, and a syntax element is further sent to select one of the four gradient filters for gradient calculation.
[0034] If GLM is enabled for a component, a syntax element is further sent to select four gradient filters. Figure 6 Gradient calculation is performed using a gradient filter in (610-640) of the algorithm.
[0035] Spatial candidate export The derivation of spatial merge candidates in VVC is the same as in HEVC, except that the positions of the first two merge candidates are swapped. For the current CU 710, from Figure 7A maximum of four merge candidates (B0, A0, B1, and A1) are selected from the positions depicted. The derived order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more adjacent CUs of positions B0, A0, B1, and A1 are unavailable (e.g., belonging to another slice or tile) or if it is intra-frame encoding / decoding. After a candidate for position A1 is added, the addition of remaining candidates is limited by redundancy checks to ensure that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency.
[0036] Temporal Candidates Derivation In this step, only one candidate is added to the list. Specifically, when exporting this temporal merge candidate for the current CU 810, a scaled motion vector is exported based on the co-located CU 820 belonging to the co-located reference image, such as... Figure 8 As shown. The list of reference images and reference indices used to export the corresponding CUs are explicitly sent in the slice header. For example... Figure 8 As shown by the dashed line, the scaled motion vector 830 of the temporal merging candidate is obtained, which is scaled from the motion vector 840 of the co-located CU using Picture Order Count (POC) distances tb and td. Here, tb is defined as the POC difference between the current image and its reference image, and td is defined as the POC difference between the co-located image and its reference image. The reference image index of the temporal merging candidate is set to zero.
[0037] The position of the time candidate is selected between candidate C0 and C1, such as... Figure 9 As shown. If the CU at position C0 is unavailable, is an intra-frame codec, or is outside the current line of the CTU, then position C1 is used. Otherwise, position C0 is used when exporting time merge candidates.
[0038] Non-adjacent spatial candidate During the development of the VVC standard, a codec tool called Non-Adjacent Motion Vector Prediction (NAMVP) was proposed in JVET-L0399 (Yu Han et al., “CE4.4.6: Improvement on Merge / Skip mode”, 12th meeting of the Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC29 / WG 11, Macau, China, October 3-12, 2018, document: JVET-L0399). According to NAMVP technology, non-adjacent spatial merge candidates are inserted into the regular merge candidate list after the TMVP (Temporal MVP). The spatial merge candidate pattern is as follows... Figure 10 As shown. The distance between a non-adjacent spatial candidate and the current codec block is based on the width and height of the current codec block. Figure 10 In this model, each small square corresponds to a NAMVP candidate, and the candidates are sorted by distance (as shown by the numbers within the square). Line buffer limits do not apply. In other words, NAMVP candidates far from the current block may need to be stored, which could require a large buffer. Summary of the Invention
[0039] A method and apparatus for encoding and decoding color images using an encoding / decoding tool that includes one or more cross-component model correlation modes are disclosed. Input data associated with a current block, comprising a first color block and a second color block, is received, wherein the input data includes pixel data to be encoded at the encoder end or data associated with the current block to be decoded at the decoder end, and the current block is located in a non-intra-frame slice or picture. Data or a location associated with reference data is determined, the location being located in a previously encoded / decoded slice / picture, wherein the reference data associated with the location is confined within a co-location codec tree unit (CTU) of the current block and one or more adjacent regions adjacent to the co-location CTU. Multiple target cross-component model parameters associated with a target inherited prediction model are derived by inheriting multiple cross-component model parameters from the reference data. The second color block is encoded or decoded using prediction data including cross-color prediction, which is generated by applying the target inherited prediction model having the multiple target cross-component model parameters to the reconstructed first color block. The first color block may correspond to a luma block, and the second color block may correspond to a chroma block.
[0040] In one embodiment, the one or more adjacent regions adjacent to the peer CTU include a top N horizontal reference line region above the peer CTU, a left N vertical reference line region adjacent to the left side of the peer CTU, a bottom N horizontal reference line region to the left of the peer CTU and aligned with the bottom of the peer CTU, or a combination thereof, where N is a positive integer. In one embodiment, N is set to a predetermined value. In another embodiment, N is set to the minimum allowed block size. In one embodiment, the minimum allowed block size is associated with a codec unit (CU), prediction unit (PU), or transform unit (TU).
[0041] In one embodiment, the one or more adjacent regions adjacent to the co-located CTU are the same as the allowed regions for the reference motion vectors of the temporal candidate in inter-frame mode. In one embodiment, the reference data in the previously encoded slice or picture is located based on the corresponding position of the current block in the previously encoded slice or picture and the motion vectors of the current block or adjacent blocks. In one embodiment, when the current block is intra-frame encoded, the motion vectors used to derive the plurality of target cross-component model parameters of the current block are set to 0.
[0042] In one embodiment, the position of the reference data is determined based on the corresponding position of the current block moved via the motion vector of the current block. In another embodiment, the position of the reference data is determined based on the corresponding position of the current block moved via the motion vector of the current block and one or more additional offsets. In one embodiment, the one or more additional offsets depend on the width of the current block, the height of the current block, or both. In one embodiment, the one or more additional offsets correspond to a horizontal offset of the width of the current block and a vertical offset of the height of the current block. In another embodiment, the one or more additional offsets correspond to a horizontal offset of half the width of the current block and a vertical offset of half the height of the current block.
[0043] In one embodiment, the one or more cross-component model parameters are optimized and used as the plurality of target cross-component model parameters. In one embodiment, the first color patch corresponds to a luminance patch, and the second color patch corresponds to a chrominance patch.
[0044] In one embodiment, only the multiple cross-component model parameters of the target co-image within the co-located CTU are allowed to be referenced by multiple time candidates. In another embodiment, only the multiple cross-component model parameters of the target co-image within the co-located CTU, the N CTUs to the left of the co-located CTU, the M CTUs to the right of the co-located CTU, or a combination thereof, are allowed to be referenced by multiple time candidates, and N and M are integers greater than 0. In yet another embodiment, only the multiple cross-component model parameters of the target co-image within the CTU row containing the co-located CTU are allowed to be referenced by multiple time candidates. In yet another embodiment, only the multiple cross-component model parameters of the target co-image corresponding to the current CTU row, the N CTU rows above the current CTU row, the M CTU rows below the current CTU row, or a combination thereof, are allowed to be referenced, and N and M are integers greater than 0. Attached Figure Description
[0045] Figure 1A An exemplary adaptive inter-frame / intra-frame video encoding and decoding system incorporating loop processing is demonstrated.
[0046] Figure 1B exhibit Figure 1A The corresponding decoder for the encoder.
[0047] Figure 2 This shows examples of the samples to the left and above the current block involved in the LM_LA mode, as well as examples of sample positions.
[0048] Figure 3 This example demonstrates how adjacent samples are classified into two groups.
[0049] Figure 4 An example showing the spatial portion of a convolutional filter.
[0050] Figure 5 This shows an example of the reference region and padding used to derive the filter coefficients.
[0051] Figure 6 This section demonstrates four gradient modes of the Gradient Linear Model (GLM).
[0052] Figure 7 Displays adjacent blocks used to derive VVC space merge candidates.
[0053] Figure 8 This example demonstrates the temporal candidate export, where scaled motion vectors are derived based on the image order count (POC) distance.
[0054] Figure 9 Show the position of the time candidate selected between candidate C0 and C1.
[0055] Figure 10Exemplary patterns of non-adjacent spatial merging candidates are shown.
[0056] Figure 11 An example of CCM information propagation is shown, where dashed blocks (i.e., A, E, G) are encoded and decoded in cross-component modes (e.g., CCLM, MMLM, GLM, CCCM).
[0057] Figure 12 This example demonstrates the inheritance of time-adjacent model parameters.
[0058] Figure 13 An example of a time candidate location is shown according to an embodiment of the present invention.
[0059] Figure 14 An example of a time-candidate available region is shown according to an embodiment of the present invention.
[0060] Figure 15A -B demonstrates two search modes for inheriting the non-closely adjacent spatial adjacency model.
[0061] Figure 16 A flowchart illustrating an exemplary video encoding / decoding system is provided, which, according to an embodiment of the invention, combines inherited temporal cross-component model parameters with buffer constraints. Detailed Implementation
[0062] The components of the present invention, as generally described and illustrated in the figures, can be arranged and designed in a variety of different configurations. Therefore, the following more detailed description of embodiments of the systems and methods of the present invention, as illustrated in the figures, is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the invention. References to “one embodiment,” “an embodiment,” or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of the invention. Therefore, the phrases “in one embodiment” or “in one embodiment” appearing throughout this specification do not necessarily refer to the same embodiment.
[0063] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. However, those skilled in the art will recognize that the invention can be practiced without using one or more specific details, or using other methods, components, etc. In other instances, well-known structures or operations are not shown or described in detail to avoid obscuring aspects of the invention. Embodiments of the invention will be best understood by referring to the accompanying drawings, wherein like parts are designated by like numbers throughout. The following description is illustrative only and only illustrates embodiments of specific selected apparatus and methods consistent with the invention declared herein.
[0064] To improve the prediction accuracy or encoding / decoding performance of cross-component prediction, various schemes related to inheriting cross-component models have been revealed.
[0065] Inheriting adjacent model parameters When a cross-component predictive codec is applied to the current block to generate a predicted signal, the cross-component model (CCM) information (more details in the section titled "Inheriting CCM Information") contains model parameters that can be inherited from neighboring blocks. Further details about neighboring blocks are described in the sections titled "Inheriting Spatially Adjacent Model Parameters," "Inheriting Temporally Adjacent Model Parameters," "Inheriting Non-Closely Adjacent Spatially Adjacent Models," and "Inheriting Model Parameters from History Tables."
[0066] The final scaling parameters for the current block are inherited from neighboring blocks and / or further optimized by dA. Once the final scaling parameters are determined, the offset parameters (e.g., in CCLM) are... This is derived based on inherited scaling parameters and / or the average of the adjacent luminance and chrominance samples of the current block. For example, if the final scaling parameters are inherited from selected adjacent blocks, and the inherited scaling parameters are... So the final scaling parameter is ( +dA). In another embodiment, the final scaling parameter is inherited from the history list. For example, the history list records the final scaling parameters of the most recent j entries from previous CCLM codec blocks. The final scaling parameter is then derived from the history list ( A selected entry is inherited, and the final scaling parameter is ( +dA). In another embodiment, the final scaling parameter is inherited from the history list or adjacent blocks, but is not further optimized by dA.
[0067] In another embodiment, after inheriting the model parameters, the offset parameters can be further optimized by dB. For example, if the final offset parameters are inherited from selected neighboring blocks, and the inherited offset parameters are... So the final offset parameter is ( + dB). In another embodiment, the final offset parameter is inherited from the history list and further optimized by dB. For example, the history list records the final offset parameters of the most recent j entries from previous CCLM codec blocks. The final offset parameter is then derived from the history list ( A selected entry is inherited, and the final offset parameter is ( + dB). In another embodiment, the final offset parameter is inherited from the history list or adjacent blocks, but is not further optimized by dB.
[0068] In another embodiment, if the inherited adjacent blocks are encoded and decoded using CCCM, then the inherited filter coefficients ( Offset parameters (e.g., in CCCM) or It can be re-derived based on the inherited parameters and the average value of the brightness and chromaticity samples of the adjacent corresponding positions of the current block.
[0069] In another embodiment, if the inherited candidate applies the GLM gradient mode to its luminance reconstruction sample, the current block should also inherit the candidate's GLM gradient mode and apply it to the current luminance reconstruction sample.
[0070] In another embodiment, if the inherited neighboring blocks are encoded and decoded using multiple cross-component models (e.g., MMLM, or multi-model CCCM), the classification threshold is also inherited to classify the neighboring samples of the current block into multiple groups, and the inherited multiple cross-component model parameters are further assigned to each group.
[0071] In another embodiment, when the current slice is a non-intra-frame slice (e.g., a P-slice or a B-slice), the cross-component model (CCM) information of the current block is exported and stored for later reconstruction of neighboring blocks using inherited neighbor model parameters. In another embodiment, when the current block is an inter-frame codec, the CCM information of the current inter-frame codec block is exported by copying the CCM information from a reference block that has CCM information in a reference picture, which is located using the motion information of the current inter-frame codec block. For example, as... Figure 11 As shown, block B in P / B image 1120 is an inter-frame codec, so the CCM information of block B is obtained by copying the CCM information from reference block A in I image 1110. It should be noted that the current block can also copy CCM information from intra-frame codec blocks in the P / B image. For example, as... Figure 11 As shown, block D in P / B image 1130 is inter-frame encoded, so the CCM information of block B is obtained by copying the CCM information from reference block E, which is intra-frame encoded in P / B image 1111. In another embodiment, if the reference block in the reference image is also inter-frame encoded, the CCM information of the reference block is obtained by copying the CCM information from another reference block in another reference image. For example, as... Figure 11As shown, in the current P / B image 1130, the current block C is inter-frame encoded, and its reference block B is also inter-frame encoded. Since the CCM information of block B is obtained by copying the CCM information from block A, the CCM information of block A is also propagated to the current block C. It should be noted that block C only needs to access block B to obtain the original CCM information from block A, because the original CCM information from block A is already stored in block B. In another embodiment, when the current block is inter-frame encoded with bidirectional prediction, if one of its reference blocks is intra-frame encoded and has CCM information, the CCM information of the current block is obtained by copying the CCM information from the intra-frame encoded reference block in the reference image. For example, suppose block F is inter-frame encoded with bidirectional prediction and has reference blocks G and H. Block G is intra-frame encoded and has CCM information. The CCM information of block F is obtained by copying the CCM information from block G, which is encoded in CCM mode. In another embodiment, when the current block is an inter-frame codec with bidirectional prediction, the CCM information of the current block is a combination of the CCM models of its reference blocks.
[0072] Inheriting CCM information In one embodiment, inherited cross-component model (CCM) information can be stored along with inherited model parameters. As mentioned earlier in this document, CCM information includes, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM), model indices indicating which model shape to use in the convolutional model, classification thresholds for multiple models, downsampling filter flags, downsampling filter indices, the number of adjacent lines used to derive the model, the template type used to derive the model, post-filter flags, or model parameters.
[0073] In one embodiment, a CCLM model can be inherited. In addition to storing model parameters, a prediction pattern can also be stored in the CCM information to indicate that the inherited model is a CCLM model.
[0074] In another embodiment, a CCLM model with nonlinear terms can be inherited. In addition to storing model parameters, prediction patterns can also be stored in the CCM information to indicate that the inherited model is a CCLM model with nonlinear terms.
[0075] In one embodiment, a CCCM model can be inherited. In addition to storing model parameters, prediction patterns can also be stored in the CCM information to indicate that the inherited model is a CCCM model. Luminance and chromaticity offsets used to adjust the CCCM model inputs can also be stored in the CCM information.
[0076] In another embodiment, CCCM models with different convolutional filter shapes can be inherited. In addition to model parameters and prediction modes, a CCCM mode index can be stored in the CCM information to indicate which convolutional filter shape the inherited CCCM model uses. For example, a CCCM model with different convolutional filter shapes may contain only horizontal spatial terms. Another example is that a CCCM model with different convolutional filter shapes may contain only vertical spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes may contain only diagonal spatial terms. Another example is that a CCCM model with different convolutional filter shapes may contain only anti-diagonal spatial terms. Yet another example is that a CCCM model with different convolutional filter shapes may contain X-shaped spatial terms.
[0077] In another embodiment, a CCCM model using unsampled samples can be inherited. In addition to storing model parameters, prediction patterns can also be stored in the CCM information to indicate that the inherited model is a CCCM model using unsampled samples.
[0078] In another embodiment, a CCCM model with multiple downsampling filters can be inherited. In addition to storing model parameters, the CCM information can also store a prediction mode indicating that the inherited model is a CCCM model with multiple downsampling filters, and a model index indicating which variant of the CCCM model with multiple downsampling filters is being inherited.
[0079] In another embodiment, a hybrid CCCM model consisting of various terms (e.g., spatial, gradient, positional, nonlinear, and bias terms) can be inherited. The gradient term can be computed in either a downsampled or non-downsampled domain. The positional term can be computed relative to the top-left corner coordinates of the current block or the current image. In addition to storing model parameters, prediction patterns can be stored in the CCM information to indicate that the inherited model is a hybrid CCCM model consisting of various terms. If multiple types of hybrid CCCM models exist, model indexes can also be stored in the CCM information to indicate which type of hybrid CCCM model is being inherited. For example, the gradient and location-based cross-component model (GL-CCCM) proposed in JVET-AB0119 (Ramin G. Youvalari et al., “Non-EE2: Gradient and location-based convolutional cross-component model (GL-CCCM) for intra prediction”, Joint Video Exploration Group (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, 28th meeting, Mainz, DE, October 20–28, 2022, document: JVET-AB0119) is a hybrid CCCM model consisting of a spatial term for the central location, two gradient terms for the horizontal and vertical directions, two location terms X and Y for the relative horizontal and vertical locations, a nonlinear term, and a bias term. In addition to storing model parameters, the prediction pattern can also be stored in the CCM information to indicate that the inherited model is a GL-CCCM model.
[0080] In one embodiment, a GLM model can be inherited. In addition to storing model parameters, the CCM information can also store a prediction mode indicating that the inherited model is a GLM model, and a downsampling filter index indicating which gradient downsampling filter the inherited GLM model uses.
[0081] In another embodiment, a GLM model with a luminance term can be inherited. In addition to storing model parameters, the prediction mode can also be stored in the CCM information to indicate that the inherited model is a GLM model with a luminance term, and the downsampling filter index can also be stored in the CCM information to indicate which gradient downsampling filter is used for the inherited GLM model with a luminance term.
[0082] In one embodiment, any type of cross-component multi-model can be inherited. In addition to storing model parameters and prediction patterns, a multi-model on / off flag can be stored in the CCM information to indicate whether the inherited CCM model is a multi-model. If the multi-model on / off flag is true, the multi-model classification threshold is also stored in the CCM information.
[0083] In one embodiment, CCM information may include information indicating how to export the inherited model. For example, CCM information may include the number of adjacent lines used to export the cross-component model and / or the template type used to export the model. For instance, a set of templates can be used to export the CCCM model. This set of templates includes templates with different positions, sizes, and shapes. CCM information may store an index of which template the inherited CCCM model was exported from. For example, the inherited CCCM model may be exported based on a top-only template, a left-only template, or both a left and top template. Another example is that the inherited CCCM model may be exported based on a 6-line template or a 2-line template.
[0084] In one embodiment, the post-filter flag can be stored in the CCM information. This information describes how the inherited model is used in its source block. If the post-filter flag is enabled, it indicates that the filter is used for predictions in the block where the inherited model resides.
[0085] Optimization of inheritance model parameters In one embodiment, the inherited model parameters can be further optimized based on the inherited CCM information. The inherited CCM information may include information on how the inherited model was derived, such as the template type and / or the number of adjacent lines used to derive the model. The optimized parameters are derived based on local information. The optimization process can follow the deriving method of the inherited model, and use the same type of template and / or the same number of adjacent lines. For example, if the inherited model is CCLM, and it is derived based on a left-side-only template (the inherited model is CCLM_L), then the offset parameters... The average value of the reconstructed samples can be derived from the adjacent left template of the current block. For example, if the inherited model is CCLM, and it is derived based on only the top template (the inherited model is CCLM_T), then the offset parameter... The average value of the samples can be reconstructed from the adjacent top template of the current block. Another example is if the inherited model is CCCM, and the model is exported using a 2-line template. Offset values (e.g., in CCCM) The samples can be reconstructed and re-exported based on the 2-row template of the current block. Another example is if the inherited model is a multi-model model (e.g., MMLM, CCCM with multiple models), and the classification threshold is based on only the left-hand template; in this case, the samples can be reconstructed and re-exported based on the left-hand template of the current block.
[0086] In one embodiment, the inherited model parameters can be further optimized using different types of templates and / or different numbers of lines, and the final model parameters are determined by the template cost. The template cost is calculated by applying candidate optimized model parameters to neighboring templates to predict template samples and computing the difference (e.g., SAD or SATD) between the predicted and reconstructed samples. For example, if the inherited model is CCLM, the optimized offset parameters... Export using the left, top, and left-top templates of the reconstructed sample for the current block, respectively. If applied... The cost of the template is The minimum value in, then It was selected as the final offset parameter.
[0087] In another embodiment, the inherited model parameters can be further optimized by predetermined values. Template cost is used to determine whether to further optimize the inherited model parameters. Template cost is calculated by applying candidate optimized model parameters to adjacent templates to predict template samples and computing the difference (e.g., SAD or SATD) between the predicted and reconstructed samples. For example, for the CCCM pattern, for each inherited model parameter... This value is obtained through Optimize and apply and The template cost is calculated to determine which is the final model parameter value.
[0088] Inheritance space adjacency model parameters In another embodiment, the inherited model parameters can come from a directly adjacent block. Models from blocks at predetermined locations are added to the candidate list in a predetermined order. For example, the predetermined location could be... Figure 7 The positions depicted in the text can be in the following predetermined order: B0, A0, B1, A1 and B2, or A0, B0, B1, A1 and B2.
[0089] In one embodiment, the predetermined position and predetermined order can be the same as the spatial candidates for the inter-frame merging mode.
[0090] In another embodiment, assuming the current block's position, width, and height are (x, y), W, and H respectively, if W is greater than or equal to TH, the predetermined position can include a position directly above the current block (W>>1) or ((W>>1) – 1) (i.e., a position directly above the current block, such as ((x + W)>>1, y-1) or ((x + W+1)>>1, y-1)), and a position directly to the left of the current block (H>>1) or ((H>>1) – 1) (i.e., a position directly to the left of the current block, such as (x-1, (y+H)>>1) or (x-1, (y+H+1)>>1)), where W and H are the width and height of the current block, and TH is a threshold, which can be 2, 4, 8, 16, 32, or 64. In another embodiment, the maximum number of models inherited from spatial neighbors is less than the number of predetermined positions. For example, if the predetermined position is as follows... Figure 7 As shown, there are 5 predefined positions. If the predefined order is B0, A0, B1, A1 and B2, and the maximum number of models inherited from spatial neighbors is 4, then a model from B2 is added to the candidate list only if the preceding block is unavailable or not encoded / decoded in the cross-component model.
[0091] Inheritance time adjacent model parameters In one embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images.
[0092] In another embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images. For example, such as Figure 12 As shown, the current block position is (x, y), and the block size is The inherited model parameters can come from blocks at positions (x', y'), (x', y'+ h / 2), (x' + w / 2, y'), (x' + w / 2, y' + h / 2), (x' + w, y'), (x', y' + h), or (x' + w, y' + h) in previously encoded / decoded slices / images, where x' = x + Δx and y' = y + Δy. In one embodiment, if the prediction mode of the current block is intra-frame, Δx and Δy are set to 0. If the prediction mode of the current block is inter-frame prediction, Δx and Δy are set to the horizontal and vertical motion vectors of the current block. In another embodiment, if the current block is bi-directional inter-frame prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 0. In yet another embodiment, if the current block is bi-directional inter-frame prediction, Δx and Δy are set to the horizontal and vertical motion vectors in reference image list 1.
[0093] In another embodiment, if the current block is a bidirectional inter-frame prediction, the inherited model parameters can come from blocks in previously encoded / decoded slices / images in the reference list. For example, if the horizontal and vertical motion vectors in reference image list 0 are... and Then the motion vector can be scaled to other reference images in reference lists 0 and 1. If the motion vector is scaled to the i-th reference image in reference list 0, then... The model can then be derived from the block in the i-th reference image in reference list 0, and Δx and Δy are set to... Another example is if the horizontal and vertical motion vectors in reference image list 0 are... and And the motion vector is scaled to the i-th reference image in reference list 1. The model can then be derived from the block in the i-th reference image in reference list 1, and Δx and Δy are set to... .
[0094] In one embodiment, if the current slice / image is a non-intra-frame slice / image, the inherited model parameters can come from blocks in previously encoded / decoded slices / images. In one embodiment, the current block position is (x, y), and the block size is... Two value sets and Defined as: , .
[0095] exist and All values in the set are positive. Let... Inherited model parameters can come from positions in previously encoded / decoded slices / images. The block.
[0096] In one sub-implementation, .For example, Inherited model parameters can come from previously encoded / decoded slices / images. Figure 13 The block at the indicated location. Figure 13 In the text, the current block is represented by a thick outline, indicating the current block's position. Represented by a black dot in the center. The positions of previously encoded / decoded slices / images are represented by "X", gray diamonds, gray squares, gray triangles, and gray circles.
[0097] In another sub-implementation, .For example, and .
[0098] In another embodiment, the current block is located at (x, y) (i.e., the top-left corner of the current block), and the block size is... Inherited model parameters can come from positions in previously encoded / decoded slices / images. The block.
[0099] In one sub-implementation, .For example, .
[0100] In another sub-implementation, .For example, and .
[0101] In one embodiment, closer The model at the location is first added to the final merge candidate list. In another embodiment, closer to The model at the location is first added to the final merge candidate list.
[0102] In one embodiment, it is set and These are two fixed positive numbers. The inherited model parameters can come from positions in previous encoder / decoder slices / images. The block.
[0103] In another embodiment, the current block position is (x, y), and the block size is .set up and These are two fixed positive numbers. The inherited model parameters can come from positions in previous encoder / decoder slices / images. The block.
[0104] In another embodiment, the current block position is (x, y), and the block size is The inherited model parameters can come from some predetermined locations in the previous codec slices / images. The block. For example, the position is within the corresponding region of the current encoded block, i.e. and Inherited model parameters can come from blocks. Another example is when the location is outside the corresponding region of the current coding block, i.e. and Inherited model parameters can come from blocks. .
[0105] In one embodiment, the previously encoded / decoded image (i.e., the same image) from which the parameter model is inherited is one of the images in the reference list.
[0106] In one embodiment, the co-position image is marked in the image / slice header. A reference list and reference index are also marked in the image / slice header. For example, the co-position image is selected as L0[0]. In another example, the co-position image is selected as L1[0].
[0107] In one embodiment, the corresponding image is selected as the image with the smallest POC difference between the current image and the reference list. For example, if the current image has a POC of 8, the images in reference list 0 have POCs of {7, 6, 5, 0}, and the images in reference list 1 have POCs of {7, 6, 5, 4}, then L0[0] (equivalent to L1[0]) is selected because it has the smallest POC difference. In one sub-embodiment, if there are two images with the smallest POC difference between them and the current image, the image with the smaller POC is selected. In another sub-embodiment, if there are two images with the smallest POC difference between them and the current image, the image with the larger POC is selected. In yet another sub-embodiment, if there are two images with the smallest POC difference between them and the current image, the image with the smaller QP difference is selected. In yet another sub-embodiment, if there are two images with the smallest POC difference between them and the current image, the image with the smaller QP is selected. In another sub-implementation, if there are two images whose POC difference with the current image is the smallest, then the image with the larger QP is selected.
[0108] In one embodiment, the corresponding image is selected as the image with the smallest QP difference between the current image and the reference list. For example, if the current image has a QP of 28, the images in reference list 0 have QPs of {19, 26, 23}, and the images in reference list 1 have QPs of {23, 22, 21}, then L0[1] is selected. In another sub-embodiment, if more than one image in the reference list has the smallest QP difference with the current image, the image with the smaller QP is selected. In another sub-embodiment, if more than one image has the smallest QP difference with the current image, the image with the larger QP is selected. In another sub-embodiment, if more than one image has the smallest QP difference with the current image, the image with the smaller POC distance is selected.
[0109] In one embodiment, the corresponding image is selected as the image with the smallest QP in the reference list. In another embodiment, the corresponding image is selected as the image with the largest QP in the reference list.
[0110] In one embodiment, the previously encoded / decoded image (i.e., the co-image) from which the inherited parameter model originates is the most recently encoded / decoded I-image. The cross-component model information of the most recently encoded / decoded I-slice / image is stored in a long-term reference buffer.
[0111] In one embodiment, the source location of the co-located image and the inherited parametric model is determined by the motion vectors of neighboring blocks. For example, if the current block position is (x, y) and the block size is... The inherited model parameters can come from blocks in the co-located image at positions (x', y'), (x', y'+ h / 2), (x'+w / 2, y'), (x'+ w / 2, y'+ h / 2), (x'+ w, y'), (x', y'+ h), or (x'+w, y'+ h), where x' = x + Δx and y' = y + Δy. Δx and Δy are set as the horizontal and vertical portions of the L0 motion vectors of the adjacent blocks, respectively, and the co-located image is an L0 reference image indicated by the L0 motion vectors of the adjacent blocks. In another embodiment, if the adjacent blocks are inter-frame bidirectional predictions, Δx and Δy are set as the horizontal and vertical portions of the L1 motion vectors of the adjacent blocks, respectively, and the co-located image is an L1 reference image indicated by the L1 motion vectors of the adjacent blocks. In one embodiment, the adjacent block is the block to the left of the current block. In another embodiment, the adjacent block is the block above the current block.
[0112] In one embodiment, the inherited model parameters are derived using luminance and chrominance reconstruction samples from co-position blocks. Let the current block position be (x, y) and the block size be... When the inherited model comes from position (x', y'), the corresponding block is the position (x', y') in the corresponding image and the block size is... A corresponding block can be located within a block. For example, a corresponding block can be located at (x, y). Another example is where, if Δx and Δy are the horizontal and vertical components of the L0 motion vectors of adjacent blocks, respectively, and the corresponding picture is an L0 reference picture indicated by the L0 motion vectors of adjacent blocks, the corresponding block can be located within the corresponding picture. .
[0113] In one embodiment, the position of the inherited parameter model from the previous codec slice / image is determined by the motion vectors of neighboring blocks. Let Δx and Δy be the horizontal and vertical displacements determined based on the motion vectors of the selected neighboring blocks, the current block position is (x, y), and the block size is... The inherited model parameters can come from the block at position (x', y'), where x' = x + Δx and y' = y + Δy, or from the block at position x' = x + w / 2 + Δx and y' = y + h / 2 + Δy.
[0114] In another embodiment, the inherited model parameters can also be derived from block positions in the pattern described in the preceding paragraphs. These positions are centered at (x', y'), where x' = x + Δx and y' = y + Δy, or x' = x + w / 2 + Δx and y' = y + h / 2 + Δy. That is, the predetermined positions are represented as... The inherited model parameters can come from ,in and The horizontal and vertical displacements are determined based on the motion vectors of the selected neighboring blocks. For example, suppose the size of the current block is... Two value sets and Defined as: ; ; All in and The values in the table are all positive. Inherited model parameters can be derived from positions in previous encoder / decoder slices / images. The block. For example, let's say... and These are two fixed positive numbers. The inherited model parameters can be derived from positions in previous encoder / decoder slices / images. The block. Another example is that inherited model parameters can come from the previous codec slice / image. Some pre-reserved locations. These locations may be... Another example is that these locations could be... .
[0115] In one embodiment, adjacent blocks can be located at predetermined positions. For example, this position could be as follows: Figure 7 The A0 position is shown. The predetermined position can also be as follows: Figure 7 The numbers A1, B0, B1, and B2 are shown. If the block at the predetermined position is not an inter-frame block, then adjacent blocks are not selected.
[0116] In another embodiment, when selecting adjacent blocks, there may be a list of predetermined positions. These positions are placed according to the order of inspection. For example, the positions may be spatial locations described in the section titled "Inherited Spatial Adjacency Model Parameters" (e.g., in...). Figure 7 (B0, A0, B1, A1, and B2). The selected adjacent block can be the first one in the list that is an inter-frame block. First, the L0 motion vector is selected. If the L0 motion vector is unavailable, the L1 motion vector is selected. Another example: first, the L1 motion vector is selected. If the L1 motion vector is unavailable, the L0 motion vector is selected.
[0117] Another example is where the positions in the list are checked in a predetermined order. For each position, the L0 motion vector is checked first, followed by the L1 motion vector. Another example is where the L1 motion vector is checked first, followed by the L0 motion vector. The selected motion vector is the first one whose reference image is a co-image. The co-image can be determined using the methods described in the preceding paragraphs of this section.
[0118] In one embodiment, the horizontal and vertical displacements Δx and Δy are determined based on the motion vectors of selected neighboring blocks. For example, if the reference image and the co-located image for the selected motion vectors are the same image, Δx equals the horizontal portion of the selected motion vector, and Δy equals the vertical portion of the selected motion vector. If the horizontal or vertical portion of the selected motion vector is a fraction, Δx equals the rounded horizontal portion of the selected motion vector, and Δy equals the rounded vertical portion of the selected motion vector. The rounding method used can be, but is not limited to, the following: rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding up by half, rounding down by half, etc.). Another example: if the reference image and the corresponding image of the selected motion vector are not the same, the reference image can be one of the images in the reference list, while the corresponding image is marked in the image / slice header. Let the POC distance between the current image and the reference image of the selected motion vector be tb, the POC distance between the current image and the corresponding image be td, and the selected motion vector be (mv_x, mv_y). Δx = mv_x (td / tb) and Δy =mv_y (td / tb). If mv_x (td / tb) or mv_y (td / tb) is a fraction, and Δx equals mv_x. (td / tb) or the result of rounding the horizontal part of the selected motion vector, Δy equals mv_y (td / tb) or the result of rounding the vertical portion of the selected motion vector. The rounding method used may be, but is not limited to, the following: rounding to negative infinity, rounding to positive infinity, rounding to zero, or rounding to the nearest integer (e.g., rounding away from zero, rounding up by half, rounding down by half, ...).
[0119] Available areas for time candidates To limit the requirements for buffer / storage resources, the available range of temporal candidates should be limited. The temporal candidates mentioned in this section refer to candidates that inherit model parameters from blocks in previously encoded / decoded slices / pictures, as described in the section titled "Inheriting Temporally Adjacent Model Parameters." For example, assuming the current block position is (x, y), the inherited parameter model could be derived from a previously encoded / decoded slice / picture at a position of (x + Δx). i , y +Δyi), where i is from 1 to M, and M is a positive integer greater than 0. Δx i and Δy i This is a predetermined displacement. Another example, assuming the current block position is (x, y), the position in the previously encoded / decoded slice / image from which the inherited parametric model comes could be (x + dx + Δx). i , y + dy + Δy i ), where i ranges from 1 to M, and M is a positive integer greater than 0. Δx i and Δy i This is a predetermined displacement. dx and dy are determined by the motion vectors of the current block's neighboring blocks. Details on how the motion vectors are determined are described in a section titled "Inheriting Temporally Adjacent Model Parameters." Another example, assuming the current block is located at (x, y), the inherited parametric model's position in the previously encoded / decoded slice / image could be (x + dx + Δx). i , y + dy + Δy i ), where i ranges from 1 to M, and M is a positive integer greater than 0. Δx i and Δy iThis is a predetermined displacement. If the prediction mode of the current block is inter-frame, dx and dy are set to the horizontal and vertical portions of the motion vector of the current block. If the horizontal or vertical portion of the motion vector is a fraction, dx is set to the rounded result of the horizontal portion of the motion vector, and dy is set to the rounded result of the vertical portion of the motion vector. The rounding method used can be, but is not limited to, the following: rounding towards negative infinity, rounding towards positive infinity, rounding towards zero, rounding to the nearest integer (e.g., rounding away from zero, rounding up by half, rounding down by half, ...), or rounding to the nearest predetermined precision (e.g., rounding to the nearest k-pixel or 1 / k-pixel precision position, where k can be 2, 4, 8, 16, or 32). If the prediction mode of the current block is IBC, dx and dy are set to the horizontal and vertical block vectors of the current block. If the horizontal or vertical portion of the block vector is a fraction, dx is set to the rounded result of the horizontal portion of the block vector, and dy is set to the rounded result of the vertical portion of the block vector. The rounding method used may be, but is not limited to, the following: rounding to negative infinity, rounding to positive infinity, rounding to zero, rounding to the nearest integer (e.g., rounding away from zero, rounding up by half, rounding down by half, ...), or rounding to the nearest predetermined precision (e.g., rounding to the nearest k-pixel or 1 / k-pixel precision position, where k may be 2, 4, 8, 16, or 32).
[0120] In one embodiment, only the Cross-Component Model (CCM) information of the CTU in the co-located image, whose position corresponds to the position of the currently encoded CTU in the current image, can be referenced by the time candidate. In another embodiment, only the CCM information of the CTU in the co-located image, whose position corresponds to the position of the currently encoded CTU, and / or the N CTUs to the left of the currently encoded CTU, and / or the M CTUs to the right of the currently encoded CTU, can be referenced by the time candidate, where N and M can be any integer greater than 0. In another embodiment, only the CCM information of the CTU row in the co-located image, whose position corresponds to the position of the currently encoded CTU row in the current image, can be referenced by the time candidate. In yet another embodiment, only the positions in the co-located image corresponding to the current CTU row and / or the N CTU rows above and / or the M CTU rows below the current CTU row can be referenced, where N and M can be any integer greater than 0. Please note that, as described under the headings “Inheriting CCM Information” and “Optimization of Inherited Model Parameters”, the CCM information mentioned in this disclosure includes, but is not limited to, prediction patterns (e.g., CCLM, MMLM, CCCM), GLM pattern indexes, model parameters, or classification thresholds.
[0121] In the following figures, the co-located CTU refers to the CTU in the co-located picture, whose position corresponds to the position of the current coded CTU in the current picture. Let the positions of the upper left corner and the lower left corner of the co-located CTU be (xL, yT) and (xL, yB) respectively. Let the picture width be w. The x and y ranges of each dashed area are defined as follows: Region 1: xL - N ≤ x < w, yT - N ≤ y < yT Region 2: xL - N ≤ x < xL, yT ≤ y ≤ yB Region 3: 0 ≤ x < xL - N, yB - N < y ≤ yB N is any positive integer, N > 0.
[0122] In one embodiment, the temporal candidate can only refer to the CCM information in the co-located CTU, Region 1, Region 2 or Region 3, as Figure 14 shown. N is set to a predetermined value. For example, N is set to the minimum block size allowed in the specification, the minimum block width allowed, or the minimum block height allowed. The block can be a CU / PU / TU. Another example, N is set to 4. In one embodiment, the area where the temporal candidate can refer to the CCM information is the same as the area where the temporal motion vector can be referred to in the inter-frame mode. That is, the available area of the temporal candidate is the same as the available area of the temporal candidate in the inter-frame merge mode that contains the motion vector information.
[0123] Inheriting Non-Adjacent Spatial NeighbouringModel For another embodiment, the inherited model parameters can come from non-adjacent spatial neighbouring blocks. The models from the predetermined positions are added to the candidate list in a predetermined order. For example, the predetermined positions and the predetermined order are the same as those of the non-adjacent spatial neighbouring candidates in the inter-frame merge mode. For example, the positions and the order can be as Figure 10 shown. The positions in the numbered squares are the predetermined positions. The numbers within each square represent the predetermined order. The distance between each position is proportional to the width and height of the current coding / decoding block.
[0124] For yet another embodiment, the maximum number of models inherited from non-adjacent spatial neighbours can be less than the number of predetermined positions. For example, if the predetermined positions are as Figure 15A -B shown, which shows two styles yang ( Figure 15A style 1510 in Figure 15B and style 1520 in ). If the maximum number of models inherited from non-adjacent spatial neighbours is Only when this is the case will search mode 2 be used (i.e., add the model from the position in mode 2).
[0125] Inheriting model parameters from history table In one embodiment, the inherited model parameters can come from a cross-component model history table. Cross-component models in the history table can be added to the candidate list in a predetermined order. In one embodiment, the order of adding historical candidates can be from the beginning to the end of the table. In another embodiment, the order of adding historical candidates can be from a specific predetermined position to the end of the table. In another embodiment, the order of adding historical candidates can be from the end to the beginning of the table. In another embodiment, the order of adding historical candidates can be from a specific predetermined position to the beginning of the table. In yet another embodiment, the order of adding historical candidates can be staggered (e.g., the first added candidate comes from the beginning of the table, the second added candidate comes from the end of the table, and so on).
[0126] In one embodiment, a single cross-component model history table can be maintained to store previous cross-component models, and the cross-component model history table can be reset at the beginning of the current image, current tile, current tile, every M CTU rows, or every N CTU, where N and M can be any value greater than 0. In another embodiment, the cross-component model history table can be reset at the end of the current image, current tile, current tile, current CTU row, or current CTU.
[0127] Inherited from the fusion model Fusion mode refers to the mode of fusing two predictions to generate a final prediction. In chroma intra-frame fusion mode, a chroma intra-frame prediction generated without using a Cross-Component Prediction (CCP) codec (e.g., CCLM, MMLM, CCCM) is fused with another chroma intra-frame prediction generated using a CCP codec. For example, a non-CCLM codec intra-frame prediction and a CCLM codec intra-frame prediction are fused together to obtain the final intra-frame prediction.
[0128] In one embodiment, when inheriting cross-component model parameters from blocks / positions encoded / decoded by chroma intra-fraction mode, the model parameters used to obtain intra-prediction for CCP encoding / decoding are inherited and further optimized.
[0129] In one embodiment, in addition to inheriting and optimizing the CCP model parameters, the encoding / decoding mode for fusion weights and intra-frame prediction of non-CCP encoding / decoding is also inherited. That is, the chroma intra-frame fusion mode is inherited.
[0130] Candidate list construction In one embodiment, the candidate list is constructed by adding candidates in a predetermined order until a maximum number of candidates is reached. The added candidates may include, but are not limited to, all of the candidates described above. For example, the candidate list may include spatially adjacent candidates, temporally adjacent candidates, historical candidates, and non-immediately adjacent candidates. In another example, the candidate list may include the same candidates as in the previous example, but the candidates are added to the list in a different order.
[0131] In another embodiment, if all predetermined neighboring and historical candidates have been added but the maximum number of candidates has not been reached, some default candidates are added to the candidate list until the maximum number of candidates is reached.
[0132] In one sub-implementation, the preset candidates include, but are not limited to, the candidates described below. Final scaling parameters. From a set And offset parameter This is derived based on adjacent luminance and chrominance samples. For example, if the average values of adjacent luminance and chrominance samples are lumaAvg and chromaAvg, then... Export.
[0133] In another embodiment, the preset candidate can be an earlier candidate with optimization of the incremental scaling parameter. For example, if the scaling parameter of the earlier candidate is... The preset candidate scaling parameters are ,in It can come from the set { The preset candidate offset parameters will be obtained through... And the average values of the adjacent luminance and chrominance samples of the current block are derived.
[0134] Send the inheritance candidate index in the list An on / off flag can be sent to indicate whether the current block inherits cross-component model parameters from adjacent blocks. This flag can be indicated by CU / CB, PU, TU / TB, per color component, or per chroma color component. High-level syntax can be indicated in the SPS, Picture Parameter Set (PPS), Picture Header (PH), or Slice Header to indicate whether the proposed method is allowed for the current sequence, picture, or slice.
[0135] The maximum allowed number of candidates is sent to indicate the maximum size of the candidate list to be merged. This number can be indicated by CU / CB, PU, TU / TB, per color component, or per chroma color component. A high-level syntax can be specified in SPS, PPS, PH, or SH to indicate whether the proposed method is permitted for use in the current sequence, image, or slice. The maximum allowed number of candidates for the proposed method can be shared with the maximum allowed number of candidates for the inter-frame merging mode.
[0136] If the current block inherits cross-component model parameters from adjacent blocks, an inheritance candidate index is sent. This index can be sent (e.g., using truncated unary codes, Exp-Golomb codes, or fixed-length codes) and shared between the current Cb and Cr blocks. For example, the index can be sent per color component. For instance, one inheritance candidate index is sent for the Cb component, and another inheritance candidate index is sent for the Cr component. Alternatively, it can use chroma intra-frame prediction syntax (e.g., IntraPredModeC[xCb][yCb]) to store the inherited candidate indices.
[0137] The cross-component prediction using buffer-constrained inherited model parameters, as described above, can be implemented at the encoder or decoder level. For example, any of the proposed candidate derivation methods can be implemented in the intra / inter-frame encoding / decoding module of the decoder (e.g., Figure 1B Intra-frame prediction 150 / MC 152 in the codec), or implemented in the intra-frame / inter-frame codec module of the encoder (e.g., Figure 1A Intra-frame prediction 110 / inter-frame prediction 112). Any of the proposed candidate derivation methods can also be implemented as circuitry connected to the intra-frame / inter-frame encoding / decoding modules in the decoder or encoder. However, the decoder or encoder can also use additional processing units to implement the required cross-component prediction processing. While the intra-frame prediction unit (e.g., Figure 1A and Figure 1B Units 110 / 112 and 150 / 152 are shown as separate processing units, which may correspond to executable software or firmware code stored on a medium, such as a hard disk or flash memory, for a central processing unit (CPU) or a programmable device (e.g., a digital signal processor (DSP) or a field programmable gate array (FPGA)).
[0138] Figure 16A flowchart illustrating an exemplary video codec system incorporating inherited temporal cross-component model parameters and buffer constraints according to an embodiment of the present invention is provided. The steps shown in the flowchart can be implemented as program code executable on one or more processors at the encoder end (e.g., one or more CPUs). The steps shown in the flowchart can also be implemented in hardware, such as one or more electronic devices or processors arranged to perform the steps in the flowchart. According to this method, in step 1610, input data associated with the current block is received, the current block comprising a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or current block-related data to be decoded at the decoder end, and the current block is located in a non-intra-frame slice or picture. In step 1620, data or positions associated with reference data in previously encoded or decoded slices or pictures are determined, wherein the positions associated with the reference data are restricted within the co-location codec tree unit (CTU) of the current block and one or more adjacent regions adjacent to the co-location CTU. In step 1630, target cross-component model parameters associated with the target inherited prediction model are derived by inheriting the cross-component model parameters of the reference data. In step 1640, the second color patch is encoded or decoded using prediction data including cross-color prediction, which is generated by applying a target inherited prediction model with target cross-component model parameters to the reconstructed first color patch.
[0139] The flowchart shown is intended to illustrate examples of video encoding and decoding according to the present invention. Those skilled in the art can modify each step, rearrange the steps, split the steps, or combine the steps to practice the invention without departing from its spirit. In this disclosure, specific syntax and semantics are used to illustrate examples of embodiments of the invention. Those skilled in the art can practice the invention by substituting equivalent syntax and semantics without departing from its spirit.
[0140] The above description is intended to enable those skilled in the art to practice the invention in the context of specific applications and their requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the specific embodiments shown and described, but rather to be given the broadest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details have been set forth to provide a thorough understanding of the invention. However, the invention can be practiced by those skilled in the art.
[0141] As described above, embodiments of the present invention can be implemented through various hardware, software code, or a combination of both. For example, embodiments of the present invention may be program code integrated into one or more circuits in a video compression chip, or integrated into video compression software to perform the processes described herein. Embodiments of the present invention may also be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. The present invention may also relate to multiple functions performed by a computer processor, digital signal processor, microprocessor, or field-programmable gate array (FPGA). These processors can be configured to perform specific tasks according to the invention, defining the specific methods embodied in the invention by executing machine-readable software code or firmware code. The software code or firmware code can be developed in different programming languages and different formats or styles. The software code can also be compiled for different target platforms. However, different code formats, styles, and languages of the software code, as well as other means of configuring the code to perform tasks according to the invention, do not depart from the spirit and scope of the invention.
[0142] This invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The examples described are to be considered illustrative in all respects, not limiting. Therefore, the scope of the invention should be indicated by the appended claims rather than the foregoing description. All variations within the meaning and equivalence of the claims should be included within its scope.
Claims
1. A color image encoding / decoding method, using multiple encoding / decoding tools including one or more cross-component model correlation modes, the method comprising: Receive input data related to the current block, which includes a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data related to the current block to be decoded at the decoder end, and the current block is located in a non-intra-frame slice or image; Determine the data or location associated with the reference data, which is located in a previously encoded slice or picture, wherein the reference data associated with the location is restricted within the co-code tree unit (CTU) of the current block and one or more adjacent regions adjacent to the co-code tree unit; By inheriting multiple cross-component model parameters from the reference data, multiple target cross-component model parameters related to the target inherited prediction model are derived. as well as The second color patch is encoded or decoded using prediction data that includes cross-color prediction, which is generated by applying the target inherited prediction model with the multiple target cross-component model parameters to the reconstructed first color patch.
2. The color image encoding and decoding method as described in claim 1, characterized in that, The one or more adjacent regions are adjacent to the peer CTU, including the top N horizontal reference line regions located above the peer CTU, the left N vertical reference line regions adjacent to the left side of the peer CTU, the bottom N horizontal reference line regions located to the left side of the peer CTU and aligned with the bottom of the peer CTU, or a combination thereof, where N is a positive integer.
3. The color image encoding and decoding method as described in claim 2, characterized in that, N is set to a predetermined value.
4. The color image encoding and decoding method as described in claim 2, characterized in that, N is set to the minimum allowed block size, minimum block height, or minimum block width.
5. The color image encoding and decoding method as described in claim 4, characterized in that, The minimum allowed block size, minimum block height, or minimum block width is related to the codec unit (CU), prediction unit (PU), or transform unit (TU).
6. The color image encoding and decoding method as described in claim 1, characterized in that, The one or more adjacent regions adjacent to the same CTU are the same regions that allow time candidates to reference motion vectors in inter-frame modes.
7. The color image encoding and decoding method as described in claim 1, characterized in that, The reference data in the previously encoded slice or image is located based on the corresponding position of the current block in the previously encoded slice or image and the motion vector of the current block or adjacent blocks.
8. The color image encoding and decoding method as described in claim 7, characterized in that, When the current block is intra-frame encoded or decoded, the motion vector used to derive the multiple target cross-component model parameters of the current block is set to 0.
9. The color image encoding and decoding method as described in claim 7, characterized in that, The position of the reference data is determined based on the corresponding position of the current block, moved by the motion vector of the current block.
10. The color image encoding and decoding method as described in claim 7, characterized in that, The position of the reference data is determined based on the corresponding position of the current block, moved by the motion vector of the current block and one or more additional offsets.
11. The color image encoding and decoding method as described in claim 10, characterized in that, The one or more additional offsets depend on the width of the current block, the height of the current block, or both.
12. The color image encoding and decoding method as described in claim 11, characterized in that, The one or more additional offsets correspond to the horizontal offset of the width of the current block and the vertical offset of the height of the current block.
13. The color image encoding and decoding method as described in claim 11, characterized in that, The one or more additional offsets correspond to a horizontal offset of half the width of the current block and a vertical offset of half the height of the current block.
14. The color image encoding and decoding method as described in claim 1, characterized in that, The one or more cross-component model parameters are optimized and used as the parameters of the multiple target cross-component models.
15. The color image encoding and decoding method as described in claim 1, characterized in that, The first color block corresponds to the luminance block, while the second color block corresponds to the chroma block.
16. The color image encoding and decoding method as described in claim 1, characterized in that, Only the multiple cross-component model parameters of the target isotope image in the isotope CTU are allowed to be referenced by multiple time candidates.
17. The color image encoding and decoding method as described in claim 1, characterized in that, The multiple cross-component model parameters of the target co-located image are allowed to be referenced by multiple time candidates only if N and M are integers greater than 0 in the case of the co-located CTU, the N CTUs to the left of the co-located CTU, the M CTUs to the right of the co-located CTU, or a combination thereof.
18. The color image encoding and decoding method as described in claim 1, characterized in that, Only in the CTU row containing the co-otope CTU are the multiple cross-component model parameters of the target co-otope image allowed to be referenced by multiple temporal candidates.
19. The color image encoding and decoding method as described in claim 1, characterized in that, Only the multiple cross-component model parameters corresponding to the target isotope image of the current CTU row, N CTU rows above the current CTU row, M CTU rows below the current CTU row, or combinations thereof, are allowed to be referenced, and N and M are integers greater than 0.
20. A color image encoding / decoding apparatus using a plurality of encoding / decoding tools including one or more cross-component model correlation modes, the apparatus comprising one or more electronic circuits or processors configured to: Receive input data related to the current block, which includes a first color block and a second color block, wherein the input data includes pixel data to be encoded at the encoder end or data related to the current block to be decoded at the decoder end, and wherein the current block is located in a non-intra-frame slice or image; Determine the data or location associated with the reference data, which is located in a previously encoded slice or picture, wherein the reference data associated with the location is restricted within the co-code tree unit (CTU) of the current block and one or more adjacent regions adjacent to the co-code tree unit; By inheriting multiple cross-component model parameters from the reference data, multiple target cross-component model parameters related to the target inherited prediction model are derived. as well as The second color patch is encoded or decoded using prediction data that includes cross-color prediction, which is generated by applying the target inherited prediction model with the multiple target cross-component model parameters to the reconstructed first color patch.