Method for video decoding and related apparatus
By updating the parameters of the CCLM mode based on the reconstructed sample subset of the co-arranged brightness blocks in the video decoding method, the problem of insufficient adaptability of the chroma block prediction model in the prior art is solved, and the encoding efficiency of image and video compression is improved.
Patent Information
- Application Number
- CN202510279271.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-28
- Filing Date
- 2022-11-02
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively utilize the reconstructed luminance samples in image and video compression from prediction models adapted to adjust chroma blocks, resulting in low encoding efficiency.
In the video decoding method, the adjustment parameters are determined based on the reconstructed samples in the juxtaposition luminance block in the current picture, and the offset parameters and slope parameters in the cross component linear model (CCLM) mode are updated, thereby adaptively adjusting the prediction model of the chroma block.
The encoding efficiency of chroma blocks is improved, making the reconstructed luminance sample content more adaptable, and improving the overall performance of video compression.
Smart Images

Figure CN120050418A_ABST
Abstract
Description
[0001] Incorporated by Reference
[0002] This application claims the benefit of priority to U.S. Patent Application No. 17 / 976,556, filed on October 28, 2022, "ADAPTIVE PARAMETER SELECTION FOR CROSS-COMPONENT PREDICTION IN IMAGE AND VIDEO COMPRESSION," which claims the benefit of priority to U.S. Provisional Application No. 63 / 305,159, filed on January 31, 2022, "ADAPTIVE PARAMETER SELECTION FOR CROSS-COMPONENT PREDICTION IN IMAGE AND VIDEO COMPRESSION." The disclosure of the prior application is hereby incorporated by reference in its entirety.
[0003] This application files a divisional application for the Chinese patent application with application number 202280009886.9, application date November 2, 2022, and invention name “Adaptive Parameter Selection for Cross-Component Prediction in Image and Video Compression”. Technical Field
[0004] The present disclosure describes techniques generally related to coding and decoding, and more particularly, methods and related apparatus for video decoding. Background Art
[0005] The background art provided herein is not necessarily prior art to the present disclosure.
[0006] An uncompressed digital image and / or video may include a series of pictures, each picture having, for example, spatial dimensions of 1920×1080 luminance samples and associated chrominance samples. This series of pictures may have a fixed or variable picture rate (also informally referred to as a frame rate) of, for example, 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bit rate requirements. For example, 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) with 8 bits per sample requires a bandwidth of nearly 1.5 Gbit / s. One hour of such video requires over 600 Gigabytes of storage space.
[0007] One purpose of image and / or video encoding and decoding can be to reduce redundancy in input image and / or video signals by compression. Compression can help reduce the bandwidth and / or storage space requirements mentioned above, in some cases by two orders of magnitude or more. Although the description herein uses video encoding / decoding as an illustrative example, the same technology can be applied to image encoding / decoding in a similar manner without departing from the spirit of the present disclosure. Both lossless compression and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique that can reconstruct an exact copy of the original signal based on the compressed original signal. When lossy compression is used, the reconstructed signal may be different from the original signal, but the distortion between the original signal and the reconstructed signal is small enough to enable the reconstructed signal to be used for the intended application. In the case of video, lossy compression is widely used. The amount of distortion tolerated depends on the application; for example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that a higher allowable / tolerable distortion can produce a higher compression ratio.
[0008] Video encoders and decoders may utilize techniques from several broad categories including, for example, motion compensation, transform processing, quantization, and entropy coding.
[0009] Video codec techniques may include techniques known as intra-frame coding. In intra-frame coding, sample values are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. When all sample blocks are encoded in intra-frame mode, the picture may be an intra-frame picture. Intra-frame pictures and their derivatives (e.g., independent decoder refresh pictures) may be used to reset decoder states, and may therefore be used as the first picture in a coded video bitstream and video session, or as a still image. Samples of intra-frame blocks may be subjected to transformation, and transform coefficients may be quantized before entropy coding. Intra-frame prediction may be a technique for minimizing sample values in a pre-transform domain. In some cases, the smaller the DC value after transformation and the smaller the AC coefficient, the fewer bits are required to represent the block after entropy coding at a given quantization step size.
[0010] For example, conventional intra-frame coding used in MPEG-2 generation coding techniques does not use intra-frame prediction. However, some newer video compression techniques include techniques that attempt to perform predictions based on, for example, surrounding sample data and / or metadata obtained during encoding and / or decoding of a block of data. Such techniques are hereinafter referred to as "intra-frame prediction" techniques. Note that, in at least some cases, intra-frame prediction uses reference data only from the current picture under reconstruction, and does not use reference data from reference pictures.
[0011] There can be many different forms of intra-frame prediction. When more than one such technique can be used in a given video coding technique, the specific technique used can be encoded as a specific intra-frame prediction mode that uses the specific technique. In some cases, an intra-frame prediction mode can have sub-modes and / or parameters, where the sub-modes and / or parameters can be encoded separately or included in a mode codeword that defines the prediction mode being used. Which codeword is used for a given mode, sub-mode and / or parameter combination can affect the coding efficiency gain through intra-frame prediction, and therefore the entropy coding technique used to convert the codeword into a bitstream can also affect the coding efficiency gain through intra-frame prediction.
[0012] Some modes of intra prediction were introduced with H.264, refined in H.265, and further refined in newer coding techniques such as the joint exploration model (JEM), versatile video coding (VVC), and the benchmark set (BMS). Using the values of neighboring samples of the already available samples, a predictor block can be formed. Sample values of neighboring samples are copied into the predictor block according to the direction. A reference to the used direction can be encoded in the bitstream or can be predicted itself.
[0013] Reference Figure 1A , depicted at the bottom right is a subset of nine predictor directions known from the 33 possible predictor directions defined in H.265 (corresponding to the 33 angular modes for the 35 intra modes). The point where the arrows intersect (101) represents the sample being predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at an angle of 45 degrees to the horizontal line to the upper right. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at an angle of 22.5 degrees to the horizontal line to the lower left of sample (101).
[0014] Still refer to Figure 1A, depicted at the top left is a square block (104) of 4×4 samples (indicated by bold dashed lines). The square block (104) includes 16 samples, each of which is labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample in the block (104) in both the Y dimension and the X dimension. Since the size of the block is 4×4 samples, S44 is at the bottom right. Also shown are reference samples that follow a similar numbering scheme. The reference samples are labeled with R, their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are adjacent to the block under reconstruction; therefore, there is no need to use negative values.
[0015] Intra-picture prediction can work by copying reference sample values from neighboring samples indicated by a signaled prediction direction. For example, assume that the coded video bitstream includes signaling that indicates a prediction direction consistent with arrow (102) for this block - that is, the sample is predicted from the sample to the upper right at a 45 degree angle from the horizontal. In this case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Sample S44 is then predicted from reference sample R08.
[0016] In some cases, the values of multiple reference samples may be combined, for example, by interpolation, in order to compute a reference sample; in particular, when the directions are not evenly divisible at 45 degrees.
[0017] As video coding technology has developed, the number of possible directions has also increased. In H.264 (2003), nine different directions could be represented. In H.265 (2013), this increased to 33. Currently, JEM / VVC / BMS can support up to 65 directions. Experiments have been conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent these possible directions with a small number of bits, at the expense of fewer possible directions. In addition, the direction itself can sometimes be predicted based on the neighboring directions used in adjacent decoded blocks.
[0018] Figure 1B A schematic diagram (110) is shown depicting 65 intra prediction directions according to JEM to illustrate that the number of prediction directions increases over time.
[0019] The mapping of the intra-frame prediction direction bit representing the direction in the coded video bit stream can be different according to different video coding techniques. The scope of such mapping can be, for example, from simple direct mapping to codewords, complex adaptive schemes and similar techniques related to the most probable mode. However, in most cases, there may be some directions that are less likely to occur in the video content than some other directions statistically. Since the target of video compression is to reduce redundancy, in the video coding techniques that work well, those less likely directions will be represented by a larger number of bits compared to more likely directions.
[0020] Image and / or video encoding and decoding may be performed using inter-picture prediction with motion compensation. Motion compensation may be a lossy compression technique and may involve a technique in which a block of sample data from a previously reconstructed picture or portion thereof (reference picture) is used to predict a reconstructed picture or picture portion after being spatially shifted in a direction indicated by a motion vector (hereinafter MV). In some cases, the reference picture may be the same as the picture under current reconstruction. The MV may have two dimensions, X and Y, or three dimensions, the third dimension being an indication of the reference picture in use (indirectly, the third dimension may be a temporal dimension).
[0021] In some video compression techniques, an MV applicable to a particular region of sample data can be predicted based on other MVs, for example, based on an MV related to another region of sample data that is spatially adjacent to the region under reconstruction and precedes the MV in decoding order. The above prediction can significantly reduce the amount of data required to encode the MV, thereby eliminating redundancy and increasing compression. MV prediction can work effectively, for example, because when encoding an input video signal derived from a camera (called natural video), there is a statistical probability that a larger region than the region to which a single MV is applicable moves in a similar direction, and therefore in some cases similar motion vectors derived from MVs of adjacent regions can be used to predict the larger region. This makes the MV obtained for a given region similar or identical to the MV predicted from the surrounding MVs, and in turn, the MV can be represented after entropy coding with a smaller number of bits than would be used if the MV was encoded directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, MV prediction itself can be lossy, for example due to rounding errors when calculating a predictor based on several surrounding MVs.
[0022] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T H.265 Recommendation, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, a technique referred to as "spatial merging" hereinafter is described with reference to FIG. 2 .
[0023] 2, the current block (201) includes samples obtained by the encoder during the motion search process, which can be predicted from a previous block of the same size that has been spatially shifted. Instead of encoding the MV directly, the MV associated with any of the five surrounding samples represented by A0, A1 and B0, B1, B2 (corresponding to 202 to 206, respectively) can be used to derive the MV from metadata associated with one or more reference pictures, such as from the most recent (in decoding order) reference picture. In H.265, MV prediction can use a predictor from the same reference picture being used by a neighboring block.
[0024] The cross-component linear model prediction (CCLM) mode is a cross-component prediction method. In CCLM, a linear model can be used to predict chrominance samples based on reconstructed luminance samples. The linear model can be established by adjacent reconstructed samples of the current block (e.g., the chrominance block to be encoded). Usually, for example, N reconstructed adjacent chrominance samples and N corresponding luminance samples are used to determine the parameters a and b of the linear model through classical linear regression theory. However, in some cases, the reference samples used to generate the linear model parameters a and b are noisy and / or not representative of the content in the actual prediction block, thereby reducing the coding efficiency. Therefore, it is necessary to develop a more content-adaptive linear model for chrominance sample prediction. Summary of the invention
[0025] Various aspects of the present disclosure provide methods and devices for video encoding / decoding. In some examples, the method for video decoding includes: decoding prediction information of a chroma block to be reconstructed in a current picture, the prediction information indicating that a cross-component linear model (CCLM) mode is applied to the chroma block; for a first area in the chroma block, determining a first adjustment parameter for adjusting an offset parameter in the CCLM mode based on a first subset of reconstructed samples in a collocated luminance block in the current picture, the first subset of reconstructed samples not including one or more samples in the collocated luminance block; updating the offset parameter based on at least the first adjustment parameter; determining a second adjustment value for modifying a slope parameter in the CCLM mode based on a second subset of reconstructed samples in the luminance block; updating the slope parameter based on at least the second adjustment value; and reconstructing the first area in the chroma block using the CCLM mode based on at least the updated offset parameter and the updated slope parameter.
[0026] Aspects of the present disclosure also provide an apparatus for video decoding, comprising a processing circuit system configured to perform the above-mentioned method for video decoding.
[0027] Aspects of the present disclosure also provide a device for video decoding, comprising: a decoding module configured to decode prediction information of a chroma block to be reconstructed in a current picture, the prediction information indicating that a cross-component linear model (CCLM) mode is applied to the chroma block; a first determination module configured to determine, for a first area in the chroma block, a first adjustment value for modifying an offset parameter in the CCLM mode based on a first subset of reconstructed samples in a luminance block juxtaposed with the chroma block in the current picture, the first subset of reconstructed samples not including one or more samples in the luminance block; a first update module configured to update the offset parameter based at least on the first adjustment value; a second determination module configured to determine a second adjustment value for modifying a slope parameter in the CCLM mode based on a second subset of reconstructed samples in the luminance block; a second update module configured to update the slope parameter based at least on the second adjustment value; and a reconstruction module configured to reconstruct the first area in the chroma block using the CCLM mode based at least on the updated offset parameter and the updated slope parameter.
[0028] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform the above-described method for video decoding.
[0029] According to the method and related devices of the present disclosure, the prediction information indicates that a cross-component linear model (CCLM) mode is applied to the chroma block. For a first area in the chroma block, a first adjustment parameter is determined based on a first subset of reconstructed samples in a juxtaposed luminance block in a current picture, an offset parameter in the CCLM model is updated based on the first adjustment value, a second adjustment value is determined based on a second subset of the reconstructed samples in the luminance block, a slope parameter in the CCLM model is updated based on the second adjustment value, and the first area in the chroma block is reconstructed using the updated CCLM model. Therefore, the present invention adjusts the parameters of the CCLM model used to reconstruct the chroma block according to the content of the reconstructed luminance samples, so that the parameters of the CCLM model used to reconstruct the chroma block are more adaptive to the content of the reconstructed luminance samples, thereby improving the coding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0031] Figure 1A is a schematic illustration of an exemplary subset of intra prediction modes;
[0032] Figure 1B is a diagram of exemplary intra prediction directions;
[0033] FIG2 shows an example of a current block ( 201 ) and surrounding samples;
[0034] Figure 3 is a schematic illustration of an exemplary block diagram of a communication system (300);
[0035] Figure 4 is a schematic illustration of an exemplary block diagram of a communication system (400);
[0036] Figure 5 is a schematic illustration of an exemplary block diagram of a decoder;
[0037] Figure 6 is a schematic illustration of an exemplary block diagram of an encoder;
[0038] Figure 7 A block diagram of an exemplary encoder is shown;
[0039] Figure 8 A block diagram of an exemplary decoder is shown;
[0040] Fig. 9 An exemplary reference sample of the current block is shown;
[0041] Fig.10 An example of reconstructed adjacent luma samples and reconstructed adjacent chroma samples used in a cross-component linear model (CCLM) mode is shown;
[0042] Fig.11 shows examples of reconstructed adjacent luma samples and reconstructed adjacent chroma samples used in CCLM mode;
[0043] FIG. 12A to FIG. 12B An example of parameter adjustment used in CCLM mode is shown;
[0044] Fig.13 An example of sample positions of selected luma samples in a collocated luma block for determining an adjustment parameter is shown;
[0045] Fig.14 Exemplary chrominance blocks and collocated luminance blocks are shown;
[0046] Fig.15 Exemplary chrominance blocks and collocated luminance blocks are shown;
[0047] Fig.16 shows a flowchart outlining an encoding process according to some embodiments of the present disclosure;
[0048] Fig.17shows a flowchart outlining a decoding process according to some embodiments of the present disclosure;
[0049] Fig.18 is a schematic illustration of a computer system according to an embodiment. DETAILED DESCRIPTION
[0050] Figure 3 An exemplary block diagram of a communication system (300) is shown. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3 In the example of , a first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can encode video data (e.g., a video picture stream captured by the terminal device (310)) for transmission to another terminal device (320) via a network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) can receive the encoded video data from the network (350), decode the encoded video data to recover the video picture, and display the video picture based on the recovered video data. Unidirectional data transmission may be common in media service applications, etc.
[0051] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, for example, during a video conference. For the bidirectional transmission of data, in the example, each of the terminal devices (330) and (340) can encode video data (e.g., a video picture stream captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) can also receive the encoded video data transmitted by the other of the terminal devices (330) and (340), and can decode the encoded video data to restore the video picture, and can display the video picture at an accessible display device based on the restored video data.
[0052] exist Figure 3In the example of , terminal devices (310), (320), (330) and (340) are respectively shown as servers, personal computers and smart phones, but the principles of the present disclosure may not be so limited. Implementations of the present disclosure are applicable to laptop computers, tablet computers, media players and / or dedicated video conferencing equipment. Network (350) represents any number of networks that transmit encoded video data between terminal devices (310), (320), (330) and (340), including, for example, wired (wired) and / or wireless communication networks. Communication network (350) can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purposes of this discussion, unless otherwise described herein below, the architecture and topology of network (350) may be irrelevant to the operation of the present disclosure.
[0053] As examples of applications of the disclosed subject matter, Figure 4 A video encoder and a video decoder in a streaming environment are shown. The disclosed subject matter may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., etc.
[0054] The streaming system may include a capture subsystem (413) that may include a video source (401), such as a digital camera, that creates, for example, an uncompressed video picture stream (402). In an example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402) is depicted as a thick line to emphasize the high amount of data when compared to the encoded video data (404) (or encoded video bitstream), which may be processed by an electronic device (420) coupled to the video source (401) and including a video encoder (403). The video encoder (403) may include hardware, software, or a combination thereof to implement or embody aspects of the disclosed subject matter as described in more detail below. The encoded video data (404) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (402), and the encoded video data (404) (or encoded video bitstream) can be stored on the streaming server (405) for future use. One or more streaming client subsystems, such as Figure 4The client subsystems (406) and (408) in the streaming server (405) can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and creates an outgoing video picture stream (411) that can be presented on a display (412) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), (407) and (409) (e.g., a video bitstream) can be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In the example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter can be used in the context of VVC.
[0055] Note that the electronic devices (420) and (430) may include other components (not shown). For example, the electronic device (420) may include a video decoder (not shown), and the electronic device (430) may also include a video encoder (not shown).
[0056] Figure 5 An exemplary block diagram of a video decoder (510) is shown. The video decoder (510) may be included in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., receiving circuitry). The video decoder (510) may replace Figure 4 The video decoder (410) in the example is used.
[0057] The receiver (531) can receive one or more encoded video sequences to be decoded by the video decoder (510). In some embodiments, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence can be received from a channel (501), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (531) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams, which can be forwarded to their respective use entities (not depicted). The receiver (531) can separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other applications, the buffer memory (515) can be external to the video decoder (510) (not depicted). In still other applications, there may be a buffer memory (not depicted) external to the video decoder (510) to, for example, prevent network jitter, and in addition there may be another buffer memory (515) internal to the video decoder (510) to, for example, handle playout timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (515) may not be needed, or the buffer memory (515) may be small. For use over a best effort packet network such as the Internet, a buffer memory (515) may be needed, which may be relatively large and may advantageously have an adaptive size, and may be implemented at least in part in an operating system or similar element (not depicted) external to the video decoder (510).
[0058] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. The categories of these symbols include: information for managing the operation of the video decoder (510), and possibly information for controlling a rendering device such as a rendering device (512) (e.g., a display screen), which is not part of the electronic device (530) but can be coupled to the electronic device (530), such as Figure 5. The control information of the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (520) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may comply with a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) may extract a subgroup parameter set for at least one subgroup of the pixel subgroups in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include: a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.
[0059] The parser (520) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).
[0060] The reconstruction of the symbol (521) may involve a number of different units depending on the type of coded video picture or portion thereof (e.g., inter- and intra-pictures, inter- and intra-blocks) and other factors. Which units are involved and how they are involved may be controlled by subgroup control information parsed from the coded video sequence by the parser (520). For clarity, such subgroup control information flow between the parser (520) and the following multiple units is not depicted.
[0061] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0062] The first unit is a sealer / inverse transform unit (551). The sealer / inverse transform unit (551) receives quantized transform coefficients and control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) from the parser (520) as (one or more) symbols (521). The sealer / inverse transform unit (551) can output blocks including sample values, which can be input into the aggregator (555).
[0063] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses surrounding reconstructed information obtained from a current picture buffer (558) to generate a block of the same size and shape as the block in reconstruction. For example, the current picture buffer (558) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (555) adds the prediction information already generated by the intra-picture prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551) on a per-sample basis.
[0064] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to an inter-frame coded and possibly motion compensated block. In such a case, the motion compensated prediction unit (553) may access the reference picture memory (557) to obtain samples for prediction. After motion compensation of the obtained samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (in this case referred to as residual samples or residual signal), thereby generating output sample information. The address within the reference picture memory (557) from which the motion compensated prediction unit (553) obtains the predicted samples may be controlled by a motion vector, which may be obtained by the motion compensated prediction unit (553) in the form of a symbol (521), which may have, for example, an X component, a Y component and a reference picture component. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (557) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0065] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream), which may be obtained by the loop filter unit (556) as symbols (521) from the parser (520). The video compression may also be responsive to meta-information obtained during decoding of a previous portion (in decoding order) of a coded picture or coded video sequence, as well as to previously reconstructed and loop filtered sample values.
[0066] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and stored in a reference picture memory (557) for use in future inter-picture prediction.
[0067] Once fully reconstructed, certain coded pictures can be used as reference pictures for future predictions. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557) and a new current picture buffer can be reallocated before starting to reconstruct a subsequent coded picture.
[0068] The video decoder (510) may perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265 Recommendation. The coded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the coded video sequence follows both the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as tools available only under the profile. For compliance, it is also required that the complexity of the coded video sequence is within the limits defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in, for example, millions of samples per second), the maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the coded video sequence.
[0069] In some embodiments, the receiver (531) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence (one or more). The additional data can be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0070] Figure 6 An exemplary block diagram of a video encoder (603) is shown. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit system). The video encoder (603) may replace Figure 4 The video encoder (403) in the example is used.
[0071] The video encoder (603) can be used to obtain the video source (601) (not Figure 6 In an example of an electronic device (620) receiving video samples, a video source (601) can capture (one or more) video images to be encoded by a video encoder (603). In another example, the video source (601) is a part of the electronic device (620).
[0072] The video source (601) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (603), the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (601) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The picture itself may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be easily understood by those skilled in the art. The following description focuses on samples.
[0073] According to an embodiment, the video encoder (603) can encode and compress the pictures of the source video sequence into an encoded video sequence (643) in real time or under any other time constraints required. Implementing the appropriate encoding speed is a function of the controller (650). In some embodiments, the controller (650) controls other functional units as described below and is functionally coupled to the other functional units. The coupling is not depicted for simplicity. The parameters set by the controller (650) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions that belong to the video encoder (603) optimized for a specific system design.
[0074] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As an extremely simplified description, in an example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (one or more) reference pictures) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to create sample data in a manner similar to the way the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (634). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (634) are also bit-accurate between the local encoder and the remote encoder. In other words, the prediction part of the encoder "regards" the sample values that are exactly the same as the sample values that the decoder "sees" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchronization (and the offset that occurs when synchronization cannot be maintained, for example due to channel errors) is also used in some related technologies.
[0075] The operation of the "local" decoder (633) can be combined with the "remote" decoder such as has been described above. Figure 5 The operation of the video decoder (510) described in detail is the same. However, in addition, briefly refer to Figure 5 , since the symbols are available and the encoding of the symbols into a coded video sequence by the entropy encoder (645) and the decoding of the symbols by the parser (520) can be lossless, the entropy decoding portion of the video decoder (510) including the buffer memory (515) and the parser (520) may not be fully implemented in the local decoder (633).
[0076] In some embodiments, the decoder technology except the parsing / entropy decoding present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is opposite to the decoder technology described comprehensively. In some aspects, a more detailed description is provided below.
[0077] In some examples, during operation, the source encoder (630) may perform motion compensated predictive coding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence designated as “reference pictures.” In this manner, the encoding engine (632) encodes the differences between pixel blocks of an input picture and pixel blocks of reference picture(s) that may be selected as prediction reference(s) for the input picture.
[0078] Based on the symbols created by the source encoder (630), the local video decoder (633) can decode the encoded video data of the picture that can be designated as the reference picture. The operation of the encoding engine (632) can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 6 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture that will be obtained by the remote video decoder.
[0079] The predictor (635) may perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) may search the reference picture memory (634) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may be used as suitable prediction references for the new picture. The predictor (635) may operate on a sample-block-by-pixel-block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (635), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (634).
[0080] The controller (650) may manage encoding operations of the source encoder (630), including, for example, setting parameters and sub-group parameters for encoding video data.
[0081] The outputs of all the above-mentioned functional units may be subjected to entropy coding in an entropy encoder (645). The entropy encoder (645) transforms the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0082] The transmitter (640) can buffer the encoded video sequence(s) created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0083] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) may assign to each encoded picture a certain encoded picture type that may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned one of the following picture types:
[0084] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and features.
[0085] A predictive picture (P picture) may be a picture that can be encoded and decoded using inter prediction or intra prediction to predict sample values of each block using at most one motion vector and a reference index.
[0086] Bi-directional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using inter-prediction or intra-prediction that uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.
[0087] The source picture may typically be spatially subdivided into a plurality of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or may be predictively coded (spatial prediction or intra-prediction) with reference to coded blocks of the same picture. Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0088] The video encoder (603) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In its operation, the video encoder (603) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0089] In some embodiments, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0090] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified as intra-prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture being encoded / decoded (which is referred to as the current picture) is divided into blocks. In the case where a block in the current picture is similar to a reference block in a reference picture that was previously encoded and buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.
[0091] In some embodiments, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but can be in the past and in the future respectively in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0092] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0093] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, a picture in a video picture sequence is divided into coding tree units (CTUs) for compression, and the CTUs in the picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. In general, a CTU includes three coding tree blocks (CTBs), each of which is a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) by a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. Depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In some embodiments, the prediction operation in the codec (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0094] Figure 7 An exemplary diagram of a video encoder (703) is shown. The video encoder (703) is configured to receive a processed block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures, and encode the processed block into an encoded picture that is part of an encoded video sequence. In an example, the video encoder (703) replaces Figure 4 The video encoder (403) in the example is used.
[0095] In the HEVC example, the video encoder (703) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (703) uses, for example, rate-distortion optimization to determine whether to best encode the processing block using intra mode, inter mode, or bidirectional prediction mode. In the case where the processing block is to be encoded in intra mode, the video encoder (703) may encode the processing block into a coded picture using intra prediction techniques; and in the case where the processing block is to be encoded in inter mode or bidirectional prediction mode, the video encoder (703) may encode the processing block into a coded picture using inter prediction or bidirectional prediction techniques, respectively. In some video coding techniques, the merge mode may be an inter-picture prediction submode, where the motion vector is derived from the predictor without the aid of an encoded motion vector component external to one or more motion vector predictors. In some other video coding techniques, there may be a motion vector component applicable to the object block. In the example, the video encoder (703) includes other components, such as a mode decision module (not shown) that determines the prediction mode of the processing block.
[0096] exist Figure 7 In the example of FIG. 7 , the video encoder ( 703 ) includes Figure 7 An inter-frame encoder (730), an intra-frame encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), an overall controller (721), and an entropy encoder (725) are shown coupled together.
[0097] The inter-frame encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., motion vectors, merge mode information, descriptions of redundant information according to inter-frame coding techniques), and calculate an inter-frame prediction result (e.g., a prediction block) based on the inter-frame prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that is decoded based on the encoded video information.
[0098] The intra encoder (722) is configured to: receive samples of a current block (e.g., a processing block); compare the block with an already encoded block in the same picture in some cases; generate quantization coefficients after transformation; and in some cases also generate intra prediction information (e.g., generate intra prediction direction information according to one or more intra coding techniques). In an example, the intra encoder (722) also calculates an intra prediction result (e.g., a prediction block) based on the intra prediction information and a reference block in the same picture.
[0099] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In an example, the overall controller (721) determines a mode of a block and provides a control signal to a switch (726) based on the mode. For example, when the mode is an intra-frame mode, the overall controller (721) controls the switch (726) to select an intra-frame mode result for use by a residual calculator (723), and controls the entropy encoder (725) to select intra-frame prediction information and include the intra-frame prediction information in a bitstream; and when the mode is an inter-frame mode, the overall controller (721) controls the switch (726) to select an inter-frame prediction result for use by a residual calculator (723), and controls the entropy encoder (725) to select inter-frame prediction information and include the inter-frame prediction information in a bitstream.
[0100] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to operate based on the residual data to encode the residual data to generate a transform coefficient. In an example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain and generate a transform coefficient. Then, the transform coefficient is subjected to quantization to obtain a quantized transform coefficient. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. In some examples, the decoded blocks are appropriately processed to generate decoded pictures, and these decoded pictures can be buffered in memory circuits (not shown) and used as reference pictures.
[0101] The entropy encoder (725) is configured to format the bitstream to include the coded blocks. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard, such as the HEVC standard. In an example, the entropy encoder (725) is configured to include overall control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, when the block is encoded in the merge sub-mode of the inter-frame mode or the bidirectional prediction mode, there is no residual information.
[0102] Figure 8An exemplary diagram of a video decoder (810) is shown. The video decoder (810) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (810) replaces Figure 4 The video decoder (410) in the example uses.
[0103] exist Figure 8 In the example of FIG. 8 , the video decoder ( 810 ) includes Figure 8 An entropy decoder (871), an inter-frame decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-frame decoder (872) are shown coupled together.
[0104] The entropy decoder (871) may be configured to reconstruct certain symbols according to the coded picture, which represent the syntax elements constituting the coded picture. Such symbols may include, for example, a mode for encoding a block (e.g., an intra-frame mode, an inter-frame mode, a bidirectional prediction mode, a merged sub-mode of the latter two, or another sub-mode) and prediction information (e.g., intra-frame prediction information or inter-frame prediction information) of certain samples or metadata that may be identified for prediction by an intra-frame decoder (872) or an inter-frame decoder (880). The symbol may also include residual information in the form of, for example, quantized transform coefficients, etc. In an example, when the prediction mode is an inter-frame mode or a bidirectional prediction mode, the inter-frame prediction information is provided to an inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, the intra-frame prediction information is provided to an intra-frame decoder (872). The residual information may be subjected to inverse quantization and provided to a residual decoder (873).
[0105] The inter-frame decoder (880) is configured to receive the inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information.
[0106] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.
[0107] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to transform the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (to include quantizer parameters (Quantizer Parameter, QP)), and the information may be provided by the entropy decoder (871) (data path is not depicted because this may only be a small amount of control information).
[0108] The reconstruction module (874) is configured to combine the residual information output by the residual decoder (873) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module as appropriate) in the spatial domain to form a reconstructed block, which can be part of a reconstructed picture, which can be part of a reconstructed video. Note that other suitable operations such as deblocking operations can be performed to improve visual quality.
[0109] Note that the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using any suitable technology. In some embodiments, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603) and (703) and the video decoders (410), (510) and (810) may be implemented using one or more processors that execute software instructions.
[0110] In intra prediction or intra prediction mode, sample values of a coding block may be predicted based on already reconstructed neighboring samples or reconstructed neighboring samples, referred to as reference samples.
[0111] An example of intra prediction is directional intra prediction. In directional intra prediction, a sample in a current block (e.g., a current sample) may be predicted using a reference sample (e.g., a prediction sample) or an interpolated reference sample. For example, the connecting line between the current sample and the prediction sample forms a given angular direction such as used in angular mode.
[0112] In another example of intra prediction, a planar mode based on sample interpolation is used. In planar mode, one or more key positions in or around the current block can be predicted using adjacent reference samples. Other positions in the current block can be predicted as a linear combination of samples at one or more key positions and reference samples. Weights (e.g., combination weights) can be determined based on the position of the current sample in the current block.
[0113] An example of calculating the planar mode, such as in VVC, is as follows.
[0114] predV[x][y]=((H-1-y)×p[x][-1]+(y+1)×p[-1][H])<<Log2(W) Equation 1
[0115] predH[x][y]=((W-1-x)×p[-1][y]+(x+1)×p[W][-1])<<Log2(H) Equation 2
[0116] pred[x][y]=(predV[x][y]+predH[x][y]+W×H)>>(Log2(W)+Log2(H)+1) Equation 3
[0117] Reference Fig. 9 , the current block (900) includes samples at positions (0, 0) to (H-1, W-1) in the current block (900). W and H are the width and height of the current block (900), respectively. As shown in Equations 1 to 3, the predicted sample value pred[x][y] of the current sample at position (x, y) (x=0, 1, ..., or W-1, and y=0, 1, ..., or H-1) in the current block (900) can be obtained as a weighted average of reference sample values (e.g., p[-1][y], p[x][-1], p[-1][H], and p[W][-1]) of reference samples at positions (-1, y), (x, -1), (-1, H), and (W, -1), respectively. The reference samples may include a reference sample at a position (-1, y) in the same row as the current sample, a reference sample at a position (x, -1) in the same column as the current sample, a reference sample at a lower left position (-1, H) relative to the current block, and a reference sample at an upper right position (W, -1) relative to the current block.
[0118] As described above, the reference sample values of the reference samples may include p[-1][y] of a reference sample located at a position (-1, y), p[x][-1] of a reference sample located at a position (x, -1), p[-1][H] of a reference sample located at a lower left position (-1, H), and p[W][-1] of a reference sample located at an upper right position (W, -1). In the example shown in Equation 1, a vertical predictor predV[x][y] is determined based on the reference samples located at the positions (x, -1) and (-1, H). In the example shown in Equation 2, a horizontal predictor predH[x][y] is determined based on the reference samples located at the positions (-1, y) and (W, -1). In Equation 3, a predicted sample value pred[x][y] is determined based on an average (e.g., a weighted average) of the horizontal predictor predH[x][y] and the vertical predictor predV[x][y].
[0119] The cross-component linear model prediction (CCLM) mode is a cross-component prediction method. In CCLM, a linear model can be used to predict chrominance samples based on reconstructed luminance samples. The linear model can be established by adjacent reconstructed samples of the current block (e.g., the chrominance block to be encoded). In some embodiments, the prediction performance is high when the luminance channel and the chrominance channel are highly linearly correlated.
[0120] In some embodiments, a chrominance block (eg, a current Cb block or a current Cr block) is predicted based on a collocated luminance block. The predicted block Pred_C of the chrominance block may be derived as follows:
[0121] Pred_C(x, y) = a × Rec_L′(x, y) + b Equation 4
[0122] The samples Pred_C(x, y) in the prediction block Pred_C of the chrominance block may be determined based on the samples in the reconstructed collocated luminance block.
[0123] Pred_C(x,y) represents the predicted chroma sample at the sampling position (x,y) of the chroma block (e.g., in the chroma channel of the current picture or the chroma picture). Rec_L'(x,y) may be determined based on the reconstructed samples in the collocated luma block (e.g., in the luma channel of the current picture or the luma picture) that have been reconstructed. Rec_L'(x,y) may represent the reconstructed luma sample in the collocated luma block or the downsampled luma sample of the collocated luma block.
[0124] In an example, for example, when the color format is 4:4:4 and the size of the chroma block is the same as the size of the collocated luma block, Rec_L' is the collocated luma block that has been reconstructed. Therefore, Rec_L'(x,y) can represent the reconstructed sample in the collocated luma block, where the reconstructed sample corresponds to the sample position (x,y) of the chroma block.
[0125] In the example, Rec_L' is different from the collocated luma block that has been reconstructed. The current chroma block is collocated with the collocated luma block, and the luma channel and the chroma channel have different resolutions. In the example, when the color format is 4:2:0, the resolution of the luma block is twice the resolution of the chroma block in both the vertical and horizontal directions. Therefore, when the color format is 4:2:0, Rec_L' can be a downsampled block of the corresponding luma block to match the size of the chroma block during the derivation of the linear model. In some embodiments, when the color format is not 4:4:4, the collocated luma block is downsampled.
[0126] The parameters (e.g., model parameters) a and b in Equation 4 may represent the slope and offset in the linear model shown in Equation 4, and may be referred to as slope parameters and offset parameters, respectively. The parameters a and b in Equation 4 may be derived based on reconstructed neighboring samples (e.g., chroma samples and luma samples) around (i) the current chroma block in the chroma channel and (ii) the collocated luma block in the luma channel. The parameters a and b may be determined using any suitable method.
[0127] In the example, the parameters a and b are determined using classical linear regression theory. The following parameters a and b can be obtained by applying the minimum linear least squares solution between (i) the reconstructed adjacent luma samples or the downsampled samples of the reconstructed adjacent luma samples and (ii) the reconstructed adjacent chroma samples:
[0128]
[0129] In equations 5 to 6, N reconstructed adjacent chroma samples Rec_C(i) and N corresponding luma samples Rec_L′(i) are used. In an example, such as when the color format is 4:4:4, the N corresponding luma samples Rec_L′(i) include N reconstructed adjacent luma samples of the luma block. In an example, such as when the color format is 4:2:0, the N corresponding luma samples Rec_L′(i) include N downsampled samples of the reconstructed adjacent luma samples of the luma block. i can be an integer from 1 to N.
[0130] The N reconstructed neighboring chroma samples used to determine parameters a and b may include any suitable already reconstructed neighboring samples of the chroma block. The reconstructed neighboring luma samples used to determine parameters a and b may include any suitable already reconstructed neighboring samples of the collocated luma block.
[0131] The prediction process of the CCLM mode may include: (1) downsampling the collocated luma block and the reconstructed adjacent luma samples of the collocated luma block to obtain Rec_L' and the downsampled adjacent luma samples, and thus matching the size of the corresponding chroma block; (2) deriving parameters a and b based on the downsampled adjacent luma samples and the reconstructed adjacent chroma samples, for example, using equations 5 to 6; and (3) applying the CCLM model (e.g., equation 4) to generate a chroma prediction block Pred_C. In some examples, step (1) is omitted when the spatial resolutions of the collocated luma block and the chroma block are the same, and step (2) is based on the reconstructed adjacent luma samples.
[0132] Fig.10 Examples of reconstructed adjacent luma samples and reconstructed adjacent chroma samples used in CCLM derivation are shown. A chroma block (1000) is being reconstructed. The width and height of the chroma block (1000) are M (e.g., 8) and N (e.g., 4), respectively. N and M can be positive integers. A luma block (e.g., a juxtaposed luma block) (1001) juxtaposed with the chroma block (1000) is used to predict the chroma block (1000). The luma block (1001) includes luma samples (1040). The luma block (1001) can have any suitable width and any suitable height. Fig.10 In the example shown in , the width and height of the luminance block (1001) are 2M (eg, 16) and 2N (eg, 8), respectively.
[0133] The adjacent chroma samples (1010) (e.g., shades of gray) of the chroma block (1000) have been reconstructed. The adjacent luma samples (1020) of the luma block (1001) have been reconstructed. The adjacent chroma samples (1010) may include top adjacent chroma samples (1011) and left adjacent chroma samples (1012). The adjacent luma samples (1020) of the luma block (1001) may include top adjacent luma samples (1021) and left adjacent luma samples (1022).
[0134] exist Fig.10 In the example of , adjacent luma samples (1020) are subsampled or downsampled to generate downsampled adjacent luma samples (1030) (e.g., shades of gray) to match the number of adjacent chroma samples (e.g., 12). Fig.10 In the example shown in , adjacent chroma samples ( 1010 ) and downsampled adjacent luma samples ( 1030 ) may be used to determine parameters a and b, such as shown in Equations 5 to 6.
[0135] In the example, the adjacent chroma samples of the chroma block (1000) include only the top adjacent chroma samples (1011) and do not include the left adjacent chroma samples (1012). Correspondingly, the adjacent luma samples of the luma block (1001) include only the top adjacent luma samples (1021) and do not include the left adjacent luma samples (1022). As described above, the top adjacent luma samples (1021) may be downsampled (e.g., the downsampled luma samples are shaded gray) to match the number (e.g., 8) of the top adjacent chroma samples (1011).
[0136] In the example, the adjacent chroma samples of the chroma block (1000) include only the left adjacent chroma samples (1012) and do not include the top adjacent chroma samples (1011). Correspondingly, the adjacent luma samples of the luma block (1001) include only the left adjacent luma samples (1022) and do not include the top adjacent luma samples (1021). As described above, the left adjacent luma samples (1022) can be downsampled (e.g., the downsampled luma samples are shaded gray) to match the number (e.g., 4) of the left adjacent chroma samples (1012).
[0137] In some examples, chroma samples in neighboring reconstructed blocks of the chroma block (1000) and corresponding luma samples in neighboring reconstructed blocks of the luma block (1001) may be used to determine parameters a and b.
[0138] The chroma block (1000) may be a Cb block in a Cb channel or a Cr block in a Cr channel. In an example, for each chroma channel (e.g., Cr or Cb), parameters a and b may be determined separately. For example, the parameters a and b of the Cr block are determined based on the chroma reconstruction neighboring samples of the Cr block, and the parameters a and b of the Cb block are determined based on the chroma reconstruction neighboring samples of the Cb block.
[0139] Fig.10 An example of reconstructed adjacent luma samples and reconstructed adjacent chroma samples used in CCLM derivation of a chroma block (1000) having a rectangular shape is shown. Fig.10 The examples in can be adapted to chroma blocks and collocated luminance blocks having any shape, such as a square shape. Fig.11 An example of reconstructed adjacent luma samples (1120) and reconstructed adjacent chroma samples (1110) used in CCLM derivation of a chroma block (1100) is shown. The width and height of the chroma block (1100) are W1 (e.g., 8) and H1 (e.g., 8), respectively, where W1 is equal to H1. A luma block (e.g., a collocated luma block) (1101) collocated with the chroma block (1100) is used to predict the chroma block (1100). The width and height of the luma block (1101) are 2W1 (e.g., 16) and 2H1 (e.g., 16), respectively.
[0140] In some embodiments, the reference samples used to generate the linear model parameters a and b are noisy and / or not very representative of the content in the actual prediction block. Therefore, such predictions may be suboptimal for coding efficiency. Therefore, the present disclosure develops a more content-adaptive linear model for chroma sample prediction.
[0141] The present disclosure describes adaptive parameter selection for cross-component prediction in image and video compression. In some embodiments, the slope parameter and / or offset parameter (e.g., parameter a and / or parameter b) used in CCLM prediction is adaptively adjusted so that the linear model can be more content-adaptive for chrominance sample prediction. For example, the linear model (e.g., parameter a and / or parameter b) is more adaptive to the content of the collocated luminance block. In an example, the content of the chrominance block is related to the content of the collocated luminance block, and therefore the CCLM prediction can adaptively predict the content of the chrominance block.
[0142] Adjustments to the slope parameter and / or offset parameter may modify a linear function (e.g., Equation 4) that maps luma sample values to chroma sample values such that chroma sample values may be mapped from luma sample values according to properties of luma sample values in a collocated luma block.
[0143] As described above, in CCLM, a model with two parameters (e.g., parameters a and b in Equation 4) is used to map luma values to chroma values. Slope parameter a and offset parameter b are used in Equation 4. Adjustments to parameters a and b (e.g., slope parameter and offset parameter) may be employed to update the linear model in Equation 4 to an updated linear model in Equation 7.
[0144] Pred_C(x, y) = a′×Rec_L′(x, y)+b′ Equation 7
[0145] In Equation 7, the updated slope parameter a' is a function of the slope parameter a, e.g., a'=f1(a), and the updated offset parameter b' is a function of the offset parameter b, e.g., b'=f2(b). The adjustment of parameters a and b can be based on local sample information, such as local luminance reconstruction sample information of a collocated luminance block.
[0146] FIG. 12A to FIG. 12B An example of parameter adjustment used in CCLM is shown. Fig. 12A A CCLM utilizing Equation 4 is shown. The predicted chroma sample values (e.g., Cb of a Cb block or Cr of a Cr block) have a linear relationship (1201) with the corresponding luma value (e.g., Y) corresponding to Equation 4, wherein the slope of the linear relationship (1201) is the slope parameter a, and the offset of the linear relationship (1201) is the offset parameter b. Fig. 12B Two linear relationships (1201) to (1202) corresponding to Equation 4 and Equation 7, respectively, are shown. In the linear relationship (1202) corresponding to Equation 7, the predicted chroma sample value (e.g., Cb of a Cb block or Cr of a Cr block) has a linear relationship (1202) with the corresponding luma value (e.g., Y), where the slope is the updated slope parameter a' and the offset is the updated offset parameter b'.
[0147] exist Fig. 12B In the example shown in , the updated slope parameter a' is a linear function of the slope parameter a, where a' is (a+u). The adjustment parameter u used to adjust the slope parameter a may be referred to as the slope adjustment parameter "u". The updated offset parameter b' may be a linear function of the offset parameter b. In the example, b'=bu×y r The adjustment parameter y used to adjust the offset parameter b r Can be called the offset adjustment parameter "y r ". Reference Fig. 12B , with adjustments (e.g., a'=a+u and b'=bu×y r ), the linear relationship (1201) (e.g., the mapping function used in Equation 4) is tilted or rotated around point (1203) to generate a linear relationship (1202). In this example, the adjustment parameter y rrepresents the brightness value (e.g., Y is y) at the point (1203) where the linear relationships (1201) and (1202) intersect. r ).
[0148] In an example, the slope parameter and the offset parameter are adjusted. In an example, one of the slope parameter and the offset parameter is adjusted. In some examples, the adjustment parameters u and y are determined or derived. r In some examples, the adjustment parameters u and y are determined r one of the.
[0149] In a' (e.g., a'=a+u) and b' (e.g., b'=bu×y r ), the adjustment parameter u may be encoded or derived. The value of the adjustment parameter u may be positive or negative. The adjustment parameter y for adjusting the offset parameter b may be determined based on the collocated luminance block. r (such as in b' = bu × y r For example, the adjustment parameter y is determined based on the selected value of the corresponding reference luma sample in the collocated luma block. r .
[0150] In CCLM, Equation 7 can be used to predict a chrominance block (e.g., chrominance block (1000)) based on a collocated luminance block (e.g., luminance block (1001)). As described above, the updated offset parameter b' can be a linear function of the offset parameter b, such as b' is (bu×y r According to an embodiment of the present disclosure, the adjustment parameter y used to adjust the offset parameter b is r The sub-set of luma samples in the collocated luma block (e.g., luma block (1001)) may be determined based on a sub-set of luma samples in the collocated luma block (e.g., luma block (1001)). In some embodiments, the sub-set of luma samples in the collocated luma block (e.g., luma block (1001)) does not include one or more samples in the collocated luma block.
[0151] Fig.13 It is shown that it can be used to determine the adjustment parameter y r 7. An example of sample positions of luma samples selected in a collocated luma block (1300). The collocated luma block (1300) is collocated with a chroma block to be encoded (e.g., to be reconstructed). The collocated luma block (1300) includes luma samples that have been reconstructed (e.g., reference luma samples). The collocated luma block (1300) or a downsampled luma block downsampled from the collocated luma block (1300) may be used as Rec_L' in Equation 7 to predict the chroma block. In addition, one or more of the luma samples selected in the collocated luma block (1300) may be used to determine the adjustment parameter y r .
[0152] The luma samples in the collocated luma block (1300) may include luma samples (e.g., selected luma samples) (1301) to (1309). Luma samples (1301) to (1309) are located at the upper left corner, upper center, upper right corner, leftmost center, center, rightmost center, lower left corner, bottom center, and lower right corner of the collocated luma block (1300), respectively.
[0153] Adjust parameter y r The sample value may be selected as one of the luma samples (1301) to (1309). Adjustment parameter y r It may be selected as an average (eg, a weighted average) of sample values of a plurality of samples among the luma samples (1301) to (1309).
[0154] In some embodiments, the adjustment parameter y r The sample value of the luma sample (1309) selected to be at the lower right corner of the collocated luma block (1300).
[0155] In some embodiments, the adjustment parameter y r The sample value selected to be the luma sample (1305) at the center of the collocated luma block (1300).
[0156] In some embodiments, the adjustment parameter y r The sample value selected is the luma sample (1306) at the rightmost center of the collocated luma block (1300).
[0157] In some embodiments, the adjustment parameter y r The sample value selected is the luma sample (1308) at the bottom center of the collocated luma block (1300).
[0158] In some embodiments, the adjustment parameter y r Selected as a sample value from a position such as the upper left corner (e.g., corresponding to luma sample (1301)), the upper right corner (e.g., corresponding to luma sample (1303)), or the lower left corner (e.g., corresponding to luma sample (1307)) of the collocated luma block (1300).
[0159] In some embodiments, the adjustment parameter y r Determined as the average (e.g., a weighted average) of sample values of certain luma samples, such as luma samples (1301), (1303), (1307), and (1309) located at four corners (e.g., upper left, upper right, lower left, and lower right) of a juxtaposed luma block (1300).
[0160] In some embodiments, the adjustment parameter y rDetermined as an average (eg, a weighted average) of sample values of a plurality of samples among luma samples (1301) to (1309) in the concatenated luma block (1300).
[0161] The brightness samples (1301) to (1309) can be used to determine the adjustment parameter y r The adjustment parameter y may also be determined using one or more other luma samples in the collocated luma block (1300). r .
[0162] The luma samples selected in the juxtaposed luma blocks are used to adjust the parameters a and b in the CCLM model. Some methods can be used to make the selection of luma samples more adaptive to different chroma blocks or to different regions (eg, sub-blocks) within a chroma block.
[0163] In some embodiments, different adjustment parameters yr are used for different chrominance blocks. For example, luma samples located at different positions in the corresponding luma block may be used as adjustment parameters yr for different chrominance blocks. r For a chrominance block (such as Fig.13 In the example, one of the luma samples (1305) at the center of the collocated luma block (1300) and (ii) the luma sample (1309) at the lower right corner of the collocated luma block (1300) may be selected to determine the adjustment parameter y for adjusting the offset parameter b. r For the chrominance block, a selection indication may be signaled or derived, for example indicating which sample(s) are used to determine the adjustment parameter y r The selection indication may be a selection index or a selection flag.
[0164] Different methods can be used to obtain the adjustment parameter y for different chromaticity blocks r In the example, the adjustment parameter y indicating the first chrominance block is determined based on a luma sample at a center position of a first luma block juxtaposed with the first chrominance block. r Therefore, the adjustment parameter y of the first chrominance block is determined based on the luma sample at the center position of the first luma block. r The adjustment parameter y indicating the second chrominance block is determined based on the luma sample at the rightmost center position of the second luma block juxtaposed with the second chrominance block. r Therefore, the adjustment parameter y of the second chrominance block is determined based on the luminance sample at the rightmost center position of the second luminance block. r .
[0165] In some embodiments, different sample values in the juxtaposed luma block may be used to determine adjustment parameters y for different regions in the chroma block. r . Fig.14 An example of a chroma block (1410) and a collocated luma block (1400) collocated with the chroma block (1410) is shown. For example, when the spatial resolution of the chroma block (1410) is the same as the spatial resolution of the collocated luma block (1400) (e.g., the color format is 4:4:4), the chroma block (1410) can be predicted based on the collocated luma block (1400). For example, when the spatial resolution of the chroma block (1410) is different from the spatial resolution of the collocated luma block (1400), the chroma block (1410) can be predicted based on a downsampled luma block downsampled from the collocated luma block (1400).
[0166] The chroma block (1410) includes a plurality of regions (also referred to as a plurality of chroma regions), such as an upper left quarter region (1411), an upper right quarter region (1412), a lower left quarter region (1413), and a lower right quarter region (1414). An adjustment parameter y for each of the plurality of regions in the chroma block (1410) may be determined based on a corresponding sample in the collocated luma block (1400). r .
[0167] For example, the center position (1401) of the juxtaposed luma block (1400) is used to determine the adjustment parameter y of the upper left quarter region (1411) of the chroma block (1410). r The rightmost center position (1402) of the juxtaposed luminance block (1400) is used to determine the adjustment parameter y of the upper right quarter region (1412) of the chrominance block (1410) r The bottom center position (1403) of the juxtaposed luminance block (1400) is used to determine the adjustment parameter y of the lower left quarter area (1413) of the chrominance block (1410). r The lower right corner position (1404) of the juxtaposed luminance block (1400) is used to determine the adjustment parameter y of the lower right quarter area (1414) of the chrominance block (1410) r .
[0168] In some embodiments, an adjustment parameter y for one of the plurality of regions in the chrominance block (1410) is determined based on a corresponding region in the collocated luminance block (1400). r For example, the corresponding region in the luminance block (1400) is concatenated with the one of the plurality of regions in the chrominance block (1410).
[0169] Reference Fig.14, the juxtaposed luminance block (1400) includes a plurality of regions (also referred to as a plurality of luminance regions) juxtaposed with a plurality of chrominance regions in the chrominance block (1410). For example, the plurality of luminance regions in the juxtaposed luminance block (1400) include an upper left quarter region (1421), an upper right quarter region (1422), a lower left quarter region (1423), and a lower right quarter region (1424) juxtaposed with an upper left quarter region (1411), an upper right quarter region (1412), a lower left quarter region (1413), and a lower right quarter region (1414) in the chrominance block, respectively.
[0170] In the example, the average value of the luma sample values in the upper left quarter region (1421) of the collocated luma block (1400) is used to determine the adjustment parameter y of the upper left quarter region (1411) in the chroma block (1410). r The average value of the luminance sample values of the upper right quarter area (1422) of the juxtaposed luminance block (1400) is used to determine the adjustment parameter y of the upper right quarter area (1412) in the chrominance block (1410) r The average value of the luminance sample values of the lower left quarter area (1423) of the juxtaposed luminance block (1400) is used to determine the adjustment parameter y of the lower left quarter area (1413) in the chrominance block (1410) r ; and the average value of the luma sample values of the lower right quarter region (1424) of the juxtaposed luma block (1400) is used to determine the adjustment parameter y of the lower right quarter region (1414) in the chroma block (1410) r .
[0171] In the example, the chroma region (e.g., (1411)) is further divided into a plurality of chroma sub-regions, and the collocated luma region (e.g., (1421)) is further divided into a plurality of luma sub-regions. The adjustment parameter y of each chroma sub-region in the region (1411) may be determined based on the average value of the luma sample values of the corresponding luma sub-region in the region (1421). r .
[0172] Fig.15 An example of a chrominance block (1510) and a collocated luma block (1500) collocated with the chrominance block (1510) is shown. Fig.14 As described in , the chrominance block (1510) can be predicted based on the collocated luma block (1500) or a downsampled luma block downsampled from the collocated luma block (1500).
[0173] The chroma block (1510) includes a plurality of chroma regions, such as an upper left quarter region (1511), an upper right quarter region (1512), a lower left quarter region (1513), and a lower right quarter region (1514). The adjustment parameter y of each chroma region in the chroma block (1510) may be determined based on an average of certain luma sample values, such as four luma sample values at four corners of a corresponding luma region in the juxtaposed luma block (1500). r .
[0174] The juxtaposed luma block (1500) includes a plurality of luma regions juxtaposed with a plurality of chroma regions in the chroma block (1510). For example, the plurality of luma regions in the juxtaposed luma block (1500) include an upper left quarter region (1521), an upper right quarter region (1522), a lower left quarter region (1523), and a lower right quarter region (1524) juxtaposed with an upper left quarter region (1511), an upper right quarter region (1512), a lower left quarter region (1513), and a lower right quarter region (1514) in the chroma block, respectively.
[0175] In the example, the average value of the luma sample values at the four corners (1501) to (1504) of the upper left quarter area (1521) of the collocated luma block (1500) is used to determine the adjustment parameter y of the upper left quarter area (1511) in the chroma block (1510). r The average value of the luminance sample values at the four corners of the upper right quarter area (1522) of the juxtaposed luminance block (1500) is used to determine the adjustment parameter y of the upper right quarter area (1512) in the chrominance block (1510) r The average value of the luminance sample values at the four corners of the lower left quarter area (1523) of the juxtaposed luminance block (1500) is used to determine the adjustment parameter y of the lower left quarter area (1513) in the chrominance block (1510) r ; and the average value of the luma sample values at the four corners of the lower right quarter area (1524) of the juxtaposed luma block (1500) is used to determine the adjustment parameter y of the lower right quarter area (1514) in the chroma block (1510) r .
[0176] In reference Figure 14 to Figure 15 In the description, the chrominance block (e.g., 1410) includes four regions (e.g., (1411) to (1414)). When the chrominance block includes M number of regions (where M is an integer greater than 1), Figure 14 to Figure 15 The description can also be applied to the chrominance block.
[0177] Return to reference Fig.14, the regions (1411) to (1414) may be sub-blocks in the chroma block (1410). The indication associated with the chroma block (1410) may indicate a CCLM mode of the chroma block (1410). Fig.14 The embodiments described in the foregoing predict each sub-block in the chroma block (1410) differently. In an example, a forward transform or an inverse transform is applied to the entire chroma block (1410) to transform the entire chroma block (1410). In an example, the chroma block (1410) is divided into a plurality of TBs, the plurality of TBs being different from the regions (1411) to (1414), and each of the plurality of TBs is transformed using an appropriate transform.
[0178] like Figure 14 to Figure 15 As described in the above, the adjustment parameters y of different regions (e.g., regions (1411) to (1414)) in the chrominance block (e.g., (1410)) are r In an example, the same offset parameter b, such as that determined using Equation 6, is used for different regions, and the offset parameter b and the corresponding adjustment parameter y for different regions may be used. r Corresponding updated offset parameters b' for different regions are obtained. In an example, different offset parameters b may be used for different regions.
[0179] Adjustment equations (e.g., a'=a+u and b'=bu×y) may be derived based on (i) reconstructed neighboring luma samples (e.g., (1020)) of the collocated luma block (e.g., luma block (1001)) and (ii) collocated luma samples (e.g., 1040) in the collocated luma block (e.g., luma block (1001)). r According to an embodiment of the present disclosure, the adjustment parameter u may be derived based on the difference between (i) the reconstructed adjacent luma samples and (ii) the collocated luma samples in the collocated luma block.
[0180] In some embodiments, the adjustment parameter u is determined based on the difference between a first average value (referred to as Rec_L'(Nei)) based on reconstructed adjacent luma samples of a collocated luma block and a second average value (referred to as Rec_L'(Col)) based on collocated luma samples in the collocated luma block. In an example, the adjustment parameter u is a linear function of the difference, for example u=M(Rec_L'(Col)-Rec_L'(Nei))+K, where M and K are constants. In an example, the adjustment parameter u is a piecewise linear function of the difference, where a mapping table can be designed to map the difference between the first average value Rec_L'(Nei) and the second average value Rec_L'(Col) to the value of the adjustment parameter u.
[0181] The first average value Rec_L'(Nei) may be an average sample value of a plurality of samples among the reconstructed adjacent luma samples of the collocated luma block. The second average value Rec_L'(Col) may be an average sample value of a plurality of samples among the collocated luma samples in the collocated luma block.
[0182] In an example, the first average value Rec_L'(Nei) is the average sample value of all reconstructed adjacent luma samples (e.g., (1020)) of the collocated luma block (e.g., (1001)). In an example, the second average value Rec_L'(Col) is the average sample value of all collocated luma samples (e.g., (1040)) in the collocated luma block (e.g., (1001)).
[0183] In some embodiments, the reconstructed neighboring luma samples of the collocated luma block used to calculate the first average value Rec_L'(Nei) are selected neighboring samples Rec_L'(i) used in calculating CCLM parameters a and b (such as used in Equations 5 to 6). Fig.10 , the selected neighboring samples Rec_L'(i) may include a subset of samples in the reconstructed neighboring luma samples (1020), such as the left neighboring luma samples (1022), the top neighboring luma samples (1021), etc.
[0184] In some embodiments, the second average value Rec_L'(Col) may be determined based on a subset of collocated luma samples selected from the collocated luma block. Fig.13 , one or more of the samples (1301) to (1309) may be used for the selected subset.
[0185] In some embodiments, the number of samples used to calculate the second average value Rec_L'(Col) is set based on the number of samples used to calculate the first average value Rec_L'(Nei). For example, the number of samples used to calculate the second average value Rec_L'(Col) is set to be equal to the number of samples used to calculate the first average value Rec_L'(Nei).
[0186] In some examples, the first average value Rec_L'(Nei) is determined based on downsampled neighboring samples of reconstructed neighboring luma samples of the collocated luma block.
[0187] In an example, the same scaling parameter a (also referred to as slope parameter a) such as determined using Equation 5 is used for different regions in the chroma block, and corresponding updated scaling parameters a' for different regions can be obtained based on the scaling parameters a for different regions and the corresponding adjustment parameters u. In an example, different scaling parameters a can be used for different regions.
[0188] Fig.16 A flow chart of an overview process (e.g., encoding process) (1600) according to an embodiment of the present disclosure is shown. Process (1600) can be performed by a device for video encoding, which may include a processing circuit system. In various embodiments, process (1600) is performed by a processing circuit system in a device (such as a processing circuit system in terminal devices (310), (320), (330) and (340), a processing circuit system that performs the functions of a video encoder (e.g., (403), (603), (703)), etc.). In some embodiments, process (1600) is implemented as software instructions, so when the processing circuit system executes the software instructions, the processing circuit system performs process (1600). The process starts at (S1601) and proceeds to (S1610).
[0189] At (S1610), for a first region in a chroma block to be encoded using a cross-component linear model (CCLM) mode in a current picture, a first adjustment parameter (e.g., y) for adjusting an offset parameter (e.g., b) in the CCLM mode may be determined based on a first subset of reconstructed samples in a collocated luma block in the current picture. r ). The first subset of reconstructed samples does not include one or more samples in the collocated luma block.
[0190] In an example, the first region in the chrominance block includes the entire chrominance block. The first subset of reconstructed samples in the collocated luma block is a sample in the collocated luma block. Fig.13 As described in , the first adjustment parameter may be determined as a sample value of a sample in the collocated luminance block.
[0191] In an example, the first region includes the entire chrominance block. The first subset of reconstructed samples includes a plurality of samples in the collocated luminance block. Fig.13 As described in , the first adjustment parameter may be determined as an average of sample values of a plurality of samples in the concatenated luminance block.
[0192] At (S1620), a first updated offset parameter may be determined based on at least the offset parameter and the first adjustment parameter.
[0193] At (S1630), the first region may be encoded using the CCLM mode based on at least the first updated offset parameter.In an example, prediction information is encoded, the prediction information indicating that the CCLM mode is applied to the chroma block.
[0194] The encoded first region and the prediction information may be included in an encoded video bitstream and sent to a decoder.
[0195] In an example, the first region includes the entire chrominance block. The prediction information further indicates samples in the collocated luma block that are included in the first subset of reconstructed samples. In an example, samples in the collocated luma block that are included in the first subset of reconstructed samples are signaled in the coded video bitstream.
[0196] Then, the process proceeds to (S1699) and terminates.
[0197] The process (1600) may be suitably adapted to various scenarios, and the steps in the process (1600) may be adjusted accordingly. One or more of the steps in the process (1600) may be adjusted, omitted, repeated, and / or combined. The process (1600) may be implemented in any suitable order. Additional steps may be added.
[0198] In an example, the chroma block also includes a second region. The first subset of reconstructed samples is a first sample in the collocated luma block. For the second region in the chroma block, a second adjustment parameter for adjusting an offset parameter in the CCLM mode may be determined based on the second sample in the collocated luma block. The second sample may be different from the first sample. A second updated offset parameter may be determined based at least on the offset parameter and the second adjustment parameter. The second region in the chroma block may be reconstructed using the CCLM mode based at least on the second updated offset parameter.
[0199] In an example, the chroma block further includes a second region. The juxtaposed luminance block includes a first luminance region and a second luminance region juxtaposed with the first region and the second region, respectively. The first subset of reconstructed samples includes a plurality of samples in the first luminance region. For the second region in the chroma block, a second adjustment parameter for adjusting an offset parameter in the CCLM mode may be determined based on the plurality of samples in the second luminance region. A second updated offset parameter may be determined based on at least the offset parameter and the second adjustment parameter. The second region in the chroma block may be reconstructed using the CCLM mode based at least on the second updated offset parameter.
[0200] In the example, the chroma block also includes a second area. The juxtaposed luminance block includes a first luminance area and a second luminance area juxtaposed with the first area and the second area, respectively. The first subset of reconstructed samples includes an upper left sample, an upper right sample, a lower left sample, and a lower right sample in the first luminance area. The first adjustment parameter can be determined as an average sample value of the upper left sample, the upper right sample, the lower left sample, and the lower right sample in the first luminance area. For the second area in the chroma block, a second adjustment parameter for adjusting the offset parameter in the CCLM mode can be determined as an average sample value of the upper left sample, the upper right sample, the lower left sample, and the lower right sample in the second luminance area. A second updated offset parameter can be determined based at least on the second offset parameter and the second adjustment parameter. The second area in the chroma block can be reconstructed based at least on the second updated offset parameter using the CCLM mode.
[0201] In an example, for the first region, the updated scaling parameter (e.g., a') may be determined as the sum of the scaling parameter used in the CCLM mode (e.g., parameter a) and the adjustment parameter used to adjust the scaling parameter (e.g., adjustment parameter u). The first updated offset parameter (e.g., b') may be determined as (bu×y r ), where b is the offset parameter, u is the adjustment parameter used to adjust the scaling parameter, and y r is the first adjustment parameter. A first region in the chroma block may be reconstructed based on the first updated offset parameter and the updated scaling parameter using the CCLM mode.
[0202] In some embodiments, for the first region, an adjustment parameter (e.g., u) for adjusting a scaling parameter (e.g., a) is determined based on (i) reconstructed luma samples in one or more adjacent luma blocks of the collocated luma block and (ii) samples in the collocated luma block. In an example, for the first region, the adjustment parameter for adjusting the scaling parameter may be determined based on a difference between an average sample value of the reconstructed luma samples in the one or more adjacent luma blocks and an average sample value of the samples in the collocated luma block.
[0203] Fig.17A flowchart of an overview process (e.g., a decoding process) (1700) according to an embodiment of the present disclosure is shown. The process (1700) can be used for a video decoder. The process (1700) can be performed by a device for video encoding, which may include a receiving circuit system and a processing circuit system. The processing circuit system in the device (such as the processing circuit system in the terminal device (310), (320), (330) and (340), the processing circuit system that performs the function of the video decoder (410), the processing circuit system that performs the function of the video decoder (510), etc.) can be configured to perform the process (1700). In some examples, the process (1700) is used for a video encoder (e.g., video encoder (403), video encoder (603)). In an example, the process (1700) is performed by a processing circuit system that performs the function of a video encoder (e.g., video encoder (403), video encoder (603)). In some embodiments, the process (1700) is implemented as software instructions, so when the processing circuit system executes the software instructions, the processing circuit system performs the process (1700). The process starts at (S1701) and proceeds to (S1710).
[0204] At (S1710), prediction information of a chroma block to be reconstructed in a current picture may be decoded. The prediction information may indicate that a cross component linear model (CCLM) mode is applied to the chroma block.
[0205] At (S1720), for a first region in the chrominance block, a first adjustment parameter (also referred to as a first adjustment value) (e.g., y) for adjusting (or modifying) an offset parameter (e.g., b) in the CCLM mode may be determined based on a first subset of samples (or reconstructed samples) in a collocated luma block in the current picture. r ). A collocated luma block is a luma block collocated with a chroma block. The first subset of samples does not include one or more samples in the collocated luma block. In an example, the samples in the collocated luma block have been reconstructed.
[0206] In an example, the first region in the chrominance block includes the entire chrominance block. The first subset of samples in the collocated luma block is a sample in the collocated luma block. Fig.13 As described in , the first adjustment parameter may be determined as a sample value of a sample in the collocated luminance block.
[0207] In an example, the first region includes the entire chrominance block. The first subset of samples includes a plurality of samples in the collocated luma block. Fig.13 As described in , the first adjustment parameter may be determined as an average of sample values of a plurality of samples in the concatenated luminance block.
[0208] In an example, the first region includes the entire chrominance block. The prediction information further indicates which of the samples in the collocated luma block are included in the first subset of samples. In an example, which of the samples in the collocated luma block are included in the first subset of samples is signaled in the coded video bitstream. In an example, it is derived which of the samples in the collocated luma block are included in the first subset of samples.
[0209] At (S1730), for a first region in a chrominance block, the first region may be adjusted based on at least an offset parameter (eg, b) and a first adjustment parameter (eg, y r ) Determine a first update offset parameter (eg, b'), such as b'=bu×y as described above r .
[0210] At (S1740), the first region in the chroma block may be reconstructed based on at least the first update offset parameter using the CCLM mode.Then, the process proceeds to (S1799) and terminates.
[0211] The process (1700) may be suitably adapted to various scenarios, and the steps in the process (1700) may be adjusted accordingly. One or more of the steps in the process (1700) may be adjusted, omitted, repeated, and / or combined. The process (1700) may be implemented in any suitable order. Additional steps may be added.
[0212] In an example, the chroma block also includes a second region. The first subset of samples is a first sample in the collocated luma block. For the second region in the chroma block, a second adjustment parameter for adjusting an offset parameter in the CCLM mode may be determined based on the second sample in the collocated luma block. The second sample may be different from the first sample. A second updated offset parameter may be determined based at least on the offset parameter and the second adjustment parameter. The second region in the chroma block may be reconstructed using the CCLM mode based at least on the second updated offset parameter.
[0213] In an example, the chroma block further includes a second region. The juxtaposed luminance block includes a first luminance region and a second luminance region juxtaposed with the first region and the second region, respectively. The first subset of samples includes a plurality of samples in the first luminance region. For the second region in the chroma block, a second adjustment parameter for adjusting an offset parameter in the CCLM mode may be determined based on the plurality of samples in the second luminance region. A second updated offset parameter may be determined based at least on the offset parameter and the second adjustment parameter. The second region in the chroma block may be reconstructed using the CCLM mode based at least on the second updated offset parameter.
[0214] In the example, the chroma block also includes a second area. The juxtaposed luminance block includes a first luminance area and a second luminance area juxtaposed with the first area and the second area, respectively. The first subset of samples includes an upper left sample, an upper right sample, a lower left sample, and a lower right sample in the first luminance area. The first adjustment parameter can be determined as an average sample value of the upper left sample, the upper right sample, the lower left sample, and the lower right sample in the first luminance area. For the second area in the chroma block, a second adjustment parameter for adjusting the offset parameter in the CCLM mode can be determined as an average sample value of the upper left sample, the upper right sample, the lower left sample, and the lower right sample in the second luminance area. A second updated offset parameter can be determined based at least on the second offset parameter and the second adjustment parameter. The second area in the chroma block can be reconstructed based at least on the second updated offset parameter using the CCLM mode.
[0215] In an example, for the first region, the updated scaling parameter (e.g., a') may be determined as the sum of the scaling parameter (also referred to as the slope parameter) (e.g., parameter a) used in the CCLM mode and the adjustment parameter (also referred to as the adjustment value) (e.g., adjustment parameter u) used to adjust the scaling parameter. The first updated offset parameter (e.g., b') may be determined as (bu×y r ), where b is the offset parameter, u is the adjustment parameter used to adjust the scaling parameter, and y r is the first adjustment parameter. A first region in the chroma block may be reconstructed based on the first updated offset parameter and the updated scaling parameter using the CCLM mode.
[0216] In some embodiments, for the first region, an adjustment parameter (e.g., u) for adjusting a scaling parameter (e.g., a) is determined based on (i) reconstructed luma samples in one or more adjacent luma blocks of the collocated luma block and (ii) samples in the collocated luma block. In an example, for the first region, the adjustment parameter for adjusting the scaling parameter may be determined based on a difference between an average sample value of the reconstructed luma samples in the one or more adjacent luma blocks and an average sample value of the samples in the collocated luma block.
[0217] In some embodiments, for a first region in a chroma block, a first adjustment value for modifying an offset parameter in a CCLM mode is determined based on a first subset of reconstructed samples in a luma block juxtaposed with the chroma block in a current picture. The first subset of reconstructed samples does not include one or more samples in the luma block. The offset parameter may be updated based on at least the first adjustment value. A second adjustment value for modifying a slope parameter in the CCLM mode may be determined based on a second subset of reconstructed samples in the luma block, and the slope parameter may be updated based on at least the second adjustment value. The first region in the chroma block may be reconstructed using the CCLM mode based on at least the updated offset parameter and the updated slope parameter.
[0218] In an example, the second subset of reconstructed samples includes the first subset of reconstructed samples.
[0219] In an example, the second adjustment value is determined based on the reconstructed samples in the entire luma block.
[0220] The embodiments in the present disclosure may be used alone or in any order. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0221] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.18 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0222] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, linking, etc. to create code comprising instructions that may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.
[0223] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0224] Fig.18 The components for the computer system (1800) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiment of the computer system (1800).
[0225] The computer system (1800) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs implemented by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, tapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to intentional human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0226] Input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1801), mouse (1802), trackpad (1803), touch screen (1810), data gloves (not shown), joystick (1805), microphone (1806), scanner (1807) and camera (1808).
[0227] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through a touch screen (1810), a data glove (not shown), or a joystick (1805), but there may also be tactile feedback devices that are not used as input devices); audio output devices (e.g., speakers (1809), headphones (not depicted)); visual output devices (e.g., screens (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output through methods such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and cigarette cans (not depicted)); and printers (not depicted).
[0228] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1820) with CD / DVD etc. media (1821), thumb drives (1822), removable hard drives or solid-state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD based devices such as security dongles (not depicted), etc.
[0229] Those skilled in the art will also appreciate that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0230] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1849) (e.g., a USB port of the computer system (1800)); other networks are typically integrated into the core of the computer system (1800) by attaching to a system bus as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smart phone computer system via a cellular network interface). The computer system (1800) can use any of these networks to communicate with other entities. Such communications may be one-way receive-only (e.g., broadcast television), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way, such as to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used on each of these networks and network interfaces as described above.
[0231] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core ( 1840 ) of the computer system ( 1800 ).
[0232] The core (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1843), hardware accelerators (1844) for certain tasks, graphics adapters (1850), etc. These devices as well as read-only memory (ROM) (1845), random access memory (1846), internal large-capacity storage devices (1847) such as internal non-user accessible hard drives, SSDs, etc. can be connected through a system bus (1848). In some computer systems, the system bus (1848) can be accessed in the form of one or more physical plugs to achieve expansion through additional CPUs, GPUs, etc. Peripheral devices can be attached to the system bus (1848) of the core directly or through a peripheral bus (1849). In an example, a screen (1810) can be connected to a graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.
[0233] The CPU (1841), GPU (1842), FPGA (1843) and accelerator (1844) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1845) or RAM (1846). Transition data can also be stored in RAM (1846), while permanent data can be stored in, for example, an internal mass storage device (1847). Fast storage and retrieval of any storage device in the storage device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1841), GPUs (1842), mass storage devices (1847), ROMs (1845), RAMs (1846), etc.
[0234] The computer readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.
[0235] The embodiments of the present application also provide a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method provided in the above embodiments.
[0236] As an example and not limitation, a computer system (1800) having an architecture, in particular a core (1840), can provide functions provided as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as described above, as well as certain storage devices of the core (1840) having a non-transitory nature, such as a mass storage device (1847) or ROM (1845) inside the core. Software that implements various embodiments of the present disclosure can be stored in such a device and executed by the core (1840). Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core (1840) - in particular the processor therein (including a CPU, GPU, FPGA, etc.) - to perform a specific process or a specific part of a specific process described herein, including defining a data structure stored in the RAM (1846) and modifying such a data structure according to a process defined by the software. Additionally or alternatively, the computer system may provide functionality due to logic hardwired or otherwise embodied in circuits (e.g., accelerators (1844)) that may operate in place of or in conjunction with software to perform specific processing or specific portions of specific processing described herein. Where appropriate, references to software may include logic, and conversely references to logic may include software. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0237] Appendix A: Acronyms
[0238] JEM: Joint Development Model
[0239] VVC: Versatile Video Coding
[0240] BMS: Benchmark Set
[0241] MV: Motion Vector
[0242] HEVC: High Efficiency Video Codec
[0243] SEI: Supplemental Enhancement Information
[0244] VUI: Video Availability Information
[0245] GOPs: Group of Pictures
[0246] TUs: Transformation Units
[0247] PUs: Prediction Units
[0248] CTUs: Coding Tree Units
[0249] CTBs: Coding Tree Blocks
[0250] PBs: prediction blocks
[0251] HRD: Hypothetical Reference Decoder
[0252] SNR: Signal to Noise Ratio
[0253] CPUs: Central Processing Units
[0254] GPUs: Graphics Processing Units
[0255] CRT: cathode ray tube
[0256] LCD: Liquid Crystal Display
[0257] OLED: Organic Light Emitting Diode
[0258] CD: Compact Disc
[0259] DVD: Digital Video Disc
[0260] ROM: Read Only Memory
[0261] RAM: Random Access Memory
[0262] ASIC: Application-Specific Integrated Circuit
[0263] PLD: Programmable Logic Device
[0264] LAN: Local Area Network
[0265] GSM: Global System for Mobile Communications
[0266] LTE: Long Term Evolution
[0267] CANBus: Controller Area Network Bus
[0268] USB: Universal Serial Bus
[0269] PCI: Peripheral Component Interconnect
[0270] FPGA: Field Programmable Gate Array
[0271] SSD: Solid State Drive
[0272] IC: Integrated Circuit
[0273] CU: Coding Unit
[0274] CCLM: Cross-Component Linear Model
[0275] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that, although not explicitly shown or described herein, those skilled in the art will be able to conceive of many systems and methods that implement the principles of the present disclosure and are therefore within its spirit and scope.
Claims
1. A video processing method in an encoder, include: For the first region in the chrominance block in the current picture, updating offset parameters in a cross-component linear model (CCLM) mode based on a first subset of reconstructed samples in a luma block collocated with the chroma block in the current picture, the first subset of reconstructed samples not including one or more reconstructed samples in the luma block; updating a slope parameter in a CCLM mode based on a second subset of the reconstructed samples in the luma block; as well as The first region in the chroma block is encoded using a CCLM mode based on the updated offset parameter and the updated slope parameter.
2. The method according to claim 1, It is characterized in that The first region in the chroma block is the entire chroma block; The first subset of the reconstructed samples in the luma block is one reconstructed sample in the luma block; and Updating the offset parameter includes updating the offset parameter based on a sample value of a reconstructed sample in the luma block.
3. The method according to claim 1, It is characterized in that The first area is the entire chrominance block; The first subset of the reconstructed samples includes a plurality of reconstructed samples in the luma block; and Updating the offset parameter includes updating the offset parameter based on an average value of sample values of the plurality of reconstructed samples in the luma block.
4. The method according to claim 1, It is characterized in that The first area is the entire chrominance block; and The encoded prediction information of the chrominance block further indicates reconstructed samples in the luma block that are included in the first subset.
5. The method according to claim 1, It is characterized in that The chroma block also includes a second region; The first subset of the reconstructed samples is a first reconstructed sample in the luma block; as well as For the second area in the chrominance block, the method further includes: determining an adjustment parameter for adjusting the offset parameter in the CCLM mode based on a second reconstructed sample in the luma block, the second reconstructed sample being different from the first reconstructed sample; determining a second updated offset parameter based on the offset parameter and the adjustment parameter; and Based on the second updated offset parameter, the second area in the chroma block is encoded using the CCLM mode.
6. The method according to claim 1, It is characterized in that The chroma block also includes a second region; The brightness block includes a first brightness area and a second brightness area respectively juxtaposed with the first area and the second area; The first subset of the reconstructed samples includes a plurality of reconstructed samples in the first luminance region; as well as For the second area in the chrominance block, the method further includes: determining an adjustment parameter for adjusting the offset parameter in the CCLM mode based on a plurality of reconstructed samples in the second luminance region; determining a second updated offset parameter based on the offset parameter and the adjustment parameter; and Based on the second updated offset parameter, the second area in the chroma block is encoded using the CCLM mode.
7. The method according to claim 1, It is characterized in that The chroma block also includes a second region; The brightness block includes a first brightness area and a second brightness area respectively juxtaposed with the first area and the second area; The first subset of the reconstructed samples includes an upper left sample, an upper right sample, a lower left sample, and a lower right sample in the first luminance region; Updating the offset parameter comprises updating the offset parameter with an average sample value of the upper left sample, the upper right sample, the lower left sample, and the lower right sample in the first brightness area; as well as For the second area in the chrominance block, the method further includes: determining an adjustment parameter for adjusting the offset parameter in the CCLM mode as an average sample value of an upper left sample, an upper right sample, a lower left sample, and a lower right sample in the second brightness area; determining a second updated offset parameter based on the offset parameter and the adjustment parameter; and Based on the second updated offset parameter, the second area in the chroma block is encoded using the CCLM mode.
8. The method according to claim 1, It is characterized in that Updating the offset parameter includes: updating the offset parameter to (bu×y r ), b is the offset parameter, u is determined based on the second subset of the reconstructed samples, y r determining based on a first subset of the reconstructed samples; Updating the slope parameter includes updating the slope parameter to (a+u), where a is the slope parameter; and Encoding the first region includes encoding the first region in the chroma block based on the updated offset parameter and the updated slope parameter using a CCLM mode.
9. The method according to claim 8, It is characterized in that Also includes: For the first region, u is determined based on (i) reconstructed luma samples in one or more neighboring luma blocks of the luma block and (ii) the second subset of the reconstructed samples in the luma block.
10. The method according to claim 9, It is characterized in that The determining u comprises: For the first region, u is determined based on a difference between an average sample value of the reconstructed luma samples in the one or more neighboring luma blocks and an average sample value of the second subset of the reconstructed samples in the luma block.
11. The method according to claim 1, It is characterized in that The second subset of the reconstructed samples includes the first subset of the reconstructed samples.
12. The method according to claim 1, It is characterized in that Updating the slope parameter comprises: The slope parameter is determined based on the reconstructed samples in the entire luma block.
13. A video processing method in a decoder, include: For the first region in the chrominance block in the current picture, updating offset parameters in a cross-component linear model (CCLM) mode based on a first subset of reconstructed samples in a luma block collocated with the chroma block in the current picture, the first subset of reconstructed samples not including one or more reconstructed samples in the luma block; updating a slope parameter in a CCLM mode based on a second subset of the reconstructed samples in the luma block; as well as The first region in the chroma block is decoded using a CCLM mode based on the updated offset parameter and the updated slope parameter.
14. A device for video decoding, It is characterized in that include: A processing circuit system, wherein the processing circuit system is configured to perform the method of any one of claims 1 to 12.
15. A device for video decoding, It is characterized in that include: a first updating module for updating an offset parameter in a cross-component linear model (CCLM) mode based on a first subset of reconstructed samples in a luma block collocated with the chroma block in the current picture, the first subset of reconstructed samples not including one or more reconstructed samples in the luma block; A second determination module, configured to update a slope parameter in a CCLM mode based on a second subset of the reconstructed samples in the luminance block; a second updating module, configured to update the slope parameter based at least on the second adjustment value; as well as An encoding module is used to encode the first area in the chroma block using a CCLM mode based on an updated offset parameter and an updated slope parameter.
16. A non-transitory computer-readable storage medium storing instructions, It is characterized in that When the instructions are executed by at least one processor, the at least one processor executes the method of any one of claims 1 to 13.
17. A computer device, It is characterized in that The computer device comprises a processor and a memory: The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 13 according to the computer program.
18. A computer program product comprising a computer program, which, when executed on a computer device, causes the computer device to execute the method according to any one of claims 1 to 13.
19. A method for storing or transmitting a video bitstream, It is characterized in that The video bit stream is generated according to the encoding method according to any one of claims 1 to 12, or is decoded based on the decoding method according to claim 13.