Video decoding method and apparatus, video encoding method and apparatus, and electronic device

CN116965024BActive Publication Date: 2026-09-11TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280011729.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-07-29
Filing Date
2022-08-01
Publication Date
2026-09-11
Estimated Expiration
2042-08-01

Smart Images

  • Figure CN116965024B_ABST
    Figure CN116965024B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a video decoding method and apparatus, a video encoding method and apparatus, and an electronic device. In the video decoding method, encoded information indicates that an adaptive upsampling filter is to be applied to a current block in a current picture. A respective class is determined for each of a plurality of sub-blocks of the current block. A respective set of filter coefficients is determined for each of the plurality of sub-blocks from a plurality of sets of filter coefficients of the adaptive upsampling filter. The respective set of filter coefficients is determined based on at least one class corresponding to the respective sub-block and a respective sampling rate of a reference pixel resampling (RPR) applied to a reference picture of the current picture. The respective sampling rate is associated with one of a plurality of phases of the RPR. The adaptive upsampling filter is applied to the current block to generate filtered reconstructed samples of the current block based on the determined respective set of filter coefficients without applying a secondary adaptive loop filter (ALF).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] References merged

[0002] This application claims priority to U.S. Patent Application 17 / 877,818, filed July 29, 2022, entitled "Adaptive Up-Sampling Filter for Luma and Chroma with Reference Picture Resampling (RPR)," and U.S. Provisional Patent Application 63 / 228,560, filed August 2, 2021, entitled "Adaptive Up-Sampling Filter for Luma and Chroma with Reference Picture Resampling (RPR)." The disclosure of the earlier applications is incorporated herein by reference in its entirety. Technical Field

[0003] This application relates to video encoding and decoding, and in particular to a video decoding method and apparatus, a video encoding method and apparatus, and an electronic device. Background Technology

[0004] The background description provided herein is intended to generally present the context of this application. To the extent described in this background section, the work of the currently named inventors, and aspects of the description that may not have qualified as prior art at the time of filing of this application, are neither expressly nor implicitly considered as prior art to this application.

[0005] Uncompressed digital video can comprise a series of images, each with a certain spatial dimension, such as a 1920×1080 luminance sample and an associated chrominance sample. This series of images can have a fixed or variable frame rate (also commonly referred to as the frame rate), for example, 60 images per second or 60 Hz. Uncompressed video has specific bitrate requirements. For example, a 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at 60Hz frame rate) requires approximately 1.5 Gbit / s of bandwidth. One hour of such video would require over 600 GB of storage space.

[0006] One objective of video encoding and decoding is to reduce redundancy in the input video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage requirements, in some cases by two or more orders of magnitude. Lossless compression, lossy compression, and combinations thereof can all be used for video encoding and decoding. Lossless compression refers to a technique that allows an exact copy of the original signal to be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be exactly the same as the original signal, but the distortion between the original and the reconstructed signal is small enough that the reconstructed signal can be used for the intended application. Lossy compression is widely used in video. The amount of distortion tolerated by lossy compression depends on the application; for example, some streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can be reflected as follows: the higher the permissible / acceptable distortion, the higher the compression ratio that can be produced.

[0007] Video encoders and decoders can use several major categories of techniques, including motion compensation, transform, quantization, and entropy coding.

[0008] Video encoding and decoding techniques can include intra-frame coding. In intra-frame coding, the representation of sample values ​​does not reference samples or other data in a previously reconstructed reference image. In some video encoding and decoding techniques, an image is spatially divided into sample blocks. When all sample blocks are encoded using intra-frame mode, the image can be an intra-frame image. Intra-frame images, and their derived images, such as images refreshed by a separate decoder, can be used to reset the decoder's state and thus can be used as the first image in an encoded video stream and video session, or as a still image. Samples in intra-frame blocks can be transformed, and the transform coefficients can be quantized before entropy coding. Intra-frame prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the transformed DC value and the smaller the AC coefficients, the fewer bits are needed to represent the entropy-coded block given the quantization step size.

[0009] Traditional intra-frame coding techniques, such as the known MPEG-2 coding technique, do not use intra-frame prediction. However, some newer video compression techniques include attempts to use, for example, neighboring sample data and / or metadata, which are obtained during the encoding and / or decoding of spatially adjacent data blocks that are decoded first. Therefore, this technique is called "intra-frame prediction." Note that, at least in some cases, intra-frame prediction uses only reference data from the current frame being reconstructed, and not reference data from a reference frame.

[0010] Intra-prediction can take many forms. When more than one of these techniques can be used in a given video coding scheme, an intra-prediction mode can be used to encode the chosen technique. In some cases, some modes have sub-modes and / or parameters, which can be encoded individually or included in the mode codeword. The codeword used for a given combination of mode / sub-mode and / or parameters will affect the coding efficiency gain through intra-prediction, and the entropy coding technique used to translate the codeword into the bitstream will also have an impact.

[0011] The H.264 standard introduced intra-frame prediction for a specific mode, which was improved upon in the H.265 standard. Newer coding techniques, such as Joint Exploration Model (JEM), Universal Video Coding (VVC), and Baseline Set (BMS), have further refined this process. Predictor blocks can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of neighboring samples are copied into the predictor block in one direction. The reference direction used can be encoded into the bitstream or can be predicted itself. Summary of the Invention

[0012] This disclosure provides methods and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes processing circuitry.

[0013] According to one aspect of this disclosure, a video decoding method executed in a video decoder is provided. In this method, encoded information of a current block in a current image can be received from an encoded video bitstream. The encoded information can indicate the application of an adaptive upsampling filter to the current block. A corresponding category for each of a plurality of sub-blocks of the current block can be determined. A corresponding set of filter coefficients for each of the plurality of sub-blocks can be determined from a plurality of filter coefficient sets of the adaptive upsampling filter. This corresponding set of filter coefficients can be determined based on at least one category corresponding to each sub-block and a corresponding sampling rate of a reference pixel resampling (RPR) applied to a reference image of the current image. The corresponding sampling rate can be associated with one of a plurality of phases of the RPR. The adaptive upsampling filter can be applied to the current block to generate filtered reconstructed samples of the current block based on the determined corresponding set of filter coefficients for the plurality of sub-blocks, without applying a secondary adaptive loop filter (ALF).

[0014] To determine the corresponding category, the category index for each sub-block can be determined based on the directionality and activity of the local gradients of the brightness samples in each sub-block. The category for each sub-block can be determined based on its corresponding category index.

[0015] In some embodiments, the number of multiple filter coefficient sets for the luminance samples of the current block can be equal to the product of N and M. N may be the number of classes associated with the current block, and M may be the number of multiple phases of the RPR associated with the luminance samples of the current block.

[0016] In some embodiments, the number of filter coefficient sets for the chroma samples of the current block can be equal to L, where L can be the number of multiple phases of the RPR associated with the chroma samples of the current block.

[0017] In the method, a first cost value for a first filter coefficient set within a plurality of filter coefficient sets can be determined. The first cost value indicates the distortion between the current block and reconstructed samples of the current block filtered based on the first filter coefficient set. A second cost value for a second filter coefficient set within the plurality of filter coefficient sets can be determined. The second cost value indicates the distortion between the current block and reconstructed samples of the current block filtered based on the second filter coefficient set. One of the first and second filter coefficient sets can be selected based on the smaller of the first and second cost values.

[0018] In some embodiments, a first set of filter coefficients may be associated with a first category index, and a second set of filter coefficients may be associated with a second category index. The second category index may be consecutive to the first category index.

[0019] In some embodiments, a first set of filter coefficients may be associated with a first phase among a plurality of phases of the RPR, and a second set of filter coefficients may be associated with a second phase among a plurality of phases of the RPR, wherein the second phase may be continuous with the first phase of the RPR.

[0020] In some embodiments, the number of multiple filter coefficient sets for the brightness samples of the current block can be equal to the product of N and P. N can be the number of categories associated with the current block, and P can be the number of regions partitioned from the current image. Accordingly, a first filter coefficient set can be associated with a first region of the current image, and a second filter coefficient set can be associated with a second region of the current image.

[0021] In the method described, multiple filter coefficient sets can be included in the adaptive parameter set of the encoded information.

[0022] The adaptive upsampling filter may further include a luminance adaptive upsampling filter and a chrominance adaptive upsampling filter. A high-resolution luminance component of the reconstructed samples of the current block can be generated based on the lower-resolution luminance component of the reconstructed samples of the current block, which serves as input to the luminance adaptive upsampling filter. Similarly, a high-resolution chrominance component of the reconstructed samples of the current block can be generated based on the lower-resolution luminance component and the lower-resolution chrominance component of the reconstructed samples of the current block, which serve as input to the chrominance adaptive upsampling filter.

[0023] According to another aspect of this disclosure, an apparatus is provided. The apparatus includes processing circuitry. The processing circuitry can be configured to perform any of the methods for video encoding / decoding.

[0024] This disclosure also provides a non-volatile computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform any of the methods for video encoding / decoding. Attached Figure Description

[0025] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, in which:

[0026] Figure 1A This is a schematic diagram of an exemplary subset of intra-frame prediction modes.

[0027] Figure 1B This is an illustration of an exemplary intra-frame prediction direction.

[0028] Figure 2 This is a schematic diagram of the current block and its surrounding space merge candidates in an example.

[0029] Figure 3 This is a simplified block diagram of a communication system according to an embodiment.

[0030] Figure 4 This is a simplified block diagram of a communication system according to an embodiment.

[0031] Figure 5 This is a simplified block diagram of the decoder according to an embodiment.

[0032] Figure 6 This is a simplified block diagram of an encoder according to an embodiment.

[0033] Figure 7 A block diagram of an encoder according to another embodiment is shown.

[0034] Figure 8 A block diagram of a decoder according to another embodiment is shown.

[0035] Figure 9 This is a schematic diagram of an adaptive loop filter (ALF) and a cross-component ALF (CC-ALF) according to some embodiments.

[0036] Figure 10A This is a first exemplary schematic diagram of an ALF for reference image resampling (RPR) according to some embodiments.

[0037] Figure 10B This is a second exemplary schematic diagram of an ALF for RPR according to some embodiments.

[0038] Figure 11 This is a schematic diagram of an adaptive upsampling filter according to some embodiments.

[0039] Figure 12 A flowchart outlining an exemplary decoding process according to some embodiments of the present disclosure is shown.

[0040] Figure 13 A flowchart outlining an exemplary coding process according to some embodiments of this disclosure is shown.

[0041] Figure 14 This is a schematic diagram of a computer system according to an embodiment. Detailed Implementation

[0042] refer to Figure 1A The image depicts, to its lower right, a subset of the 33 possible predictor directions of the H.265 standard (corresponding to 33 angular modes of 35 intra-frame modes), with 9 known predictor directions. The convergence point (101) of each arrow represents a sample being predicted. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted based on one or more samples at the upper right corner, which is at a 45-degree angle to the horizontal axis. Similarly, arrow (103) indicates that sample (101) is predicted based on one or more samples at the lower left corner, which is at a 22.5-degree angle to the horizontal axis.

[0043] Still referencing Figure 1A As shown, Figure 1AThe upper left corner depicts a square block (104) with 4×4 samples (represented by a bold dashed line). The square block (104) comprises 16 samples, each labeled with "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is the second (counting from top to bottom) sample in the Y dimension and the first (counting from left to right) sample in the X dimension. Similarly, sample S44 is the fourth sample in both the X and Y dimensions within the block (104). Because the block size is 4×4 samples, S44 is located in its lower right corner. Figure 1A Reference samples are further illustrated, following a similar numbering method. The reference samples are labeled with R, their Y position (e.g., row index) relative to the block (104), and their X position (e.g., column index). In the H.264 and H.265 standards, the predicted samples are adjacent to the blocks being reconstructed; therefore, negative values ​​are not required.

[0044] Intra-frame image prediction can be performed by appropriately copying reference sample values ​​from adjacent samples, based on the prediction direction indicated by the signal. For example, suppose the encoded video stream contains signaling that, for the block, indicates a prediction direction consistent with arrow (102), i.e., the samples in the block are predicted based on one or more reference samples at the upper right corner, which are at a 45-degree angle to the horizontal direction. In this case, samples S41, S32, S23, and S14 are predicted based on the same reference sample R05. Sample S44 is predicted based on reference sample R08.

[0045] In some cases, for example by interpolation, the values ​​of multiple reference samples can be combined to calculate a single reference sample; especially when the direction is not divisible by 45 degrees.

[0046] With the development of video coding technology, the number of possible directions is also increasing. In the H.264 standard (2003), nine different directions could be represented. In the H.265 standard (2013), this increased to 33 directions. At the time of this invention, JEM / VVC / BMS can support up to 65 directions. Some experiments have been conducted to identify the most likely directions. Some entropy coding techniques use a small number of bits to represent these most likely directions, accepting a certain cost for less likely directions. Furthermore, sometimes these directions can be predicted based on the adjacent directions used by adjacent decoded blocks.

[0047] Figure 1B A schematic diagram (110) depicting 65 intra-frame prediction directions according to JEM is shown to illustrate how the number of prediction directions increases over time.

[0048] Depending on the video coding technique, the mapping of intra-prediction direction bits used to represent direction in the encoded video bitstream may also differ; for example, the range can be from a simple, direct mapping of the prediction direction of the intra-prediction mode to codewords, to a complex adaptive scheme involving the most probable mode and similar techniques. However, in all these cases, statistically, some directions are less likely to appear in the video content compared to others. Since the purpose of video compression is to reduce redundancy, in high-performance video coding techniques, these less likely directions are represented using more bits compared to the more probable directions.

[0049] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be a lossy compression technique and can involve using sample data blocks from a previously reconstructed picture or a portion thereof (the reference picture), spatially shifted in a direction indicated by a motion vector (hereinafter referred to as MV), to predict the newly reconstructed picture or picture partition. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or three dimensions, with the third dimension indicating the reference picture in use (the latter may indirectly be a temporal dimension).

[0050] In some video compression techniques, a motion vector (MV) applicable to a given sample data region can be predicted based on other MVs, for example, based on MVs related to another sample data region adjacent to the region being reconstructed and preceding the MV in the decoding order. This significantly reduces the amount of data required to encode the MV, thereby eliminating redundancy and improving compression. For example, MV prediction works effectively because when encoding an input video signal (called the natural video) originating from a camera, there is a statistical probability that regions larger than the region applicable to a single MV move in similar directions, and therefore, in some cases, predictions can be made using similar motion vectors derived from MVs of adjacent regions. This makes the MV found in a given region similar to or the same as the MV predicted based on surrounding MVs, and after entropy encoding, the number of bits representing the MV can be less than the number of bits used in the case of direct encoding of the MV. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself may be lossy, for example, due to rounding errors when calculating predictors based on several surrounding MVs.

[0051] H.265 / HEVC (ITU-T H.265 Recommendation, “High Efficiency Video Coding”, December 2016) describes various MV prediction mechanisms. Among the various MV prediction mechanisms provided by H.265, this application describes a technique called “spatial combining”.

[0052] refer to Figure 2 As shown, the current block (201) includes samples found by the encoder during motion search that can be predicted from the previous block (already spatially shifted) of the same size as the current block. The MV is not directly encoded, but rather derived from metadata associated with one or more reference images (e.g., the most recent (in decoding order) reference image) using the MV associated with any one of the five surrounding samples (denoted as A0, A1, B0, B1, B2 (202 to 206 respectively)). In H.265, MV prediction can use predictors from the same reference image used by its neighboring blocks.

[0053] Figure 3 A simplified block diagram of a communication system (300) according to an embodiment disclosed in this application is illustrated. The communication system (300) includes a plurality of terminal devices that can communicate with each other via, for example, a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). Figure 3 In the example, the first pair of terminal devices (310) and (320) perform one-way data transmission. For example, terminal device (310) may encode video data (e.g., a video image stream captured by terminal device (310)) for transmission over a network (350) to another terminal device (320). The encoded video data may be transmitted in the form of one or more encoded video streams. Terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to recover the video images, and display the video images based on the recovered video data. One-way data transmission is likely to be common in applications such as media services.

[0054] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) performing bidirectional transmission of encoded video data, which may occur, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video images captured by the terminal device) for transmission over a network (350) to the other terminal device (330) and (340). Each of the terminal devices (330) and (340) may also receive encoded video data transmitted by the other terminal device (330) and (340), and may decode the encoded video data to recover video images, and may display the video images on an accessible display device based on the recovered video data.

[0055] exist Figure 3 In the examples, terminal devices (310), (320), (330), and (340) may be illustrated as servers, personal computers, and smartphones, but the principles disclosed herein are not limited thereto. The embodiments disclosed herein are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (350) refers to any number of networks, including, for example, wired (connected) and / or wireless communication networks, that transmit encoded video data between terminal devices (310), (320), (330), and (340). Communication networks (350) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks (LANs), wide area networks (WANs), and / or the Internet. For the purposes of this discussion, unless explained below, the architecture and topology of the network (350) may be irrelevant to the operation of this disclosure.

[0056] As an example of the application of the subject matter disclosed in this application Figure 4 The placement of a video encoder and video decoder in a streaming environment is illustrated. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0057] The streaming system may include an acquisition subsystem (413) that may include a video source (401) such as a digital camera, which creates, for example, an uncompressed video image stream (402). In one example, the video image stream (402) includes samples captured by a digital camera. The video image stream (402) is depicted as a thick line to emphasize its higher data volume compared to encoded video data (404) (or encoded video bitstream). The video image stream (402) may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter as described in more detail below. Encoded video data (404) (or encoded video bitstream (404)) is depicted as a thin line to emphasize its lower data volume compared to the video picture stream (402), which can be stored on a streaming server (405) for future use. One or more streaming client subsystems, for example, Figure 4 Client subsystems (406) and (408) can access a streaming server (405) to retrieve copies (407) and (409) of encoded video data (404). Client subsystem (406) may include, for example, a video decoder (410) in an electronic device (430). The video decoder (410) decodes the incoming copy (407) of the encoded video data and produces an output video picture stream (411) that can be displayed on a display (412) (e.g., a screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (404), video data (407), and video data (409) (e.g., a video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the subject matter disclosed in this application can be used in the context of the VVC standard.

[0058] It should be noted that electronic devices (420) and (430) may include other components (not shown). For example, electronic device (420) may include a video decoder (not shown), and electronic device (430) may also include a video encoder (not shown).

[0059] Figure 5This is a block diagram of a video decoder (510) according to an embodiment disclosed in this application. The video decoder (510) may be disposed in an electronic device (530). The electronic device (530) may include a receiver (531) (e.g., a receiving circuit). The video decoder (510) may be used in place of Figure 4 The video decoder (410) in the example.

[0060] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510); in the same embodiment or another embodiment, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not depicted). The receiver (531) may separate the encoded video sequences from other data. To prevent network jitter, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as "parser (520)"). In some applications, the buffer memory (515) is part of the video decoder (510). In other cases, the buffer memory (515) may be located outside the video decoder (510) (not depicted). In other cases, an external buffer memory (not depicted) may be provided for the video decoder (510) to prevent network jitter, for example, and another buffer memory (515) may be configured internally for, for example, handling broadcast timing. When the receiver (531) receives data from a store / forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be necessary, or it may be made smaller. For use on best-effort packet networks such as the Internet, a buffer memory (515) may also be required; this buffer memory may be relatively large and adaptive in size, and may be at least partially implemented in the operating system or a similar component (not depicted) external to the video decoder (510).

[0061] The video decoder (510) may include a parser (520) to reconstruct symbols (521) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (510) and potential information for controlling a display device (512) (e.g., a display screen), which is not part of the electronic device (530) but may be coupled to it, such as... Figure 5As shown. The control information for the display device may be a parameter set fragment (not depicted) of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI). The parser (520) can parse / decode the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (520) can extract a subgroup parameter set of at least one subgroup of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (520) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0062] The parser (520) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (515) to create symbols (521).

[0063] Depending on the type of encoded video picture or portion of encoded video picture (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (521) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (520). For the sake of brevity, the flow of such subgroup control information between the parser (520) and the various units described below is not described.

[0064] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and may at least be partially integrated with each other. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into the functional units described below.

[0065] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives quantization transform coefficients as symbols (521) and control information from the parser (520), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values, which can be input into the aggregator (555).

[0066] In some cases, the output samples of the scaler / inverse transform unit (551) may belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed images but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses surrounding reconstructed information extracted from the current picture buffer (558) to generate blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (558) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (555) adds the predictive information generated by the intra-picture prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551) based on each sample.

[0067] In other cases, the output samples of the scaler / inverse transform unit (551) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (553) may access the reference image memory (557) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (521) belonging to the block, these samples may be added by the aggregator (555) to the output of the scaler / inverse transform unit (551) (referred to in this case as residual samples or residual signals) to generate output sample information. The motion compensation prediction unit (553) may obtain the predicted samples from the address in the reference image memory (557) under motion vector control, and the motion vector is available to the motion compensation prediction unit (553) in the form of the symbols (521), which may include, for example, X, Y and reference image components. Motion compensation may also include interpolation of sample values ​​extracted from the reference image memory (557) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.

[0068] The output samples of the aggregator (555) can be employed by various loop filtering techniques in the loop filter unit (556). The video compression technique may include an in-loop filtering technique controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and these parameters can be used as symbols (521) from the parser (520) in the loop filter unit (556), but may also be responsive to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0069] The output of the loop filter unit (556) can be a sample stream, which can be output to a display device (512) and stored in a reference image memory (557) for subsequent inter-frame image prediction.

[0070] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (520)) are identified as reference images, the current image buffer (558) can become part of the reference image memory (557), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.

[0071] The video decoder (510) can perform decoding operations according to a predetermined video compression technique as specified in standards such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference image size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.

[0072] In this embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be a portion of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant stripes, redundant images, forward error correction codes, etc.

[0073] Figure 6 This is a block diagram of a video encoder (603) according to an embodiment disclosed in this application. The video encoder (603) is disposed in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used to replace... Figure 4 The video encoder (403) in the example.

[0074] The video encoder (603) can obtain data from the video source (601) (not) Figure 6 In one example, a portion of the electronic device (620) receives a video sample, the video source being capable of capturing video images to be encoded by a video encoder (603). In another embodiment, the video source (601) is a portion of the electronic device (620).

[0075] A video source (601) can provide a sequence of source video samples encoded by a video encoder (603) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb4:2:0, YCrCb4:4:4). In a media service system, the video source (601) can be a storage device storing previously prepared video. In a video conferencing system, the video source (601) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0076] According to an embodiment, the video encoder (603) can encode and compress images of a source video sequence into an encoded video sequence (643) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (650) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be used with other suitable functions related to the video encoder (603) optimized for a particular system design.

[0077] In some embodiments, the video encoder (603) operates in an encoding loop. As a simplified description, in one example, the encoding loop may include a source encoder (630) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols in a manner similar to how the (remote) decoder creates sample data, to create sample data (because in the video compression techniques considered in the subject matter disclosed in this application, any compression between the symbols and the encoded video stream is lossless). The reconstructed sample stream (sample data) is input to a reference image memory (634). Since decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image memory (634) are also bit-precisely corresponding between the local encoder and the remote encoder. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values ​​that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0078] The operation of the “local” decoder (633) can be combined with, for example, the above-mentioned Figure 5 The video decoder (610) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 5 When symbols are available and the entropy encoder (645) and parser (520) are able to encode / decode the symbols into the encoded / decoded video sequence without loss, the entropy decoding portion of the video decoder (510), including the buffer (515) and parser (520), may not be fully implemented in the local decoder (633).

[0079] It can be observed that any decoder technique other than parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in essentially the same functional form. For this reason, this application focuses on decoder operation. The description of encoder techniques can be simplified because encoder techniques are inverses of the fully described decoder techniques. More detailed descriptions are only required in certain areas, and are provided below.

[0080] During operation, in some examples, the source encoder (630) may perform motion-compensated predictive coding, predictively encoding the input image with reference to one or more previously encoded images from the video sequence designated as "reference images." In this manner, the encoding engine (632) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.

[0081] The local video decoder (633) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (630). The operation of the encoding engine (632) can be a lossy process. When the encoded video data can be decoded by the video decoder (633), Figure 6 During decoding (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.

[0082] The predictor (635) can perform a prediction search against the encoding engine (632). That is, for a new image to be encoded, the predictor (635) can search in the reference image memory (634) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (635) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (635), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (634).

[0083] The controller (650) can manage the encoding operations of the source encoder (630), including, for example, setting parameters and subgroup parameters for encoding video data.

[0084] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (645). The entropy encoder (645) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.

[0085] The transmitter (640) can buffer the encoded video sequence created by the entropy encoder (645) in preparation for transmission via a communication channel (660), which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter (640) can combine the encoded video data from the video encoder (603) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0086] The controller (650) manages the operation of the video encoder (603). During encoding, the controller (650) can assign a specific encoded image type to each encoded image, which may affect the encoding techniques applicable to the corresponding image. For example, images can typically be assigned to any of the following image types:

[0087] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.

[0088] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and a reference index to predict sample values ​​for each block.

[0089] A bidirectional predictive image (B-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values ​​for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.

[0090] Source images are typically spatially subdivided into multiple sample blocks (e.g., each with 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.

[0091] The video encoder (603) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (603) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0092] In this embodiment, the transmitter (640) may transmit additional data and encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant images and stripes, SEI messages, VUI parameter set fragments, etc.

[0093] The acquired video can serve as multiple source images (video images) presented in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.

[0094] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, for example, a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.

[0095] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.

[0096] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, for example, 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, for example, inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0097] Figure 7 A diagram of a video encoder (703) according to another embodiment disclosed in this application is shown. The video encoder (703) is used to receive processing blocks (e.g., prediction blocks) of sample values ​​within a current video image in a video image sequence, and to encode said processing blocks into an encoded image that is part of an encoded video sequence. In this embodiment, the video encoder (703) is used instead of Figure 4The video encoder (403) in the example.

[0098] In the HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8×8 sample prediction block. The video encoder (703) uses, for example, rate-distortion (RD) optimization to determine whether to use intra-frame mode, inter-frame mode, or bidirectional prediction mode to encode the processing block. When encoding the processing block in intra-frame mode, the video encoder (703) can use intra-frame prediction techniques to encode the processing block into an already encoded picture; and when encoding the processing block in inter-frame mode or bidirectional prediction mode, the video encoder (703) can use inter-frame prediction or bidirectional prediction techniques to encode the processing block into an already encoded picture, respectively. In some video coding techniques, the merging mode can be an inter-frame picture prediction sub-mode, in which motion vectors are derived from one or more motion vector prediction values ​​without relying on already encoded motion vector components outside the prediction values. In some other video coding techniques, motion vector components applicable to the subject block may exist. In embodiments, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the processing block mode.

[0099] exist Figure 7 In the example, the video encoder (703) includes, for example, Figure 7 The inter-frame encoder (730), intra-frame encoder (722), residual calculator (723), switch (726), residual encoder (724), general controller (721) and entropy encoder (725) are shown coupled together.

[0100] An inter-frame encoder (730) is configured to receive samples of the current block (e.g., the processing block), compare the block with one or more reference blocks in a reference image (e.g., blocks in previous and later images), generate inter-frame prediction information (e.g., redundancy descriptions based on inter-frame coding techniques, motion vectors, merging mode information), and compute inter-frame prediction results (e.g., predicted blocks) based on the inter-frame prediction information using any suitable technique. In some examples, the reference image is a decoded reference image based on encoded video information.

[0101] The intra encoder (722) is used to receive samples of the current block (e.g., the processing block), in some cases compare the block with an encoded block in the same image, generate quantization coefficients after transformation, and in some cases also (e.g., based on intra prediction direction information of one or more intra coding techniques) generate intra prediction information. In one example, the intra encoder (722) also calculates an intra prediction result (e.g., a predicted block) based on the intra prediction information and a reference block in the same image.

[0102] A general controller (721) determines general control data and controls other components of the video encoder (703) based on the general control data. In an embodiment, the general controller (721) determines the mode of a block and provides control signals to a switch (726) based on the mode. For example, when the mode is an intra-frame mode, the general controller (721) controls the switch (726) to select an intra-frame mode result for use by the residual calculator (723) and controls the entropy encoder (725) to select intra-frame prediction information and add the intra-frame prediction information to the bitstream; and when the mode is an inter-frame mode, the general controller (721) controls the switch (726) to select an inter-frame prediction result for use by the residual calculator (723) and controls the entropy encoder (725) to select inter-frame prediction information and add the inter-frame prediction information to the bitstream.

[0103] A residual calculator (723) is used to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). A residual encoder (724) is used to operate on the residual data to encode the residual data to generate transform coefficients. In an embodiment, the residual encoder (724) is used to transform the residual data from the time domain to the frequency domain and generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is used to perform an inverse transform and generate decoded residual data. The decoded residual data can be used appropriately by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and inter-frame prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and intra-frame prediction information. The decoded blocks are processed appropriately to generate a decoded image, and in some embodiments, the decoded image may be buffered in a memory circuit (not shown) and used as a reference image.

[0104] An entropy encoder (725) is used to format the bitstream to produce encoded blocks. The entropy encoder (725) generates various information according to a suitable standard such as the HEVC standard. In an embodiment, the entropy encoder (725) is used to obtain general control data, selected prediction information (e.g., intra-frame prediction information or inter-frame prediction information), residual information, and other suitable information in the bitstream. It should be noted that, according to the disclosed subject matter, residual information is not present when blocks are encoded in a merged sub-mode of inter-frame mode or bidirectional prediction mode.

[0105] Figure 8A diagram of a video decoder (810) according to another embodiment disclosed in this application is shown. The video decoder (810) is used to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In one example, the video decoder (810) is used instead of Figure 4 The video decoder (410) in the example.

[0106] exist Figure 8 In the example, the video decoder (810) includes, for example, Figure 8 The entropy decoder (871), inter-frame decoder (880), residual decoder (873), reconstruction module (874), and intra-frame decoder (872) are shown coupled together.

[0107] An entropy decoder (871) can be used to reconstruct certain symbols from an encoded image, representing the syntax elements constituting the encoded image. Such symbols may include, for example, the block-encoded mode (e.g., intra-frame mode, inter-frame mode, bidirectional prediction mode, a combined sub-mode of the latter two, or another sub-mode), prediction information (e.g., intra-frame prediction information or inter-frame prediction information) that can be identified for prediction by either the intra-frame decoder (872) or the inter-frame decoder (880), residual information in the form of, for example, quantized transform coefficients, and so on. In one example, when the prediction mode is an inter-frame prediction mode or a bidirectional prediction mode, inter-frame prediction information is provided to the inter-frame decoder (880); and when the prediction type is an intra-frame prediction type, intra-frame prediction information is provided to the intra-frame decoder (872). Residual information may be provided to the residual decoder (873) via inverse quantization.

[0108] The inter-frame decoder (880) is used to receive inter-frame prediction information and generate inter-frame prediction results based on the inter-frame prediction information.

[0109] The intra-frame decoder (872) is used to receive intra-frame prediction information and generate prediction results based on the intra-frame prediction information.

[0110] The residual decoder (873) performs inverse quantization to extract the dequantized transform coefficients and processes the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require some control information (to obtain the quantizer parameters QP), and this information may be provided by the entropy decoder (871) (the data path is not indicated because this may only be low-level control information).

[0111] The reconstruction module (874) is used to combine the residual output by the residual decoder (873) with the prediction result (which may be output by the inter-frame prediction module or the intra-frame prediction module) in the spatial domain to form a reconstructed block, which may be a part of a reconstructed image, which in turn may be a part of a reconstructed video. It should be noted that other suitable operations, such as deblocking, may be performed to improve visual quality.

[0112] It should be noted that any suitable technology can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In one embodiment, one or more integrated circuits can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810). In another embodiment, one or more processors executing software instructions can be used to implement the video encoder (403), video encoder (603), and video encoder (703), as well as the video decoder (410), video decoder (510), and video decoder (810).

[0113] This disclosure includes embodiments relating to adaptive upsampling filters for the luma and chroma components of the current block when Reference Image Resampling (RPR) is enabled. The adaptive upsampling filter eliminates the two-stage filtering operation that includes upsampling filtering and adaptive loop filtering. Furthermore, a cross-component filter can be applied within the adaptive upsampling filter for the chroma channel to further improve the chroma quality between the upsampled full-size image and the original full-size image.

[0114] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) released the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (Revision 1), 2014 (Revision 2), 2015 (Revision 3), and 2016 (Revision 4). In 2015, the two standards organizations jointly established the JVET (Joint Video Exploration Team) to explore the potential of developing the next video coding standard beyond HEVC. In October 2017, the two standards organizations issued a joint call for proposals (CfP) for video compression capabilities beyond HEVC. As of February 15, 2018, 22 CfP responses were submitted regarding Standard Dynamic Range (SDR), 12 regarding High Dynamic Range (HDR), and 12 regarding 360 video categories. In April 2018, all received CfP responses were evaluated at the 122nd MPEG meeting / 10th JVET meeting. As a result of this meeting, JVET officially launched the standardization process for next-generation video coding beyond HEVC, naming the new standard Multifunctional Video Coding (VVC), and JVET was renamed the Joint Video Experts Group. In 2020, ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) released the VVC video coding standard (version 1).

[0115] For example, in VVC, the classic block-based hybrid video coding architecture can be applied. Within this architecture, new tools can be incorporated into each basic building block to further improve compression.

[0116] In VVC, the Quadtree Plus Multi-Type Tree (QT+MTT) scheme can use quad partitions followed by binary and ternary partitions as the partitioning structure, instead of the quadtree with multiple partitioning types used in HEVC. Furthermore, the separate partitioning tree structure can support the luma and chroma channels separately. For inter-frame frames, the luma and chroma channels within a single CTU can share the same coding tree structure. However, for intra-frame frames, the luma and chroma channels can have separate trees to improve the coding efficiency of the chroma channel.

[0117] In inter-frame prediction, for each inter-frame prediction coding unit (CU), motion parameters can be used to generate inter-frame prediction samples. Motion parameters may include motion vectors, reference picture indices, reference picture list usage indices, and / or additional information required for new coding features of the VVC. Motion parameters can be explicitly or implicitly signaled. When a CU is coded in skip mode, the CU can be associated with a PU and may not require significant residual coefficients, encoded motion vector increments, and / or reference picture indices. When a CU is coded in merge mode, the motion parameters of the CU can be obtained from neighboring CUs. Neighboring CUs may include spatial and temporal candidates, as well as additional scheduling (or additional candidates) introduced, for example, in the VVC. Merging mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to merging mode is explicit transmission of motion parameters, where, for each CU, motion vectors, the corresponding reference picture index for each reference picture list, reference picture list usage flags, and / or other required information can be explicitly signaled.

[0118] In VVC, the VVC Test Model (VTM) reference software may include some new and improved inter-frame predictive coding tools, which may include one or more of the following:

[0119] (1) Extended merge forecast

[0120] (2) Combined Motion Vector Difference (MMVD)

[0121] (3) AMVP mode with symmetric MVD signaling

[0122] (4) Affine Motion Compensation Prediction

[0123] (5) Sub-block-based temporal motion vector prediction (SbTMVP)

[0124] (6) Adaptive Motion Vector Resolution (AMVR)

[0125] (7) Sports field storage: 1 / 16 th Luminance sample MV storage and 8×8 motion field compression

[0126] (8) Bidirectional prediction with CU-level weights (BCW)

[0127] (9) Bidirectional optical flow (BDOF)

[0128] (10) Decoder-side motion vector correction (DMVR)

[0129] (11) Combined Inter-Frame and Intra-Frame Prediction (CIIP)

[0130] (12) Geometric Partitioning (GPM)

[0131] In intra-frame prediction, for each intra-frame prediction CU, the samples of the intra-frame CU can be predicted based on reference samples in the neighboring blocks to the left and / or above the current CU. The neighboring blocks can be those that have already been decoded within the same image before intra-loop filtering. HEVC can have 35 intra-frame prediction modes, including a planar mode, a reference sample averaging mode (also known as DC mode), and 33 orientation angle modes. For example, in VVC, the intra-frame prediction modes can be extended as follows:

[0132] (1) Including 93 intra-image orientation prediction angles in the Wide Angle Intra-Frame Prediction (WAIP) mode

[0133] (2) Two sets of 4-tap interpolation filters

[0134] (3) Location-related prediction combination (PDPC)

[0135] (4) Multiple Reference Rows (MRL)

[0136] (5) Cross-component linear model (CCLM)

[0137] (6) Intra-Frame Sub-Partition (ISP)

[0138] To achieve better energy compression of residual data and further reduce quantization errors of transform coefficients, new tools can be introduced, for example, in VVC, as follows:

[0139] (1) Non-square transformation

[0140] (2) Multiple transformation selection (MTS) including explicit MTS and implicit MTS

[0141] (3) Low-frequency non-separable transform (LFNST)

[0142] (4) Subblock Transformation (SBT)

[0143] (5) Relevant Quantification (DQ)

[0144] (6) Joint Coding of Chromaticity Residues (JCCR)

[0145] In VVC, remapping operations and three in-loop filters can be applied sequentially to the reconstructed frame to eliminate different types of artifacts. For example, a new sample-based process, Luminance Mapping with Chroma Scaling (LMCS), can be performed first. Then, a deblocking filter can be used to reduce block artifacts. A Sample Adaptive Offset (SAO) filter can then be applied to the deblocked image to attenuate ringing and banding artifacts. Finally, an ALF can be applied to reduce other potential distortions introduced by the transform and quantization processes. The ALF in VVC can include two operations. The first operation can be based on an ALF adapted to block-based filters for both luminance and chroma samples, and the second operation can be based on a CC-ALF that applies only to chroma samples.

[0146] Filter shapes can be applied to block-based ALF. In the example, two filter shapes can be applied. The first filter shape can be a 7×7 rhombus applied to the luma component, and the second filter shape can be a 5×5 rhombus applied to the chroma component. One of up to 25 filters for each block (e.g., a 4×4 block) can be selected based on the directionality and activity of the local gradients of the individual blocks (e.g., a 4×4 block). Based on the directionality and activity of the local gradients, each block (e.g., a 4×4 block) can be classified and categorized into one of 25 classes. Each class can have a corresponding filter coefficient assignment. Before filtering, geometric transformations such as 90-degree rotation and diagonal or vertical flips can be applied to the filter shape (e.g., a 7×7 rhombus or a 5×5 rhombus) based on the gradient values ​​computed for the 4×4 block. Applying geometric transformations before performing the filtering process is equivalent to applying geometric transformations to samples in the filter support region. By aligning the block's directionality, ALF can be performed on the block in a more similar (or consistent) manner.

[0147] In addition to 4×4 block-level filter adaptation for luma, ALF also supports CTU-level filter adaptation. Each CTU can use a filter set computed from the current slice, a filter set signaled at the encoded slice, or a filter set from a set of 16 offline-trained filters. Within each CTU, the selected filter set can be applied to each 4×4 block of the CTU. Filter coefficients and clipping indices can be carried in the ALF Adaptation Parameter Set (APS). The ALF APS can include up to eight chroma filters and a luma filter set that can include up to 25 filters. The ALF APS can also include an index i for each of the 25 luma categories. c By merging different categories, the number of bits associated with the filter coefficients can be reduced.

[0148] CC-ALF can use luminance sample values ​​to correct chrominance sample values ​​during the ALF process. Figure 9 An exemplary combination of ALF and CC-ALF in VVC can be illustrated. Figure 9 As shown, the linear filtering operation can use the luminance sample (902) of the current block as input to CC-ALF filters (904) and (906) to generate correction values ​​(e.g., ΔR) for the chrominance sample values ​​(908) and (910). cb and ΔR cr The luminance sample (902) can be filtered based on SAO (918) to reduce ringing and banding artifacts. The chrominance sample value (908) can be generated by ALF (912) based on the Cb component (914) of the chrominance sample of the current block, as input. The chrominance sample value (910) can be generated by ALF (912) based on the Cr component (916) of the chrominance sample of the current block, as input. For each chrominance component (e.g., (914) and (916)), corrections can be generated independently. After correction, the chrominance sample value (908) can be compared with the correction value ΔR. cb The sum of these values ​​generates the Cb component (920) of the reconstructed sample for the current block, which can be based on the chromaticity sample value (910) and the correction value ΔR. cr The sum of these components generates the Cr component (922) of the reconstructed sample of the current block. Furthermore, when the luminance sample (902) is filtered by ALF (926), the Y (or luminance) component of the reconstructed sample of the current block can be generated.

[0149] Reference pixel (or image) resampling (RPR) allows for resolution changes without encoding the intra-frame random access point (IRAP) image. RPR can also be used in applications requiring scaling of the entire video area or a region of interest. By using RPR, image resolution changes are allowed, enabling inter-image prediction based on a reference image with a different resolution than the current image to be decoded. When the reference image's resolution differs from the current image's resolution, RPR can require resampling of the reference image used for inter-image prediction.

[0150] The scaling factor of the RPR can be limited. For example, for scaling from the reference image to the current image, the scaling factor can be limited from 1 / 2 (or downsampling with a factor of 2) to 8 (or upsampling with a factor of 8). Three sets of resampling filters with different cutoff frequencies can be specified to handle various scaling factors between the reference and current images. The three sets of resampling filters can be applied to scaling factors ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Furthermore, for each set of resampling filters, 16 phases can be applied to luma and 32 phases to chroma. Each phase can correspond to the corresponding sampling rate of the RPR. The number of phases in the resampling filters can be the same as the number of phases in the interpolation filters used in motion compensation.

[0151] ALF (Advanced Programming Filtering) can be used to post-filter reconstructed images (e.g., in VVC) to reduce distortion between the reconstructed and original images. ALF can reduce distortion between low-resolution reconstructed images and low-resolution (or downsampled) original images. When RPR (Reconstructed Persistent Range) is enabled, an upsampling process can be applied to generate an upsampled reconstructed image. Therefore, an additional ALF process stage can be applied to reduce distortion between the upsampled reconstructed image and the original full-size image.

[0152] Figure 10A A first exemplary embodiment of applying ALF to RPR is shown. Figure 10A As shown, the core decoding process (1002) can generate reconstructed samples of the current image. At (1004), the reconstructed samples of the current image can be filtered by post-loop in-circuit filters. Post-loop in-circuit filters can include a deblocking filter (DBF), SAO, ALF-1, and / or CC-ALF-1. DBF can improve visual quality and prediction performance by smoothing sharp edges between blocks. SAO can attenuate ringing and banding artifacts. ALF-1 and CC-ALF-1 can be applied to reduce other potential distortions introduced by the transform and quantization processes. At (1006), upsampling can be performed based on RPR to generate a high-resolution reference image for inter-frame prediction. Accordingly, an upsampled reconstructed image can be generated. At (1008), ALF-2 and CC-ALF-2 can be applied to reduce the distortion between the upsampled reconstructed image and the original full-size image. It should be noted that the reconstructed sample of the current image filtered at (1004) can be stored as a reference image for future images in the decoded image buffer (DPB) (1010).

[0153] Figure 10B A second exemplary embodiment of applying ALF to RPR is shown. Figure 10BAs shown, the core decoding process (1012) can generate reconstructed samples of the current image. At (1014), the reconstructed samples of the current image can be filtered by a deblocking filter (DBF), SAO, ALF-1, and CC-ALF-1. At (1016), ALF-1 can be applied to reduce the distortion between the reconstructed image and the original full-size image. At (1018), upsampling can be performed to generate a high-resolution reference image for inter-frame prediction. Accordingly, an upsampled reconstructed image can be generated. At (1020), CC-ALF-2 can be applied to reduce the distortion between the upsampled reconstructed image and the original full-size image. The reconstructed samples of the current image filtered at (1014) can be stored as reference images for future images in the decoded image buffer (DPB) (1022).

[0154] exist Figure 10A and Figure 10B In this context, when RPR is enabled, an ALF post-filter stage can be applied to the reconstructed upsampled image. Two filters, such as an upsampling filter (e.g., (1006)) and an ALF / CC-ALF (e.g., (1008)), can be cascaded in the ALF post-filter stage to enhance the reconstructed image when RPR is enabled. Due to the use of a series of cascaded filters, the image is increased... Figure 10A and Figure 10B The delay and design complexity of the ALF post-filter stage in the embodiment shown in the figure.

[0155] In this disclosure, an adaptive upsampling filter for luminance and chrominance can be used instead of a cascaded filter, such as a cascaded filter having an upsampling filter, an ALF, and a CC-ALF. The adaptive upsampling filter can combine (or integrate) the functions of a cascaded filter, such as one or more of an upsampling filter, an ALF, and a CC-ALF. Accordingly, a cascaded structure may not be necessary in the adaptive upsampling filter.

[0156] In some embodiments, the adaptive upsampling filter may include a luminance adaptive upsampling filter and a chrominance adaptive upsampling filter, wherein the luminance adaptive upsampling filter may use a low-resolution luminance signal as input and a high-resolution luminance signal as output, and the chrominance adaptive upsampling filter may use both a low-resolution luminance signal and a low-resolution chrominance signal as input and a high-resolution chrominance signal as output.

[0157] An adaptive upsampling filter can include adaptive filter coefficients to implement an upscaling (or upsampling) process, reducing distortion between the upsampled reconstructed image and the original full-size image when RPR is enabled. The filter coefficients of the adaptive upsampling filter can be adapted to the class of the CU (or sub-blocks) and the phase of the RPR. Therefore, the corresponding filter coefficients can be applied to each CU based on its corresponding class and phase.

[0158] Based on the directionality and activity of the local gradients within a small patch (e.g., a 4×4 patch) used for the luminance channel (or luminance samples), each patch within the CU of the current image of a low-resolution image can be classified into one of N categories. In some embodiments, determining the category of a patch (or sub-patch) of the CU used for the adaptive upsampling filter can be similar to determining the category of a patch (or sub-patch) of the CU used for the ALF. Thus, for example, N could be 25 for the luminance samples of the current image.

[0159] For a luminance sample of the current image, multiple phases, such as M phases, can be applied to an adaptive upsampling filter to achieve the resampling function of RPR. Each of these phases can correspond to a specific sampling rate of RPR. Accordingly, an M-phase filter coefficient set can be introduced for the adaptive upsampling filter to satisfy all candidate phases. Therefore, for the adaptive upsampling filter, up to M×N filter coefficient sets can be introduced based on N categories with M phases for the luminance samples of the entire image (or the entire current image). For example, each filter coefficient set can include corresponding filter coefficients and a corresponding limiting value index. When RPR is enabled, the filter coefficient set for the adaptive upsampling filter used for the luminance component can be signaled in the bitstream used for the entire encoded image (or the encoded current image).

[0160] The adaptive filter coefficients of the adaptive upsampling filter can also be applied to the chroma components of the current image. For example, when applying RPR, L phases can be applied to the adaptive upsampling filter. Accordingly, when RPR is enabled, up to L sets of filter coefficients can be signaled throughout the bitstream used for the entire encoded image.

[0161] The directionality and activity of the local gradients for each block used in ALF (e.g., in VVC) can be directly applied to classify each block of the adaptive upsampling filter. Based on the directionality and activity of the local gradients for each block, a corresponding classification index (or class index) can be obtained. The class corresponding to each block can be determined based on the classification index. For example, the class corresponding to each block could be one of 25 candidate classes used in ALF in VVC.

[0162] The use of adaptive upsampling filter coefficients can be signaled implicitly or explicitly. In some embodiments, a flag such as an adaptive upsampling flag may be signaled to indicate whether adaptive upsampling filter coefficients are used. If the flag is false, a default set of filter coefficients can be used for the current picture.

[0163] In some embodiments, n sets of filter coefficients for N categories can be applied to the adaptive upsampling filter by using a merging method, where n<N. Therefore, for each phase, only n filter coefficients can be used. Accordingly, for the upsampling filter, only M×n sets of filter coefficients can be signaled in the code stream.

[0164] In some embodiments, sets of filter coefficients of different phases can be combined (or merged). In an example, two sets of filter coefficients can be combined (or merged) by using the merging method. The sets of filter coefficients can be arbitrarily selected. The merging method can select one of the two sets of filter coefficients based on a cost value. For example, a first cost value of a first set of filter coefficients can be determined, where the first cost value can indicate distortion associated with the current block and reconstructed samples of the current block filtered based on the first set of filter coefficients, for example, rate distortion or difference. A second cost value of a second set of filter coefficients can be determined, where the second cost value can indicate distortion associated with the current block and reconstructed samples of the current block filtered based on the second set of filter coefficients. One of the first set of filter coefficients and the second set of filter coefficients can be selected based on the smaller one of the first cost value and the second cost value. After a recursive merging operation is performed on the two filters by using a rate-distortion optimization method at an encoder side, some sets of filter coefficients can be selected from N sets of filter coefficients, where N can correspond to the number of categories associated with the current picture. Finally, the optimal (or selected) sets of coefficients can be signaled in one or more code streams. In addition, a mapping table can also be signaled. The mapping table can include a plurality of filter coefficient set indexes. Each filter coefficient set index can correspond to a corresponding category.

[0165] In another example, two sets of filter coefficients having two adjacent category indexes can be combined. For example, after a recursive merging operation using rate-distortion optimization at the encoder side, the set of filter coefficients for category 0 can also be used for category 1, category 2, category 3 and category 4, where a cost value of the set of filter coefficients for category 0 can be smaller than cost values of the sets of filter coefficients for category 2, category 3 and category 4. After the recursive merging operation is performed on the sets of filter coefficients of all category indexes, only n sets of filter coefficients can be signaled in one or more code streams, where n is smaller than the number N of categories.

[0166] In another example, two sets of filter coefficients having two different phases may be further combined. After performing the filter coefficient merging operation, only m×n sets of filter coefficients may be signaled in one or more code streams, where m<M, n<N, and M and N are the number of phases and classes respectively. Based on the merging operation, two sets of filter coefficients having two arbitrary phases can be combined by using rate-distortion optimization. The reduced sets of filter coefficients and a mapping table corresponding to phases and filter coefficient set indices may be signaled. In some embodiments, only sets of filter coefficients of adjacent (or consecutive) phases (e.g., two adjacent phases) may be merged. For example, one set of filter coefficients among n sets of filter coefficients may be applied to both phase 1 and phase 2.

[0167] In some embodiments, a low-resolution encoded picture may be divided into a plurality of sub-regions. Each sub-region may have a geometric shape, for example, a rectangular shape. In an example, the encoded picture may be partitioned into P rectangular regions. For each sub-region (e.g., rectangular region) within the encoded picture, different sets of filter coefficients may be used, for example, up to N sets of filter coefficients for N classes. Therefore, a total of P×N sets of filter coefficients of the adaptive upsampling filter may be applied to the P rectangular regions partitioned from the encoded picture.

[0168] Sets of filter coefficients corresponding to N classes and P sub-regions may be merged. The set of filter coefficients for a rectangular region i may also be used for other rectangular regions, for example, a rectangular region j, where i≠j. For each rectangular region, an index of the set of filter coefficients may be signaled to indicate which set of filter coefficients is used for the corresponding rectangular region.

[0169] In some embodiments, the region-based approach may be adopted for luma and chroma respectively. For example, the region-based approach may be used only for the luma channel.

[0170] When RPR is enabled, the sets of filter coefficients of the adaptive upsampling filter may be signaled in an adaptive parameter set (APS). The sets of filter coefficients of the adaptive upsampling filter in the APS may be used for further encoded pictures with RPR. In some embodiments, the sets of filter coefficients of the adaptive upsampling filter may be signaled only when the current encoded picture is used as a reference picture for future pictures.

[0171] In the present disclosure, in order to eliminate the Figure 10A and Figure 10B two-stage filtering process for RPR shown in , a single adaptive upsampling filter is provided to replace the upsampling filter coupled with ALF, so as to enhance the upsampled reconstructed full-size picture for both the luma channel and the chroma channel when RPR is enabled. Figure 11 An exemplary adaptive upsampling filter (1100) is shown. By using the adaptive upsampling filter (1100) on both the luminance and chrominance channels, the intensity of the chrominance can be reduced. Figure 10A and Figure 10B The cascaded filter shown in the diagram reduces the delay and complexity, for example, reducing the process from two stages to one stage. Figure 11 As shown, the adaptive upsampling filter (1100) may include a luminance adaptive upsampling filter (1106) and a chrominance adaptive upsampling filter (1108). The core decoding process (1102) can generate reconstructed samples of the current image. At (1104), the reconstructed samples of the current image can be filtered by a deblocking filter (DBF), SAO, ALF-1, and CC-ALF-1. The lower-resolution luminance component of the reconstructed samples of the current block can be transferred to the luminance adaptive upsampling filter (1106) to generate a high-resolution luminance component of the filtered reconstructed samples of the current block. The lower-resolution luminance component and the lower-resolution chrominance component of the reconstructed samples of the current block can also be transferred to the chrominance adaptive upsampling filter (1108) to generate a high-resolution chrominance component of the filtered reconstructed samples of the current block.

[0172] To adapt the filter coefficients of the adaptive upsampling filter (1100) to the brightness channel, sub-block-level filter adaptation can be achieved using classification based on the directionality and activity of the local gradients of the sub-blocks. For example, each CU can be divided into multiple sub-blocks. The size of each sub-block can be any number, for example, 4×4. For each sub-block, it can be classified based on a corresponding category index. The category index can be derived based on the directionality of the local gradients of the sub-block and 2D Laplacian activity. Each sub-block can be classified into one of the predefined categories based on the category index. Based on the determined category and sampling rate (or phase of RPR), the corresponding set of filter coefficients for each sub-block can be determined. The coefficient set can be applied to the sub-block to perform the filtering process. Accordingly, filtered reconstructed samples can be generated. The typical number of predefined categories can be 25, which can be the same as the number of categories in an ALF, for example, in a VVC. Based on sub-block-level adaptation, adaptive upsampling filters with different sub-blocks having different local gradient characteristics can be implemented.

[0173] Furthermore, due to the high correlation between the luminance and chrominance channels, luminance samples from low-resolution images can also be used to compensate for chrominance distortion between the upsampled image and the original image. For example, ... Figure 11As shown, both the lower-resolution luminance component and the lower-resolution chrominance component of the reconstructed sample of the current block can be transmitted to the chrominance adaptive upsampling filter (1108) to generate the high-resolution chrominance component of the filtered reconstructed sample of the current block. Accordingly, a cross-component filter can be formed in the chrominance adaptive upsampling filter (1108). The cross-component filter can directly obtain high-frequency information from the low-resolution image as a compensation offset value for the upsampled chrominance sample.

[0174] Figure 12 A flowchart outlining an exemplary decoding process (1200) according to some embodiments of the present disclosure is shown. Figure 13 A flowchart outlining an exemplary encoding process (1300) according to some embodiments of this disclosure is shown. The proposed processes can be used individually or in combination in any order. Furthermore, each of the processes (or embodiments), encoders, and decoders can be implemented by a processing circuit system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-volatile computer-readable medium.

[0175] The operations of the processes (e.g., (1200) and (1300)) can be combined or arranged in any number or order as needed. In an embodiment, two or more of the operations of the processes (e.g., (1200) and (1300)) can be performed in parallel.

[0176] The processes (e.g., (1200) and (1300)) can be used to reconstruct and / or encode blocks to generate predicted blocks for the blocks in the reconstruction. In various embodiments, the processes (e.g., (1200) and (1300)) are executed by processing circuitry, such as processing circuitry in terminal devices (310), (320), (330), and (340); processing circuitry that performs the functions of a video encoder (403); processing circuitry that performs the functions of a video decoder (410); processing circuitry that performs the functions of a video decoder (510); processing circuitry that performs the functions of a video encoder (603), etc. In some embodiments, the processes (e.g., (1200) and (1300)) are implemented as software instructions, so that the processing circuitry executes the processes (e.g., (1200) and (1300)) when the processing circuitry executes the software instructions.

[0177] like Figure 12 As shown, process (1200) can start from (S1201) and proceed to (S1210). At (S1210), the encoded information of the current block in the current picture can be received from the encoded video bitstream. The encoded information can indicate that an adaptive upsampling filter should be applied to the current block.

[0178] At (S1220), the corresponding category of each of the multiple sub-blocks of the current block can be determined.

[0179] At (S1230), a corresponding set of filter coefficients for each of the multiple sub-blocks can be determined from the multiple set of filter coefficients of the adaptive upsampling filter. This corresponding set of filter coefficients can be determined based on at least one category corresponding to each sub-block and the corresponding sampling rate of the reference pixel resampling (RPR) applied to a reference image for the current image. The corresponding sampling rate can be associated with one of the multiple phases of the RPR.

[0180] At (S1240), an adaptive upsampling filter can be applied to the current block to generate filtered reconstructed samples of the current block based on the determined set of corresponding filter coefficients for multiple sub-blocks, without applying a secondary ALF.

[0181] To determine the corresponding category, the category index for each sub-block can be determined based on the directionality and activity of the local gradients of the brightness samples in each sub-block. The category for each sub-block can be determined based on its corresponding category index.

[0182] In some embodiments, the number of multiple filter coefficient sets for the luminance samples of the current block can be equal to the product of N and M. N may be the number of classes associated with the current block, and M may be the number of multiple phases of the RPR associated with the luminance samples of the current block.

[0183] In some embodiments, the number of filter coefficient sets for the chroma samples of the current block can be equal to L, where L can be the number of multiple phases of the RPR associated with the chroma samples of the current block.

[0184] In process (1200), a first cost value for a first filter coefficient set within a plurality of filter coefficient sets can be determined. The first cost value indicates the distortion between the current block and reconstructed samples of the current block filtered based on the first filter coefficient set. A second cost value for a second filter coefficient set within the plurality of filter coefficient sets can be determined. The second cost value indicates the distortion between the current block and reconstructed samples of the current block filtered based on the second filter coefficient set. One of the first and second filter coefficient sets can be selected based on the smaller of the first and second cost values.

[0185] In some embodiments, a first set of filter coefficients may be associated with a first category index, and a second set of filter coefficients may be associated with a second category index. The second category index may be consecutive to the first category index.

[0186] In some embodiments, a first set of filter coefficients may be associated with a first phase among a plurality of phases of the RPR, and a second set of filter coefficients may be associated with a second phase among a plurality of phases of the RPR, wherein the second phase may be continuous with the first phase of the RPR.

[0187] In some embodiments, the number of multiple filter coefficient sets for the brightness samples of the current block can be equal to the product of N and P. N can be the number of categories associated with the current block, and P can be the number of regions partitioned from the current image. Accordingly, a first filter coefficient set can be associated with a first region of the current image, and a second filter coefficient set can be associated with a second region of the current image.

[0188] In process (1200), multiple filter coefficient sets can be included in the adaptive parameter set of the encoded information.

[0189] The adaptive upsampling filter may further include a luminance adaptive upsampling filter and a chrominance adaptive upsampling filter. A high-resolution luminance component of the reconstructed samples of the current block can be generated based on the lower-resolution luminance component of the reconstructed samples of the current block, which serves as input to the luminance adaptive upsampling filter. Similarly, a high-resolution chrominance component of the reconstructed samples of the current block can be generated based on the lower-resolution luminance component and the lower-resolution chrominance component of the reconstructed samples of the current block, which serve as input to the chrominance adaptive upsampling filter.

[0190] Procedure (1200) may be modified as appropriate. One or more steps in procedure (1200) may be improved and / or omitted. One or more additional steps may be added. Any suitable implementation order may be used.

[0191] like Figure 13 As shown, process (1300) can start from (S1301) and proceed to (S1310). At (S1310), the corresponding category of each of the multiple sub-blocks can be determined. The multiple sub-blocks can come from the current block in the current image.

[0192] At (S1320), multiple filter coefficient sets of the adaptive upsampling filter can be determined, corresponding to the category of the current block and multiple phases of the reference pixel resampling (RPR) applied to the resolution change of the current image.

[0193] At (S1330), the corresponding filter coefficient set for each sub-block can be determined from multiple filter coefficient sets based on at least one category corresponding to each sub-block and the corresponding sampling rate of the reference pixel resampling (RPR) applied to the reference image of the current image. The corresponding sampling rate can be associated with one of the multiple phases of the RPR.

[0194] At (S1340), an adaptive upsampling filter can be applied to the current block to generate filtered reconstructed samples of the current block based on the corresponding filter coefficient sets of the determined multiple sub-blocks, without applying a secondary ALF.

[0195] At (S1350), the encoded information for the current block can be generated. The encoded information can indicate multiple filter coefficient sets of the adaptive upsampling filter.

[0196] Then, the process proceeds to (S1399) and terminates.

[0197] The procedure (1300) can be modified as appropriate. One or more steps in the procedure (1300) can be improved and / or omitted. One or more additional steps can be added. Any suitable implementation order can be used.

[0198] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 A computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0199] The computer software can be encoded using any suitable machine code or computer language, which can be subjected to assembly, compilation, linking or similar mechanisms to create code including instructions that can be executed directly or by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., through interpretation, microcode execution, etc.

[0200] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0201] Figure 14 The components shown for the computer system (1400) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing embodiments of this application. Nor should the configuration of the components be construed as having any dependency or requirement on any one or combination of components shown in the exemplary embodiments of the computer system (1400).

[0202] The computer system (1400) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users via, for example, tactile input (e.g., key presses, swipes, data glove movements), audio input (e.g., voice, taps), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0203] The input human-machine interface device may include one or more of the following (only one of each is depicted): keyboard (1401), mouse (1402), trackpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), and camera (1408).

[0204] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback from a touchscreen (1410), a data glove (not shown), or a joystick (1405), but tactile feedback devices that do not act as input devices may also exist), audio output devices (e.g., speakers (1409), headphones (not depicted)), visual output devices (e.g., screens (1410), including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touchscreen input capability, each with or without tactile feedback capability—some of which are capable of outputting two-dimensional or greater than three-dimensional visual outputs in a manner such as stereoscopic flat image output; virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).

[0205] The computer system (1400) may also include human-accessible storage devices and associated media of the storage devices, such as optical media, including CD / DVD ROM / RW (1420) having media such as CD / DVD (1421), thumb drives (1422), removable hard disk drives or solid-state drives (1423), legacy magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated devices based on ROM / Application-Specific Integrated Circuits (ASICs) / Programmable Logic Devices (PLDs), such as security protection devices (not depicted), etc.

[0206] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0207] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network may be, for example, wireless, wired, or optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include, for example, local area networks such as Ethernet and wireless LANs; cellular networks including Global System for Mobile Communications (GSM), 3G, 4G, 5G, and LTE; wired or wireless wide area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicular and industrial networks including Controller Area Network Bus (CANBus), etc. Some networks typically require external network interface adapters attached to certain general-purpose data ports or peripheral buses (1449) (e.g., the Universal Serial Bus (USB) port of the computer system (1400)); others are typically integrated into the core of the computer system (1400) by attaching to system buses as described below (e.g., integrated into a PC computer system via an Ethernet interface, or integrated into a smartphone computer system via a cellular network interface). By using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receiving (e.g., broadcasting TV), one-way transmitting (e.g., a CANBus connected to a CANBus device), or bidirectional, such as connecting to other computer systems using a local area digital network or a wide area digital network. Certain protocols and protocol stacks can be used on each of those networks and network interfaces described above.

[0208] The aforementioned human-machine interface device, human-accessible storage device, and network interface can be attached to the core (1440) of the computer system (1400).

[0209] The core (1440) may include one or more central processing units (CPU) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units (1443) in the form of field-programmable gate areas (FPGAs), hardware accelerators (1444) for certain tasks, graphics adapters (1450), and so on. These devices, along with read-only memory (ROM) (1445), random access memory (1446), internal mass storage devices such as internal non-user-accessible hard disk drives, solid-state drives (SSDs), etc. (1447), may be connected via a system bus (1448). In some computer systems, the system bus (1448) may be accessed via one or more physical connectors to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1449) to the core's system bus (1448). In one example, a screen (1410) may be connected to a graphics adapter (1450). Architectures for the peripheral bus include peripheral device interconnect (PCI), USB, and so on.

[0210] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions, which, when combined, constitute the aforementioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Transient data can also be stored in RAM (1446), while permanent data can be stored, for example, in an internal mass storage device (1447). Fast storage and retrieval of any memory device can be achieved using a cache memory, which can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage devices (1447), ROMs (1445), RAMs (1446), etc.

[0211] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be designed and constructed specifically for the purposes of this application, or may belong to a class well-known and available to those skilled in the art of computer software.

[0212] For example, but not as a limitation, a computer system having an architecture (1400), and particularly a core (1440), can provide functionality resulting from the execution of software embodied in one or more tangible computer-readable media by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media can be media associated with certain storage devices of a non-transitory nature (e.g., internal core storage device (1447) or ROM (1445)) of the user-accessible mass storage devices described above and the core (1440). Software implementing various embodiments of this application can be stored in such devices and executed by the core (1440). Depending on specific needs, the computer-readable media may include one or more memory devices or chips. The software can cause the core (1440), and more specifically, the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to processes defined by the software. Alternatively or as an alternative, the computer system may provide functionality generated by logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1444)), which may operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may cover logic, and vice versa. Where appropriate, references to computer-readable media may cover circuitry storing software for execution (e.g., integrated circuits (ICs)), circuitry embodying logic for execution, or both. This application covers any suitable combination of hardware and software.

[0213] Appendix A: Abbreviations

[0214] JEM: Joint Exploration Model

[0215] VVC: Versatile Video Coding

[0216] BMS: Benchmark Set

[0217] MV: Motion Vector

[0218] HEVC: High Efficiency Video Coding

[0219] SEI: Supplementary Enhancement Information

[0220] VUI: Video Usability Information

[0221] GOP: Groups of Pictures

[0222] TU: Transform Unit

[0223] PU: Prediction Unit

[0224] CTU: Coding Tree Unit

[0225] CTB: Coding Tree Block

[0226] PB: Prediction Block

[0227] HRD: Hypothetical Reference Decoder

[0228] SNR: Signal Noise Ratio

[0229] CPU: Central Processing Unit

[0230] GPU: Graphics Processing Unit

[0231] CRT: Cathode Ray Tube

[0232] LCD: Liquid-Crystal Display

[0233] OLED: Organic Light-Emitting Diode

[0234] CD: Compact Disc

[0235] DVD: Digital Video Disc

[0236] ROM: Read-Only Memory

[0237] RAM: Random Access Memory

[0238] ASIC: Application-Specific Integrated Circuit

[0239] PLD: Programmable Logic Device

[0240] LAN: Local Area Network

[0241] GSM: Global System for Mobile Communications

[0242] LTE: Long-Term Evolution

[0243] CANBus: Controller Area Network Bus

[0244] USB: Universal Serial Bus

[0245] PCI: Peripheral Component Interconnect

[0246] FPGA: Field Programmable Gate Array

[0247] SSD: Solid-state drive

[0248] IC: Integrated Circuit

[0249] CU: Coding Unit

[0250] Although this application describes several exemplary embodiments, various modifications, arrangements, and alternative equivalents are possible within the scope of this application. Therefore, it should be understood that those skilled in the art can design various systems and methods, though not expressly shown or described herein, that embody the principles of this application, within the spirit and scope of the application.

Claims

1. A method of video decoding, the method comprising: The method includes: Receive encoded information of the current block in the current image from the encoded video stream, wherein the encoded information indicates that an adaptive upsampling filter should be applied to the current block; Determine the corresponding category of each of the multiple sub-blocks of the current block; Based on at least one category corresponding to each of the plurality of sub-blocks and the corresponding sampling rate of the reference pixel resampling (RPR) applied to the reference image of the current image, a corresponding filter coefficient set for each sub-block is determined from a plurality of filter coefficient sets of the adaptive upsampling filter, wherein the corresponding sampling rate is associated with one phase of a plurality of phases of the reference pixel resampling (RPR); and The adaptive upsampling filter is applied to the current block to generate filtered reconstructed samples of the current block based on the determined set of corresponding filter coefficients for the plurality of sub-blocks, without applying a secondary adaptive loop filter (ALF) to the reconstructed samples.

2. The method of claim 1, wherein, The determination of the corresponding category further includes: Based on the directionality and activity of the local gradient of the brightness samples in each of the multiple sub-blocks, the corresponding category index of each sub-block is determined, and Based on the category index corresponding to each of the multiple sub-blocks, the corresponding category of each sub-block is determined.

3. The method of claim 1, wherein, The number of multiple filter coefficient sets used for the luminance sample of the current block is equal to the product of N and M, where N is the number of classes associated with the current block and M is the number of multiple phases of the reference pixel resampling (RPR) associated with the luminance sample of the current block.

4. The method of claim 1, wherein, The number of filter coefficient sets used for the chroma samples of the current block is equal to L, where L is the number of multiple phases of the reference pixel resampling (RPR) associated with the chroma samples of the current block.

5. The method of claim 1, wherein, Further includes: A first cost value is determined for a first filter coefficient set in the plurality of filter coefficient sets, the first cost value indicating the distortion between the current block and the reconstructed samples of the current block filtered based on the first filter coefficient set; Determine a second cost value for a second filter coefficient set within the plurality of filter coefficient sets, the second cost value indicating the distortion between the current block and reconstructed samples of the current block filtered based on the second filter coefficient set; as well as Based on the smaller of the first cost value and the second cost value, one of the first filter coefficient set and the second filter coefficient set is selected.

6. The method according to claim 5, characterized in that: The first filter coefficient set is associated with the first category index corresponding to the sub-block. The second filter coefficient set is associated with the second category index corresponding to the sub-block, and the second category index is continuous with the first category index.

7. The method according to claim 5, characterized in that: The first set of filter coefficients is associated with a first phase of a plurality of phases of the reference pixel resampling (RPR), and The second set of filter coefficients is associated with a second phase of a plurality of phases of the reference pixel resampling (RPR), the second phase being continuous with the first phase of the reference pixel resampling (RPR).

8. The method of claim 5, wherein, The number of filter coefficient sets used for the brightness samples of the current block is equal to the product of N and P, where N is the number of categories associated with the current block and P is the number of regions partitioned from the current image.

9. The method according to claim 8, characterized in that: The first set of filter coefficients is associated with a first region of the current image. The second set of filter coefficients is associated with a second region of the current image.

10. The method according to any one of claims 1 to 9, characterized in that, The plurality of filter coefficient sets are included in the adaptive parameter set of the encoded information.

11. The method according to any one of claims 1 to 9, characterized in that: The adaptive upsampling filter further includes a luminance adaptive upsampling filter and a chrominance adaptive upsampling filter; Based on the lower resolution luminance component of the reconstructed sample of the current block as input to the luminance adaptive upsampling filter, a high resolution luminance component of the filtered reconstructed sample of the current block is generated. Based on the lower-resolution luminance component and lower-resolution chrominance component of the reconstructed sample of the current block as input to the chrominance adaptive upsampling filter, a high-resolution chrominance component of the filtered reconstructed sample of the current block is generated.

12. A video decoding apparatus, comprising: include: The processing circuit is configured as follows: Receive encoded information of the current block in the current image from the encoded video stream, wherein the encoded information indicates that an adaptive upsampling filter should be applied to the current block; Determine the corresponding category of each of the multiple sub-blocks of the current block; Based on at least one category corresponding to each of the plurality of sub-blocks and the corresponding sampling rate of the reference pixel resampling (RPR) of the reference image applied to the current image, a corresponding filter coefficient set for each sub-block is determined from a plurality of filter coefficient sets of the adaptive upsampling filter, wherein the corresponding sampling rate is associated with one of a plurality of phases of the reference pixel resampling (RPR). as well as The adaptive upsampling filter is applied to the current block to generate filtered reconstructed samples of the current block based on the determined set of corresponding filter coefficients for the plurality of sub-blocks, without applying a secondary adaptive loop filter (ALF) to the reconstructed samples.

13. The apparatus of claim 12, wherein, The processing circuit is configured as follows: Based on the directionality and activity of the local gradient of the brightness samples in each of the multiple sub-blocks, the corresponding category index of each sub-block is determined, and Based on the category index corresponding to each of the multiple sub-blocks, the corresponding category of each sub-block is determined.

14. The apparatus according to claim 12, characterized in that: The number of filter coefficient sets used for the luminance samples of the current block is equal to the product of N and M, where N is the number of classes associated with the current block, and M is the number of phases of the reference pixel resampling (RPR) associated with the luminance samples of the current block. The number of filter coefficient sets used for the chroma samples of the current block is equal to L, where L is the number of multiple phases of the reference pixel resampling (RPR) associated with the chroma samples of the current block.

15. The apparatus of claim 12, wherein, The processing circuit is configured as follows: A first cost value is determined for a first filter coefficient set in the plurality of filter coefficient sets, the first cost value indicating the distortion between the current block and the reconstructed samples of the current block filtered based on the first filter coefficient set; Determine a second cost value for a second filter coefficient set within the plurality of filter coefficient sets, the second cost value indicating the distortion between the current block and reconstructed samples of the current block filtered based on the second filter coefficient set; as well as Based on the smaller of the first cost value and the second cost value, one of the first filter coefficient set and the second filter coefficient set is selected.

16. The apparatus according to claim 15, characterized in that: The first filter coefficient set is associated with the first category index corresponding to the sub-block. The second filter coefficient set is associated with the second category index corresponding to the sub-block, and the second category index is continuous with the first category index.

17. The apparatus according to claim 15, characterized in that: The first set of filter coefficients is associated with a first phase of a plurality of phases of the reference pixel resampling (RPR). The second set of filter coefficients is associated with a second phase of a plurality of phases of the reference pixel resampling (RPR), the second phase being continuous with the first phase of the reference pixel resampling (RPR).

18. A method of video encoding, comprising: The method includes: Determine the category of each of the multiple sub-blocks of the current block in the current image; Based on the multiple categories of the current block and the multiple phases of the reference pixel resampling (RPR) applied to the resolution change of the current image, a multiple set of filter coefficients for the adaptive upsampling filter is determined. Based on at least one category corresponding to each of the multiple sub-blocks of the current block and the corresponding sampling rate of the reference pixel resampling (RPR) applied to the reference image of the current image, a corresponding filter coefficient set for each sub-block is determined from the multiple filter coefficient sets, wherein the corresponding sampling rate of each sub-block is associated with one of the multiple phases of the reference pixel resampling (RPR). The adaptive upsampling filter is applied to the current block to generate filtered reconstructed samples of the current block based on the corresponding filter coefficient sets of the determined multiple sub-blocks, without applying a secondary adaptive loop filter (ALF) to the reconstructed samples. Generate encoded information for the current block, which indicates a set of multiple filter coefficients for the adaptive upsampling filter.

19. A video encoding device, characterized in that, include: The processing circuit is configured as follows: Receive encoded information of the current block in the current image from the encoded video stream, wherein the encoded information indicates that an adaptive upsampling filter should be applied to the current block; Determine the corresponding category of each of the multiple sub-blocks of the current block; Based on at least one category corresponding to each of the plurality of sub-blocks and the corresponding sampling rate of the reference pixel resampling (RPR) of the reference image applied to the current image, a corresponding filter coefficient set for each sub-block is determined from a plurality of filter coefficient sets of the adaptive upsampling filter, wherein the corresponding sampling rate is associated with one of a plurality of phases of the reference pixel resampling (RPR). as well as The adaptive upsampling filter is applied to the current block to generate filtered reconstructed samples of the current block based on the determined set of corresponding filter coefficients for the plurality of sub-blocks, without applying a secondary adaptive loop filter (ALF) to the reconstructed samples.

20. An electronic device, characterized in that, It includes a memory for storing computer-readable instructions; and a processor for reading the computer-readable instructions and executing the method according to any one of claims 1 to 11, 18.

21. A method for storing video streams, characterized in that, The method of claim 18 is used to generate a video stream and to store the video stream.

22. A method for transmitting a video stream, characterized in that, The method of claim 18 is used to generate a video stream and to transmit the video stream.

23. A computer-readable storage medium storing a computer program / instructions and a video stream thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 18 to generate the video stream.

Citation Information

Patent Citations

  • Signaling filters for video processing

    US20210092458A1

  • Loop filter design for adaptive resolution video coding

    WO2020252745A1