Method and apparatus for video coding

Adaptive reduced-resolution coding techniques enhance video encoding and decoding efficiency by employing block-level flags, scaled motion vectors, and interleaved lists, addressing redundancy and bandwidth challenges in high-resolution video streams.

JP7697641B2Active Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024049106
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-28
Filing Date
2024-03-26
Publication Date
2025-06-24
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies face challenges in efficiently reducing redundancy and bandwidth requirements while maintaining acceptable video quality, particularly in high-resolution video streams, due to limitations in intra prediction modes and motion vector prediction mechanisms.

Method used

The implementation of reduced-resolution coding techniques, including block-level flags for adaptive downsampling and upsampling of reference blocks, scaled motion vectors, and interleaved motion vector candidate lists, to enhance compression efficiency and quality.

Benefits of technology

This approach allows for improved compression ratios and reduced bandwidth usage without significant quality loss, by adaptively applying reduced-resolution coding at the block level, leveraging advanced motion vector scaling and prediction methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697641000010
    Figure 0007697641000010
  • Figure 0007697641000011
    Figure 0007697641000011
  • Figure 0007697641000012
    Figure 0007697641000012
Patent Text Reader

Abstract

To provide a method, an apparatus, and a storage medium for video encoding / decoding with block-level super-resolution coding.SOLUTION: A decoder decodes a video bitstream to obtain a reduced-resolution residual block for a current block, determines a block-level flag that is set to a predefined value indicating that the current block is coded with reduced-resolution coding, generates a reduced-resolution prediction block for the current block by down-sampling a full-resolution reference block for the current block on the basis of the block-level flag, generates a reduced-resolution reconstructed block for the current block on the basis of the reduced-resolution prediction block and the reduced-resolution residual block, and generates a full-resolution reconstructed block for the current block by up-sampling the reduced-resolution reconstructed block.SELECTED DRAWING: Figure 23
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] Incorporation by reference This application claims priority to U.S. Patent Application No. 17 / 488,027, filed September 28, 2021, entitled "Methods and Apparatus for Video Coding", which claims priority to U.S. Provisional Application No. 63 / 137,350, filed January 14, 2021, entitled "Hybrid Resolution Prediction for CU-Based Super-Resolution Coding". The disclosures of the prior applications are hereby incorporated by reference in their entirety.

[0002]

[0002] Technical field This disclosure generally describes embodiments related to video coding.

Background Art

[0003]

[0003] The background description provided herein is for the purpose of generally presenting the context of the disclosure. Work done under the present inventor's name is not admitted as prior art to this disclosure, either expressly or by implication, to the extent that the work is not described in this background section and is not otherwise eligible to be considered prior art at the time of filing in a manner that would render the description of the work as prior art.

[0004]

[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate (informally known as the frame rate), for example 60 pictures per second, i.e., 60 Hz. Uncompressed video has significant bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.

[0005]

[0005] One of the purposes of video encoding and decoding can be said to be the reduction of redundancy in the input video signal by compression. Compression can, in some cases, reduce the aforementioned bandwidth or storage space requirements by a factor of two or more. Both lossless and lossy compression, as well as combinations thereof, can be used. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is widely used. The amount of allowable distortion depends on the application. For example, users of certain consumer streaming applications may be able to tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect the fact that higher allowable / tolerable distortion can result in a higher compression ratio.

[0006]

[0006] Video encoders and decoders can utilize techniques from several broad categories including, for example, motion compensation, transformation, quantization, and entropy coding.

[0007]

[0007] Video codec technology can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference images or other data. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are coded in the intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in a coded video bitstream and video session or as a still image. It is possible to apply a transformation to the samples of an intra block, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique for minimizing the sample values in the domain before transformation. In some cases, the smaller the DC value and AC coefficients after transformation, the fewer the number of bits required to represent the block with a given quantization step size after entropy coding.

[0008]

[0008] Traditional intra coding, such as that known in MPEG-2 generation coding technology, does not use intra prediction. However, some new video compression techniques include techniques that attempt to predict sample values from, for example, neighboring sample data and / or metadata obtained during the encoding and / or decoding of spatially adjacent and previously decoded data blocks. Such techniques are hereinafter referred to as "intra prediction" techniques. It should be noted that in at least some cases, intra prediction uses only reference data from the current picture being reconstructed and not from a reference picture.

[0009]

[0009] There can be various numerous forms of intra prediction. In a given video coding technology, if more than one such technology can be used, the technology used can be coded in an intra prediction mode. In some cases, the mode can have sub - modes and / or parameters, which can be coded individually or included in the mode codeword. The codeword used for a given combination of mode, sub - mode, and / or parameter has an effect on the coding efficiency gain through intra prediction, and the same is true for the entropy coding technology used to convert the codeword into a bitstream.

[0010]

[0010] Certain intra prediction modes were introduced in H.264, improved in H.265, and further improved in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using adjacent sample values belonging to already available samples. The sample values of the adjacent samples are copied into the predictor block according to a certain direction. The reference for the direction in use may be coded in the bitstream or may itself be predicted.

[0011] [

[0011] ] Referring to FIG. 1A, what is shown at the lower right is a subset of 9 predictor directions out of the 33 possible predictor directions of H.265 (corresponding to 33 of the 35 intra - mode angular modes). The point (101) where the arrows converge represents the sample to be predicted. The arrows indicate the direction from which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples at an angle of 45 degrees from the horizontal and towards the upper right. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples at an angle of 22.5 degrees from the horizontal and towards the lower left of sample (101).

[0012] [

[0012] ] Continuing to refer to FIG. 1A, at the upper left, a square block (104) of 4×4 samples is shown (indicated by the thick dashed line). The square block (104) contains 16 samples, each labeled with an "S", its position in the Y dimension (e.g., row index), and its position in the X dimension (e.g., column index). For example, sample S21 is at the second sample in the Y dimension (from the top) and the first sample in the X dimension (from the left). Similarly, sample S44 is the fourth sample of block (104) in both the Y and X dimensions. Since the block size is 4×4 samples, S44 is at the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with an "R", its Y position (e.g., row index) with respect to block (104), and its X position (column index). In both H.264 and H.265, the predicted samples are adjacent to the block during reconstruction; thus, negative values need not be used.

[0013]

[0013] Intra-picture prediction can be processed by copying the reference sample value from adjacent samples as needed according to the signaled prediction direction. For example, assuming that the coded video bitstream includes signaling indicating a prediction direction that matches the arrow (102) for this block, that is, the samples are predicted from one or more prediction samples at an angle of 45 degrees from the horizontal direction and upward to the right. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. And sample S44 is predicted from reference sample R08.

[0014]

[0014] In some cases, especially when the direction cannot be evenly divided at 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation, to calculate a certain reference sample.

[0015]

[0015] The number of possible directions is increasing as video coding technology develops. In H.264 (2003), nine different directions could be represented. This was increased to 33 in H.265 (2013), and JEM / VVC / BMS at the time of this disclosure can support up to 65 directions. Experiments are conducted to identify the most likely directions, and in entropy coding, a certain technique is used to represent the more likely directions with fewer bits and accept a penalty for the less likely directions. Furthermore, the direction itself can often be predicted from the adjacent directions used in adjacent, already decoded blocks.

[0016]

[0016] FIG. 1B shows a schematic (105) indicating 65 intra prediction directions by JEM, showing the gradually increasing number of prediction directions.

[0017]

[0017] The mapping of intra prediction direction bits in a coded video bitstream representing a direction may vary for each video coding technique; for example, it may range from a simple direct mapping of codewords to intra prediction modes of prediction directions, to complex adaptive schemes including the most likely modes, or similar techniques. However, in all cases, there may be certain directions in the video content that are statistically less likely to occur than certain other directions. Since the goal of video compression is to reduce redundancy, in well - operating video coding techniques, less likely directions are represented with more bits than more likely directions.

[0018]

[0018] Motion compensation can be a lossless compression technique, and it can be associated with a technique in which, after spatially shifting in the direction indicated by a motion vector (hereinafter referred to as MV), a block of sample data from a previously reconstructed picture or a part thereof (reference picture) is used for prediction of a newly reconstructed picture or a part of the picture. In some cases, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions X and Y, or three dimensions, where the third dimension indicates the reference picture in use (the latter can be considered, indirectly, as the temporal dimension).

[0019]

[0019] In some video compression techniques, the MVs applicable to a particular area of sample data can be predicted from other MVs, for example, those related to other areas of sample data spatially adjacent to the area being reconstructed and preceding that MV in the decoding order. By doing so, the amount of data required to code the MVs can be significantly reduced, thereby removing redundancy and enhancing compression. For example, when coding an input video signal derived from a camera (known as natural video), there is a statistical likelihood that areas larger than the area to which a single MV is applicable move in a similar direction, and thus in some cases, it is possible to predict using a similar motion vector derived from the MVs of adjacent areas, so MV prediction may function effectively. This results in an MV that is found to be similar to or the same as the MV predicted from surrounding MVs for a given area, and it can be represented with fewer bits than when directly coding the MV after entropy coding. In some cases, MV prediction may be an example of lossless compression of the signal (i.e., the MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself may be lossy due to rounding errors, for example, when calculating predictors from several surrounding MVs.

[0020]

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec.H.265, “High Efficiency Video Coding”, December 2016). Among the many MV prediction mechanisms provided by H.265, the one described in this application is a technique that will hereafter be called “spatial merge”.

[0021]

[0021] Referring to FIG. 1, the current block (111) contains samples discovered by the encoder during the motion search process so that it can be predicted from a previous block of the same size that has been spatially shifted. Instead of directly coding the MV, the MV can be derived from the metadata associated with one or more reference pictures, for example, using the MV associated with any of five neighboring samples (112 to 116 respectively) denoted as A0, A1, and B0, B1, B2, from the latest reference picture (in the decoding order). In H.265, MV prediction can use predictors from the same reference picture as that used by adjacent blocks.

Summary of the Invention

[0022]

[0022] Aspects of the present disclosure provide an apparatus for video encoding / decoding. The apparatus includes a processing circuit that decodes a video bitstream to obtain a reduced-resolution residual block for a current block. The processing circuit determines that a block-level flag is set to a predefined value. The predefined value indicates that the current block is coded in reduced-resolution coding. Based on the block-level flag, the processing circuit generates a reduced-resolution prediction block for the current block by downsampling a full-resolution reference block of the current block. The processing circuit generates a reduced-resolution reconstruction block for the current block based on the reduced-resolution prediction block and the reduced-resolution residual block. The processing circuit generates a full-resolution reconstruction block for the current block by upsampling the reduced-resolution reconstruction block.

[0023]

[0023] In an embodiment, the processing circuit determines the size of a prediction block at a reduced resolution based on the size of a reference block at full resolution and the downsampling rate of the current block.

[0024]

[0024] In an embodiment, the processing circuit decodes a block-level flag for the current block from a video bitstream. The block-level flag indicates that the current block is coded with reduced resolution coding.

[0025]

[0025] In an embodiment, the processing circuit decodes one of a filter coefficient or an index of the filter coefficient from a video bitstream. The filter coefficient is used when upsampling a reconstructed block at a reduced resolution.

[0026]

[0026] In an embodiment, the processing circuit scales the motion vector of a first adjacent block of the current block based on a scaling factor that is the ratio of the downsampling rates of the current block and the first adjacent block. The processing circuit constructs a first motion vector candidate list for the current block. The first motion vector candidate list includes the scaled motion vectors of the first adjacent blocks.

[0027]

[0027] In an embodiment, the processing circuit determines a scaled motion vector based on a shift operation in response to the scaling factor being a power of two. In one example, when the scaling factor is 2 N , in order to obtain the scaled motion vector, the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded. In another example, when the scaling factor is 2 N , first, the motion vector is added with a rounding factor (e.g., 2 N-1 ), and then the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded to obtain the scaled motion vector.

[0028]

[0028] In an embodiment, the processing circuit determines a scaled motion vector based on a look-up table in response to the scaling factor not being a power of two.

[0029]

[0029] In an embodiment, the processing circuit determines the priority of the scaled motion vector within the first motion vector candidate list based on the downsampling rates of the current block and the first adjacent block.

[0030]

[0030] In an embodiment, the processing circuit constructs a second motion vector candidate list for the current block based on one or more second adjacent blocks of the current block. Each of the one or more second adjacent blocks has the same downsampling rate as the current block. The processing circuit constructs a third motion vector candidate list for the current block based on one or more third adjacent blocks of the current block. Each of the one or more third adjacent blocks has a different downsampling rate from the current block.

[0031]

[0031] In an embodiment, the processing circuit scans the third motion vector candidate list based on the number of motion vector candidates within the second motion vector candidate list being less than a specified number.

[0032]

[0032] In an embodiment, the processing circuit determines a fourth motion vector candidate list for the current block by merging the second motion vector candidate list and the third motion vector candidate list in an interleaved manner.

[0033]

[0033] In an embodiment, the processing circuit determines the affine parameters of the current block based on the downsampling rate of the current block.

[0034]

[0034] Aspects of the present disclosure disclose a method for video encoding / decoding. In the method, a video bitstream is decoded to obtain a reduced-resolution residual block for a current block. A block-level flag set to a pre-defined value is determined. The pre-defined value indicates that the current block is coded in reduced-resolution coding. Based on the block-level flag, a reduced-resolution prediction block for the current block is generated by downsampling a full-resolution reference block of the current block. Based on the reduced-resolution prediction block and the reduced-resolution residual block, a reduced-resolution reconstruction block is generated for the current block. By upsampling the reduced-resolution reconstruction block, a full-resolution reconstruction block is generated for the current block.

[0035]

[0035] Also, aspects of the present disclosure provide a non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one or a combination of methods for video encoding / decoding.

Brief Description of the Drawings

[0036]

[0036] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1A

[0037] Figure 1A is a schematic diagram of an exemplary subset of intra prediction modes.

Figure 1B

[0038] Figure 1B is an illustration of an exemplary intra prediction direction.

Figure 1C

[0039] Figure 1C is a schematic diagram of a current block and its surrounding spatial merge candidates in one example.

Figure 2

[0040] Figure 2 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.

Figure 3

[0041] Figure 3 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment.

Figure 4

[0042] Figure 4 is a schematic diagram of a simplified block diagram of a decoder according to an embodiment.

Figure 5

[0043] Figure 5 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment.

Figure 6

[0044] Figure 6 shows a block diagram of an encoder according to another embodiment.

Figure 7

[0045] Figure 7 shows a block diagram of a decoder according to another embodiment.

Figure 8

[0046] Figure 8 shows an exemplary block partitioning according to some embodiments of the present disclosure.

Figure 9A

[0047] Figure 9A shows an exemplary block partitioning using a quad-tree-plus-binary-tree (QTBT) and a corresponding tree structure according to an embodiment of the present disclosure.

Figure 9B

[0047] Figure 9B shows an exemplary block partitioning using a quad-tree-plus-binary-tree (QTBT) and a corresponding tree structure according to an embodiment of the present disclosure.

Figure 10

[0048] Figure 10 shows an exemplary nominal angle according to an embodiment of the present disclosure.

Figure 11

[0049] Figure 11 shows the positions of the upper, left, and left part samples with respect to one pixel in the current block according to an embodiment of the present disclosure.

Figure 12

[0050] Figure 12 shows an exemplary recursive filter intra mode according to an embodiment of the present disclosure

Figure 13

[0051] FIG. 13 shows an exemplary multi-layer reference frame structure according to an embodiment of the present disclosure.

Figure 14

[0052] FIG. 14 shows an exemplary candidate motion vector list construction process according to an embodiment of the present disclosure.

Figure 15

[0053] FIG. 15 shows an exemplary motion field estimation process according to an embodiment of the present disclosure.

Figure 16A

[0054] FIGS. 16A and 16B show exemplary overlapping regions (shaded regions) predicted using upper adjacent blocks and left adjacent blocks, respectively.

Figure 16B

[0054] FIGS. 16A and 16B show exemplary overlapping regions (shaded regions) predicted using upper adjacent blocks and left adjacent blocks, respectively.

Figure 17

[0055] FIG. 17 shows an exemplary two-step warping process in which vertical shearing is performed following horizontal shearing.

Figure 18

[0056] FIG. 18 shows an overall loop filtering pipeline including frame-level super-resolution in AV1.

Figure 19

[0057] FIG. 19 shows an exemplary implementation using block-level flags according to an embodiment of the present disclosure.

Figure 20

[0058] FIG. 20 shows an exemplary reference of spatially adjacent motion vectors according to an embodiment of the present disclosure.

Figure 21

[0059] FIG. 21 shows an exemplary reference of temporally adjacent motion vectors according to an embodiment of the present disclosure.

Figure 22

[0060] FIG. 22 shows an exemplary reference of spatially adjacent motion vectors referred to for affine motion prediction according to an embodiment of the present disclosure.

Figure 23

[0061] FIG. 23 shows an exemplary flowchart according to an embodiment.

Figure 24

[0062] FIG. 24 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0037]

[0063] I. VIDEO DECODER AND ENCODER SYSTEM

[0064] FIG. 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present disclosure. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via a network (250). In the example of FIG. 2, the first pair of terminal devices (210) and (220) perform unidirectional data transmission. For example, the terminal device (210) can code video data (e.g., a stream of video pictures captured by the terminal device (210)) for transmission to other terminal devices (220) via the network (250). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (220) can receive the coded video data from the network (250), decode the coded video data to restore the video picture, and display the video picture according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0038]

[0065] In another example, the communication system (200) includes, for example, a second pair of terminal devices (230) and (240) that perform bidirectional transmission of coded video data that may occur during a video conference. With respect to the bidirectional transmission of data, for example, each of the terminal devices (230) and (240) can code video data (e.g., a stream of video pictures captured by a terminal device) for transmission to the other of the terminal devices (230) and (240) via the network (250). Each of the terminal devices (230) and (240) can also receive the coded video data transmitted by the other of the terminal devices (230) and (240), can decode the coded video data to restore the video picture, and can display the video picture on an accessible display device according to the restored video data.

[0039]

[0066] In the example of FIG. 2, the terminal devices (210), (220), (230), and (240) are shown as servers, personal computers, and smartphones, but the principles of the present disclosure are not so limited. Embodiments of the present disclosure find applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. The network (250) represents any number of networks that carry the coded video data between the terminal devices (210), (220), (230), and (240), including, for example, wired (wired) and / or wireless communication networks. The communication network (250) can exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the present disclosure, the architecture and topology of the network (250) may not be important for the operation of the present disclosure, unless otherwise described below.

[0040]

[0067] FIG. 3 shows the arrangement of a video encoder and a video decoder in a streaming environment as an application example of the disclosed subject matter. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storage on digital media (including CDs, DVDs, memory sticks, etc.) of compressed video.

[0041]

[0068] The streaming system can include a video source (301), such as a digital camera, and may include a capture subsystem (313) capable of generating a stream of, for example, uncompressed video pictures (302). In one example, the stream of video pictures (302) includes samples taken by a digital camera. The stream of video pictures (302), drawn as a thick line to emphasize a larger amount of data when compared to the encoded video data (304) (or coded video bitstream), can be processed by an electronic device (320) including a video encoder (303) coupled to the video source (301). The video encoder (303) includes hardware, software, or a combination thereof and can operate or implement aspects of the disclosed subject matter as detailed below. The encoded video data (304) (or encoded video bitstream (304)), drawn as a thin line to emphasize a smaller amount of data when compared to the stream of video pictures (302), can be stored in a streaming server (305) for future use. One or more streaming client subsystems, such as the client subsystems (306) and (308) of FIG. 3, can access the streaming server (305) to retrieve copies (307) and (309) of the encoded video data (304). The client subsystem (306) can include, for example, a video decoder (310) within an electronic device (330). The video decoder (310) decodes an incoming copy (307) of the encoded video data and generates an output stream of video pictures (311) that can be rendered on a display (312), such as a display screen, or other rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., video bitstreams) can be encoded according to a particular video coding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0042]

[0069] Note that electronic devices (320) and (330) can include other components (not shown). For example, electronic device (320) can include a video decoder (not shown), and electronic device (330) can also include a video encoder (not shown).

[0043]

[0070] FIG. 4 shows a block diagram of a video decoder (410) according to an embodiment of the present disclosure. The video decoder (410) can be included in an electronic device (430). The electronic device (430) can include a receiver (431) (e.g., a receiving circuit). The video decoder (410) can be used in place of the video decoder (310) in the example of FIG. 3.

[0044]

[0071] Receiver (431) is capable of receiving one or more coded video sequences to be decoded by video decoder (410); in the same or another embodiment, if the decoding of each coded video sequence is independent of other coded video sequences, it is possible to receive one coded video sequence at a time. The coded video sequence can be received from channel (401), which may be a hardware / software link to a storage device storing the encoded video data. Receiver (431) is capable of receiving the encoded video data together with other data, such as coded audio data and / or auxiliary data streams, and these data can be transferred using respective entities (not shown). Receiver (431) can separate the coded video sequence from other data. To handle network jitter, buffer memory (415) may be coupled between receiver (431) and entropy decoder / parser (420) (hereinafter referred to as "parser (420)"). In certain applications, buffer memory (415) is part of video decoder (410). In other cases, it may be outside video decoder (410) (not shown). In yet another example, for example, to handle network jitter, there may be a buffer memory (not shown) outside video decoder (410), and furthermore, for example, to handle playback timing, there may be another buffer memory (415) inside video decoder (410). If receiver (431) is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from a synchronous network, buffer memory (415) may not be required or can be made smaller.For use in a best - effort packet network such as the Internet, buffer memory (415) may be required, which may be relatively large and advantageously may be of an adaptable size and may be implemented at least partially in an operating system or similar element (not shown) outside the video decoder (410).

[0045]

[0072] The video decoder (410) can include a parser (420) to reconstruct symbols (421) from the coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (410), and potentially information for controlling a rendering device (412) (e.g., a display screen) that is not an essential part of the electronic device (430) but can be coupled to the electronic device (430), as shown in FIG. 4. The control information for the rendering device may be in the form of supplementary enhancement information (SEI message) or a video user utility information (VUI) parameter set fragment (not shown). The parser (420) can parse / entropy decode the received coded video sequence. The coding of the coded video sequence can conform to a video coding technology or standard and can follow various principles including variable length coding, Huffman coding, arithmetic coding with or without context influence, etc. The parser (420) can extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroups can include a group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (420) can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0046]

[0073] The parser (420) can perform an entropy decoding / parsing process on the video sequence received from the buffer memory (415) to generate symbols (421).

[0047]

[0074] The reconstruction of symbol (421) can include a plurality of different units depending on the type of the coded video picture or a part thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is included can be controlled by subgroup control information parsed by a parser (420) from the coded video sequence. Such a flow of subgroup control information between the parser (420) and a plurality of subsequent units is not depicted for clarity.

[0048]

[0075] The video decoder (410) can be conceptually subdivided into a plurality of functional units as described below, in addition to the functional blocks already described. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0049]

[0076] The first unit is a scaler / inverse transform unit (451). The scaler / inverse transform unit (451) receives, as symbol (421), not only the quantized transform coefficients but also control information (including the transform to be used, block size, quantization factor, quantization scaling matrix, etc.) from the parser (420). The scaler / inverse transform unit (451) can output a block including sample values that can be input to an aggregator (455).

[0050]

[0077] In some cases, the output samples of the scaler / inverse transform (451) may be related to intra-coded blocks: i.e., blocks that do not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (452). In some cases, the intra-picture prediction unit (452) uses the already reconstructed surrounding information fetched from the buffer (458) of the current picture to generate blocks of the same size and shape as the block being reconstructed. The current picture buffer (458) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (455) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (452) to the output sample information as provided by the scaler / inverse transform unit (451).

[0051]

[0078] Otherwise, the output samples of the scaler / inverse transform unit (451) may be related to inter-coded motion-compensable blocks. In such cases, the motion compensation prediction unit (453) can access the reference picture memory (457) to retrieve the samples used for prediction. According to the symbols (421) related to the block, after motion-compensating the retrieved samples, these samples are added by the aggregator (455) to the output of the scaler / inverse transform unit (451) (in this case, called the residual samples or residual signal) to generate output sample information. The address in the reference picture memory (457) from which the motion compensation prediction unit (453) retrieves the prediction samples can be controlled by the motion vectors available to the motion compensation prediction unit (453) in the form of, for example, X, Y, and symbols (421) that can have reference picture components. Also, motion compensation can include interpolation of sample values taken from the reference picture memory (457), a motion vector prediction mechanism, etc., when exact sub-sample motion vectors are used.

[0052]

[0079] The output samples of the aggregator (455) can be affected by various loop filtering techniques within the loop filter unit (456). The video compression technology includes loop filter techniques that are included in the coded video sequence (also called the coded video bitstream) and are controlled by parameters made available to the loop filter unit (456) as symbols (421) from the parser (420), but can respond to meta information obtained during the decoding of previous parts of the coded picture or coded video sequence (in decoding order), and can also respond to previously reconstructed loop-filtered sample values.

[0080] The output of the loop filter unit (456) can be output not only to the rendering device (412), but also can be a sample stream that can be stored in the reference picture memory (457) for future inter-picture prediction.

[0053]

[0081] Once a given coded picture is completely reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is completely reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (420)), the buffer (458) of the current picture can become part of the reference picture memory (457), and the buffer of the new current picture can be reallocated before starting the reconstruction of the subsequent coded picture.

[0082] The video decoder (410) is capable of performing a decoding operation according to a predetermined video compression technique in a standard such as ITU-T Rec.H.265. The coded video sequence can comply with the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence complies with both the syntax of the video compression technique or standard and the profile as documented in the video compression technique or standard. Specifically, the profile can select specific tools as the only tools available under that profile from all the tools available in the video compression technique or standard. Also, for compliance, it is necessary that the complexity of the coded video sequence falls within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the virtual reference decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.

[0054]

[0083] In an embodiment, the receiver (431) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (410) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0055]

[0084] Figure 5 shows a block diagram of a video encoder (503) according to an embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) can be used in place of the video encoder (303) in the example of FIG. 3.

[0056]

[0085] The video encoder (503) can receive video samples from a video source (501) (not part of the electronic device (520) in the example of FIG. 5) that can capture video images to be coded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).

[0057]

[0086] The video source (501) can provide a source video sequence to be coded by the video encoder (503) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (501) may be a storage device that stores pre-prepared videos. In a video conferencing system, the video source (501) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed as a sequence. Each picture itself can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0058]

[0087] According to an embodiment, the video encoder (503) can code and compress pictures of a source video sequence into a coded video sequence (543) in real time or under any other arbitrary time constraints required by an application. Enforcing an appropriate coding speed is one function of the controller (550). In some embodiments, the controller (550) controls other functional units and is functionally coupled to other functional units as described below. The coupling is not depicted for clarity. Parameters set by the controller (550) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (550) can be configured to have other appropriate functions related to the video encoder (503) optimized for a particular system design.

[0059]

[0088] In some embodiments, the video encoder (503) is configured to operate in a coding loop. As an extremely simplified explanation, in one example, the coding loop can include a source coder (530) (which is responsible for generating symbols such as a symbol stream based on the input picture and reference pictures to be coded), and a (local) decoder (533) incorporated in the video encoder (503). The decoder (533) reconstructs the symbols to generate sample data in the same way as a (remote) decoder does (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input into the reference picture memory (534). Since the decoding of the symbol stream results in a bit-exact result independent of the decoder's location (local or remote), the content in the reference picture memory (534) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" exactly the same sample values as the decoder would "see" as reference picture samples when the decoder uses prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is also used in some related technologies.

[0060]

[0089] The operation of the "local" decoder (533) can be assumed to be the same as that of a "remote" decoder such as the video decoder (410) that has already been described in detail above in connection with FIG. 4. However, referring briefly to FIG. 4, it is possible to assume that symbols are available and that the encoding / decoding of the symbol-coded video sequence by the entropy coder (545) and the parser (420) is lossless. Therefore, the entropy decoding section of the video decoder (410) including the buffer memory (415) and the parser (420) may not be fully realized in the local decoder (533).

[0061]

[0090] An observation that can be made at this point is that any decoder technology other than parsing / entropy decoding existing in the decoder must necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter focuses on the operation of the decoder. The description of the encoder technology can be omitted because it is the reverse of the decoder technology that has been comprehensively described. More detailed descriptions are required only in specific areas and are given below.

[0062]

[0091] During operation, the source coder (530) can perform motion-compensated predictive coding in which, in some examples, it predicts and codes an input picture with reference to one or more previously coded pictures from the video sequence designated as "reference pictures". In this way, the coding engine (532) codes the difference between the pixel block of the input picture and the pixel block of the reference picture that can be selected as a prediction reference for the input picture.

[0063]

[0092] The local video decoder (533) can decode the coded video data of a picture that can be specified as a reference picture based on the symbols generated by the source coder (530). The operation of the coding engine (532) may advantageously be a lossless process. If the coded video data can be decoded by a video decoder (not shown in FIG. 5), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (533) can repeat the decoding process that can be performed by the video decoder in the reference picture, causing the reconstructed reference picture to be stored in the reference picture cache (534). In this way, the video encoder (503) can locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the video decoder at the remote end (assuming no transmission errors).

[0064]

[0093] The predictor (535) can perform a prediction search for the coding engine (532). That is, for a new picture to be coded, the predictor (535) can search the reference picture memory (534) for sample data (as a candidate reference pixel block) or predetermined metadata (reference picture motion vectors, block shapes, etc.), which may serve as an appropriate prediction reference for the new picture. The predictor (535) can operate on a sample block - pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (534) as determined by the search result obtained by the predictor (535).

[0065]

[0094] The controller (550) can manage the coding operations of the source coder (530), including, for example, the setting of parameters and subgroup parameters used to encode video data.

[0066]

[0095] All outputs of the aforementioned functional units can be entropy-coded in the entropy coder (545). The entropy coder (545) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0067]

[0096] The transmitter (540) can buffer the coded video sequence as created by the entropy coder (545) and prepare it for transmission via the communication channel (560), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (540) can merge the coded video data from the video coder (503) with other data to be transmitted, such as, for example, coded audio data and / or auxiliary data streams (sources not shown).

[0068]

[0097] The controller (550) can manage the operation of the video encoder (503). During coding, the controller (550) can assign a specific coded picture type to each of the coded pictures, which may affect the coding techniques applicable to each picture. For example, a picture may often be assigned as one of the following picture types.

[0069]

[0098] An Intra Picture (I Picture) can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of Intra Pictures, including, for example, Independent Decoder Refresh (“IDR”) Pictures. Those skilled in the art are aware of these variations of I Pictures, as well as their respective uses and characteristics.

[0070]

[0099] A Predicted Picture (P Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0071]

[0100] A Bi - Directionally Predicted Picture (B Picture) can be encoded and decoded using Intra prediction or Inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of one block.

[0072]

[0101] The source picture is typically spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples) and can be coded block by block. The blocks can be predictive coded with reference to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictive coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be predictive coded by spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictive coded by spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0073]

[0102] The video encoder (503) can perform coding operations in accordance with a predetermined video coding technology or standard such as ITU-T Rec.H.265. In this operation, the video encoder (503) can execute various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. The coded video data can thus conform to the syntax specified by the video coding technology or standard being used.

[0074]

[0103] In an embodiment, the transmitter (540) can transmit additional data along with the coded video. The source coder (530) can include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0075]

[0104] Video can be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture under encoding / decoding, called the current picture, is partitioned into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously coded and is still buffered in the video, the block within the current picture can be coded by a vector called an MV. The MV points to the reference block within the reference picture and can have a third dimension to identify the reference picture when multiple reference pictures are used.

[0076]

[0105] In some embodiments, it is possible to use dual-prediction techniques for inter-picture prediction. According to the dual-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture in decoding order within the video (although they may be in the past and future respectively in display order), are used. A block within the current picture can be coded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0077]

[0106] Further, in order to improve coding efficiency, it is possible to use merge-mode techniques for inter-picture prediction.

[0078]

[0107] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and CTUs within a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is partitioned into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0079]

[0108] FIG. 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) receives a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and is configured to encode the processing block into a coded picture that is part of a coded video sequence. In one example, the video encoder (603) is used instead of the video encoder (303) of the example of FIG. 3.

[0080]

[0109] In an example of HEVC, a video encoder (603) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (603) uses an intra mode, an inter mode, or a bi-prediction mode to determine, for example using rate distortion optimization, whether the processing block is best coded. If the processing block is to be coded in the intra mode, the video encoder (603) can use intra prediction techniques to encode the processing block for the coded picture; if the processing block is to be coded in the inter mode or the bi-prediction mode, the video encoder (603) can use inter prediction techniques or bi-prediction techniques respectively to encode the processing block for the coded picture. In certain video coding techniques, the merge mode can be an inter prediction picture sub-mode, in which case the motion vector is derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.

[0081]

[0110] In the example of FIG. 6, the video encoder (603) includes an inter encoder (630), an intra encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general purpose controller (621), and an entropy encoder (625) coupled together as shown in FIG. 6.

[0082]

[0111] The inter - encoder (630) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and blocks in a subsequent picture), generate inter - prediction information (e.g., a description of redundant information by an encoding technique, a motion vector, merge - mode information), and calculate an inter - prediction result (e.g., a predicted block) based on the inter - prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0083]

[0112] The intra - encoder (622) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block with blocks already coded within the same picture, generate quantized coefficients after transformation, and optionally also generate intra - prediction information (e.g., intra - prediction direction information according to one or more intra - coding techniques). In one example, the intra - encoder (622) also calculates an intra - prediction result (e.g., a predicted block) based on the intra - prediction information and reference blocks within the same picture.

[0084]

[0113] The General Controller (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In one example, the General Controller (621) determines the mode of a block and provides a control signal to the switch (626) based on that mode. For example, if the mode is the intra mode, the General Controller (621) controls the switch (626) to select the intra mode result for use by the residual calculator (623) and controls the entropy encoder (625) to select the intra prediction information and include the intra prediction information in the bitstream; if the mode is the inter mode, the General Controller (621) controls the switch (626) to select the inter prediction result for use by the residual calculator (623) and controls the entropy encoder (625) to select the inter prediction information and include the inter prediction information in the bitstream.

[0085]

[0114] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra-encoder (622) or the inter-encoder (630). The residual encoder (624) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (624) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra-encoder (622) and the inter-encoder (630). For example, the inter-encoder (630) can generate a decoded block based on the decoded residual data and the inter-prediction information, and the intra-encoder (622) can generate a decoded block based on the decoded residual data and the intra-prediction information. The decoded block is appropriately processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.

[0086]

[0115] The entropy encoder (625) is configured to format the bitstream to include the encoded block. The entropy encoder (625) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other appropriate information in the bitstream. It should be noted that there is no residual information when coding a block in either the inter-mode or the merge sub-mode of the bi-prediction mode according to the disclosed subject matter.

[0087]

[0116] FIG. 7 shows a diagram of a video decoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive a coded picture that is part of a coded video sequence and decode the coded picture to generate a reconstructed picture. In one example, the video decoder (710) is used in place of the video decoder (310) in the example of FIG. 3.

[0088]

[0117] In the example of FIG. 7, the video decoder (710) includes an entropy decoder (771), an inter decoder (780), a residual decoder (773), a reconstruction module (774), and an intra decoder (772) coupled together as shown in FIG. 7.

[0089]

[0118] The entropy decoder (771) can be configured to reconstruct from the coded picture certain symbols that represent syntax elements making up the coded picture. Such symbols can include, for example, the mode in which a block is coded (e.g., the latter two in intra mode, inter mode, bi-prediction mode, merge submode or another submode), prediction information (e.g., intra prediction information or inter prediction information) that can identify certain samples or metadata used for prediction by the intra decoder (772) or the inter decoder (780) respectively, residual information (e.g., in the form of quantized transform coefficients), etc. In one example, when the prediction mode is inter or bi-prediction mode, inter prediction information is provided to the inter decoder (780); when the prediction type is intra prediction type, intra prediction information is provided to the intra decoder (772). The residual information can be inverse quantized and provided to the residual decoder (773).

[0090]

[0119] The inter-decoder (780) is configured to receive inter-prediction information and generate an inter-prediction result based on the inter-prediction information.

[0091]

[0120] The intra-decoder (772) is configured to receive intra-prediction information and generate a prediction result based on the intra-prediction information.

[0092]

[0121] The residual decoder (773) is configured to perform inverse quantization to extract non-quantized transform coefficients, and process the non-quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require certain control information (including quantization parameter (QP)), and such information may be provided by the entropy decoder (771) (since this may be only a small amount of control information, the data path is not depicted).

[0093]

[0122] The reconstruction module (774) is configured to combine, in the spatial domain, the residual as the output by the residual decoder (773) and the prediction result (which may be output by the inter or intra prediction module in some cases) to form a reconstructed block, and the reconstructed block is part of the reconstructed picture, and the reconstructed picture may be part of the reconstructed video. Note that other appropriate processing such as deblocking processing may be performed to improve the visual quality.

[0123] Note that the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using any suitable technology. In an embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more processors that execute software instructions.

[0094]

[0124] II. Block Partition

[0125] FIG. 8 shows an exemplary block partition according to some embodiments of the present disclosure.

[0095]

[0126] In some related cases such as VP9 proposed by AOMedia (Alliance for Open Media), four types of partition trees can be used. As shown in FIG. 8, starting from the 64×64 level and descending to the 4×4 level, there are some additional restrictions for blocks of 8×8 or less. The partition designated by R can be called a recursive partition. That is, until the lowest 4×4 level is reached, the same partition tree is repeated at a lower scale.

[0096]

[0127] In some related cases such as AV1 proposed by AOMedia and based on VP9, the partition tree can be extended to 10 structures as shown in FIG. 8, and the maximum coding block size (referred to as superblock in VP9 / AV1 terms) is increased to start from 128x128. It should be noted that the 4:1 / 1:4 rectangular partition is included in AV1 but not in VP9. The rectangular partition cannot be further subdivided. Furthermore, when using partitions below the 8×8 level, AV1 can support more flexibility because in some cases, it is possible to perform inter prediction on 2×2 chroma blocks.

[0097]

[0128] In some related cases such as HEVC, in order to adapt to various local features, by using a quadtree structure shown as a coding tree, a CTU can be divided into CUs. The decision of whether to code a picture area using inter-picture (temporal) or intra-picture (spatial) prediction can be made at the CU level. Each CU can be further divided into one, two, or four PUs according to the PU partition type. Within one PU, the same prediction process can be applied, and the related information can be sent to the decoder on a PU basis. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be partitioned into TUs according to another quadtree structure such as the coding tree related to the CU. One of the important features of the HEVC structure is having multiple partition concepts including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square in shape, but a PU can be square or rectangular in the case of an inter-prediction block. In HEVC, one coding block can be further divided into four square sub-blocks, and the transformation process can be executed for each sub-block, i.e., TU. Each TU can be further recursively divided (e.g., using quadtree partitioning) into smaller TUs. Quadtree partitioning can be called a residual quadtree (RQT).

[0098]

[0129] At the picture boundary, HEVC performs implicit quadtree partitioning, and as a result, the block can continue to perform quadtree partitioning until the block size conforms to the picture boundary.

[0099]

[0130] In some related cases such as VVC, a quad-tree-plus-binary-tree (QTBT) partitioning structure can be applied. The quad-tree-plus-binary-tree (QTBT) structure removes the concept of multiple partition types (i.e., removes the distinction between the concepts of CU, PU, and TU), and supports rich flexibility for the CU partition shape.

[0100]

[0131] Figures 9A and 9B show exemplary block partitioning according to an embodiment of the present disclosure using QTBT and corresponding tree structures. Solid lines indicate QT splits, and dotted lines indicate BT splits. At each split (i.e., non-leaf) node of the BT, one flag is signaled to indicate which split type (i.e., horizontal or vertical) is used. In Figure 9B, 0 indicates a horizontal split and 1 indicates a vertical split. In the case of a QT split, there is no need to specify the split type because a QT split always splits the block in both the horizontal and vertical directions, generating four sub-blocks of the same size.

[0101]

[0132] In the QTBT structure, a CU can have either a square or rectangular shape. As shown in Figures 9A and 9B, a CTU is first partitioned by the QT structure. A QT leaf node can be further split by the BT structure. There are two split types for BT splitting: symmetric horizontal splitting and symmetric vertical splitting. A BT leaf node is a CU, and the segmentation into two CUs is used for prediction and transformation processing without any further partitioning. Therefore, a CU, a PU, and a TU can have the same block size in the QTBT structure.

[0102]

[0133] A CU can contain CBs of different color components as in JEM. For example, in the case of P and B slices using a 4:2:0 chroma format, one CU can contain one luma CB and two chroma CBs. In other examples, a CU can contain a single-component CB. For example, in the case of an I slice, one CU can contain only one luma CB or only two chroma CBs.

[0103]

[0134] The following parameters are defined for the QTBT partitioning method: CTU size (the size of the root node of the QT, such as in HEVC), MinQTSize (the minimum allowable QT leaf node size), MaxBTSize (the maximum allowable BT root node size), MaxBTDepth (the maximum allowable BT depth), MinBTSize (the minimum allowable BT leaf node size).

[0104]

[0135] In an example of the QTBT partitioning structure, the CTU size is set as 128×128 luma samples, there are two corresponding 64×64 blocks of chroma samples, MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize is set to 4×4 (for both width and height), and MaxBTDepth is set to 4. QT partitioning is first applied to the CTU to generate QT leaf nodes. The QT leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf QT node is 128×128, the size exceeds MaxBTSize (i.e., 64×64), so it will not be further split by BT. Otherwise, the leaf QT node can be further split by the BT tree. Therefore, the QT leaf node is also the root node of the BT and has a BT depth of 0. When the BT depth reaches MaxBTDepth (i.e., 4), no further splitting is considered. If a BT node has a width equal to MinBTSize (i.e., 4), no further horizontal splitting is considered. Similarly, if a BT node has a height equal to MinBTSize, no further vertical splitting is considered. The leaf nodes of the BT are further processed by prediction and transformation processing without any other partitioning. For example, the maximum CTU size is 256×256 luma samples as in JEM.

[0105]

[0136] III. Prediction in AV1

[0137] In some related cases such as VP9, eight direction modes corresponding to angles from 45 degrees to 207 degrees are supported. In some related cases such as AV1, to utilize the spatial redundancy of more diverse directional textures, the directional intra mode is extended to an angle set with finer granularity. The original eight angles are slightly modified and referred to as nominal angles, and these eight angles are named V_PRED, H_PRED, D45_PRED, D135_PRED, D113_PRED, D157_PRED, D203_PRED, and D67_PRED.

[0106]

[0138] Figure 10 shows exemplary nominal angles according to an embodiment of the present disclosure. In some related cases such as AV1, since each nominal angle can be associated with seven finer angles, there can be a total of 56 direction angles. The predicted angle can be represented by the nominal intra angle plus the angle delta. The angle delta is equal to the coefficient multiplied by a step size of 3 degrees. The coefficient can be in the range of -3 to 3. In AV1, eight nominal modes are first signaled together with five non-angular smooth modes. Then, if the current mode is an angle mode, an index is further signaled to indicate the angle delta with respect to the corresponding nominal angle. To implement the direction prediction mode in AV1 in a general way, all 56 directional intra prediction angles in AV1 can be realized using a unified direction predictor, which projects each pixel to a reference sub-pixel position and interpolates the reference sub-pixel by a 2-tap bilinear filter.

[0107]

[0139] In some related cases such as AV1, there are five non-directional smooth intra prediction modes, which are DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of the top-left adjacent samples is used as the predictor for the block to be predicted. In PAETH prediction, first, the top, left, and top-left reference samples are taken, and then the value closest to (top + left - top-left) is set as the predictor for the pixel to be predicted.

[0108]

[0140] Figure 11 shows the positions of the top, left, and top-left samples for one pixel in the current block according to an embodiment of the present disclosure. In the case of the SMOOTH, SMOOTH_V, and SMOOTH_H modes, the block is predicted using second-order interpolation in the vertical or horizontal direction, or the average in both directions.

[0109]

[0141] In some related cases such as AV1, an intra prediction mode based on recursive filtering can be used.

[0110]

[0142] Figure 12 shows an exemplary recursive filter intra mode according to an embodiment of the present disclosure.

[0111]

[0143] To capture the decaying spatial correlation with references at the edge, the FILTER INTRA mode is designed for luma blocks. In AV1, five filter intra modes are defined, and each mode is represented by a set of eight 7-tap filters that reflect the correlation between pixels in a 4x2 patch and seven neighbors of the patch. For example, the weighting coefficients of the 7-tap filters are position-dependent. As shown in FIG. 12, an 8×8 block is divided into eight 4×2 patches, which are denoted by B0, B1, B2, B3, B4, B5, B6, and B7. For each patch, seven neighbors denoted by R0~R7 are used to predict the pixels within each patch. In the case of patch B0, all neighbors have already been reconstructed. However, in other patches, not all neighbors have been reconstructed, and the predicted value of the immediately preceding neighbor is used as the reference value. For example, since not all neighbors of patch B7 have been reconstructed, the predicted samples of the neighbors of patch B7 (i.e., B5 and B6) are used instead.

[0112]

[0144] In some related cases such as VP9, three reference frames can be used for inter prediction. The three reference frames include the LAST (nearest past), GOLDEN (distant past), and ALTREF (temporally filtered future) frames.

[0113]

[0145] In some related cases such as AV1, an extended reference frame can be used. For example, in addition to the three reference frames used in VP9, AV1 can use four additional types of reference frames. The four additional reference frames include the LAST2, LAST3, BWDREF, and ALTREF2 frames. The LAST2 and LAST3 frames are two recent past frames, and the BWDREF and ALTREF2 frames are two future frames. Further, the BWDREF frame is a look-ahead frame coded without temporal filtering and is more useful as a backward reference at relatively short distances. The ALTREF2 frame is an intermediate filtered future reference frame between the GOLDEN frame and the ALTREF frame.

[0114]

[0146] FIG. 13 shows an exemplary multi-layer reference frame structure according to an embodiment of the present disclosure. In FIG. 13, an adaptive number of frames share the same GOLDEN and ALTREF frames. The BWDREF frame is a look-ahead frame directly coded without applying temporal filtering and is thus more applicable as a backward reference at relatively short distances. The ALTREF2 frame functions as an intermediate filtered future reference between the GOLDEN frame and the ALTREF frame. All new references can be selected by a single prediction mode or can be paired to form a composite mode. AV1 provides a rich set of reference frame pairs, resulting in both bidirectional and unidirectional composite predictions, and thus various videos with dynamic temporal correlation characteristics can be coded in a more adaptive and optimal way.

[0115]

[0147] FIG. 14 shows an exemplary candidate motion vector list construction process according to an embodiment of the present disclosure. Spatial and temporal reference motion vectors can be classified into two categories based on where they appear (the nearest spatial neighbors, and others). In some related cases, motion vectors from the immediately preceding upper, left, and upper-right adjacent blocks of the current block can have a higher correlation with the current block than other blocks, and thus are considered to have a higher priority. Within each category, the motion vectors are ranked within the spatial and temporal search ranges in descending order of their appearance counts. Motion vector candidates with higher appearance counts can be considered "popular" in the local region, i.e., having a higher prior probability. The two categories are concatenated to form a ranked list.

[0116]

[0148] FIG. 15 shows an exemplary motion field estimation process according to an embodiment of the present disclosure.

[0117]

[0149] In some related cases such as AV1, dynamic spatial and temporal motion vector references can be used. For example, in order to efficiently code motion vectors, a motion vector reference selection method can be incorporated. In the motion vector reference selection method, the spatial neighborhood can extend more widely than that used in VP9. Furthermore, a temporal motion vector reference candidate can be found using a motion field estimation process. The motion field estimation process may operate in three stages: motion vector buffering, motion trajectory creation, and motion vector projection. First, for each of the coded frames, the reference frame index and the related motion vector of each coded frame can be stored. The stored information can be referenced by the next coding frame and can generate the motion field of the next coding frame. Motion field estimation can examine, for example, in FIG. 15, a motion trajectory (e.g., MV ref2 aiming from a block in a reference frame Ref2 to another reference frame Ref0 Ref2 ). Next, the motion field estimation process searches through all motion trajectories passing through each 64×64 processing unit at 8×8 block resolution through a collocated 128×128 area. Then, at the coding block level, when the reference frame is determined, motion vector candidates can be derived by linearly projecting the passing motion trajectories onto the desired reference frame, for example, converting MV ref2 to MV0 or MV1 in FIG. 15.

[0118]

[0150] Once all candidate motion vectors are found, the candidate motion vectors can be sorted, merged, ranked, and increased to four final candidates. Then, it is possible to signal the index of the reference motion vector selected from the list, and optionally, it is possible to code the motion vector difference.

[0119]

[0151] In some related cases such as AV1, in order to reduce the prediction error around the block boundary by combining predictions obtained from adjacent motion vectors, block - based prediction can be combined with secondary predictors from the top and left ends by applying 1 - D filters in the vertical and horizontal directions respectively. This method can be referred to as overlapped block motion compensation (OBMC).

[0120]

[0152] FIGS. 16A and 16B respectively show exemplary overlapping regions (shaded regions) predicted using the upper - adjacent block (2) and the left - adjacent block (4). The shaded region of the prediction block (0) can be predicted by recursively generating mixed prediction samples with a 1 - D raised - cosine filter.

[0121]

[0153] In some related cases such as AV1, two affine prediction models called global warped motion compensation and a local warped motion compensation can be used. The former signals a frame - level affine model between the frame and its reference, and the latter implicitly processes varying local motion with minimal overhead. The local motion parameters can be derived at the block level by using 2D motion vectors from the neighborhood of the cause. This affine model is realized by continuous horizontal and vertical shearing operations based on an 8 - tap interpolation filter at 1 / 64 pixel accuracy.

[0122]

[0154] FIG. 17 shows an exemplary two - stage warping process where vertical shearing follows horizontal shearing. In FIG. 17, the affine model is realized by local warping motion compensation that first performs horizontal shearing and then vertical shearing.

[0123]

[0155] IV. Frame-based Super-Resolution in AV1

[0156] Figure 18 shows the overall loop filtering pipeline including frame-level super-resolution in AV1. On the encoder side, the source frame can first be downscaled in a non-normative way and encoded at a lower resolution. On the decoder side, a deblocking filter and a constrained directional enhancement filter (CDEF) are applied to remove coding artifacts while preserving edges at a low resolution. Then, a linear upsampling filter can be applied only along the horizontal direction to obtain a full-resolution reconstruction. Optionally, a loop restoration filter can then be applied at full resolution to recover the high-frequency details lost during downsampling and quantization.

[0124]

[0157] In some related cases such as AV1, super-resolution is a special mode signaled at the frame level. Each coded frame can use a horizontal-only super-resolution mode at any resolution within the ratio constraint range. Whether to apply linear upsampling after decoding and the scaling ratio used can be signaled. The upsampling ratio can potentially have nine possible values given as d / 8 for d = 8, 9,... and 16. The corresponding downsampling ratio before encoding can be assumed to be 8 / d.

[0125]

[0158] When the output frame dimensions WxH and the upsampling ratio d are given, both the encoder and the decoder can calculate the low-resolution coded frame dimensions as wxH, where the reduced width is w = (8W + d / 2) / d. The linear upscaling process takes in a frame of reduced resolution wxH and outputs a frame of dimensions WxH specified in the frame header. The canonical horizontal linear upscaler in AV1 uses a 1 / 16-phase linear 8-tap filter for interpolation of each row.

[0126]

[0159] V. CU-based Super-Resolution Coding

[0160] In some related cases such as AV1, super-resolution is performed at the frame level. That is, super-resolution is applied to all areas of a picture with a certain scaling ratio. However, the statistics of the signals in different areas within a picture can vary greatly. Therefore, applying downsampling and / or upsampling to all areas may not necessarily result in a good rate-distortion trade-off.

[0127]

[0161] Adapting the application of downsampling and / or upsampling to picture regions can be performed using preprocessing approaches and / or postprocessing approaches such as using mask and / or segment information to select the picture regions for downsampling and / or upsampling. However, this process cannot guarantee that the improvement in rate-distortion performance will exceed other coding methods that do not use super-resolution.

[0128]

[0162] The present disclosure includes a method of block-level super-resolution coding.

[0129]

[0163] According to an aspect of the present disclosure, mixed-resolution prediction can be used to emulate frame-level super-resolution adapted for use at the block level.

[0130]

[0164] In one embodiment, a block-level flag can be used to specify whether reduced-resolution coding is used for a coding block. For example, when the block-level flag is set to a first predefined value (e.g., 1), reduced-resolution coding is enabled for that block. Prediction sample generation, such as the prediction procedure in AV1 as described above, can be performed by generating a low-resolution prediction block for the coding block using reference samples or pictures at reduced resolution. Then, a low-resolution reconstruction block for the coding block can be generated using the low-resolution prediction block. Finally, the reduced-resolution reconstruction block can be upsampled to a full-resolution reconstruction block for the coding block.

[0131]

[0165] In one embodiment, when the block-level flag is set to a second predefined value (e.g., 0), prediction sample generation is performed at full resolution (or original resolution), followed by full-resolution reconstruction.

[0132]

[0166] Figure 19 shows an exemplary implementation using a block-level flag according to an embodiment of the present disclosure. When the block-level flag is on, for example, when the block-level flag is equal to a first predefined value, the source block (1901) on the encoder side can be downsampled by the downsampler module (1920) to generate a downsampled source block (1902). The downsampled source block (1902) can be combined with a low-resolution prediction block (1903) to generate a downsampled residual block (1904). Then, the downsampled residual block (1904) can be coded into a coded video bitstream by a module (or multiple modules not shown in Figure 19) including a conversion, quantization, and entropy coding process. To decode the source block (1901), the coded video bitstream received on the decoder side can be processed by an entropy decoding, inverse quantization, and inverse conversion process to generate a downsampled residual block (1911). The downsampled residual block (1911) can be combined with a reduced-resolution prediction block (1912) to generate a downsampled reconstruction block (1913). The downsampled reconstruction block (1913) can be upsampled by the upsampling module (1930) to generate a full-resolution reconstruction block (1914) for the source block (1901). Note that the reduced-resolution prediction block (1912) can be generated by the downsampler module (1940) by downsampling a full-resolution reconstruction block for a reference block of the source block (1901).

[0133]

[0167] If the block-level flag is off, for example, if the block-level flag is equal to a second predefined value, the down-sample modules (1920), (1940) and the up-sampler module (1930) of FIG. 19 are not applied. The down-sampled or low-resolution blocks (1902)-(1904) and (1911)-(1913) can be the corresponding full-resolution ones.

[0134]

[0168] In one embodiment, for a coding block having a size of MxN, the reference block having a size of MxN is down-sampled along the horizontal and vertical directions at down-sampling rates D x and D y respectively, and a low-resolution prediction block having a size of (M / D x )×(N / D y ) can be generated. Exemplary values of M and N can include, but are not limited to, 256, 128, 64, 32, 16, and 8. The down-sampling rates D x and D y are integers including 2, 4, and 8, but are not limited thereto.

[0169] In one embodiment, the block-level flag can be adaptively signaled or estimated for each CU, super-block, prediction block, transform block, tile, coded segment, frame, or sequence basis.

[0135]

[0170] In one embodiment, regarding the up-sampler module (1930) used to up-sample the low-resolution reconstruction block, the up-sampling filter coefficients can be directly signaled, or the index of the set of filter coefficients among a plurality of sets of predefined coefficients can be signaled.

[0136]

[0171] According to an aspect of the present disclosure, in order to be referenced by the motion vector of the current block, the motion vectors of spatially and / or temporally adjacent blocks of the current block can be scaled at the same or different resolutions as the current block.

[0137]

[0172] For example, if the current block is coded along the horizontal and vertical directions with sampling ratios (or down-sampling rates) D x and D y respectively, the motion vectors of spatially and / or temporally adjacent blocks having sampling ratios D ref,x and D ref,y can be scaled for each of the horizontal and vertical components

[0138]

Number

[0139]

[0173] FIG. 20 shows an exemplary reference of spatially adjacent motion vectors according to an embodiment of the present disclosure.

[0140]

[0174] In an embodiment, based on a motion vector reference method such as the spatially adjacent motion vector reference method in AV1, the reference motion vector list construction process can search for adjacent regions in the order of (1) to (8) shown in FIG. 20 in WxH luma sample units. The upper WxH region T ij , the left WxH region L ij , the top left WxH region TL, and the top right WxH region TR, the motion vectors for are scaled by

[0141]

Number

[0142]

[0175] FIG. 21 shows a reference to exemplary temporally adjacent motion vectors according to an embodiment of the present disclosure.

[0143]

[0176] In one embodiment, based on a motion vector reference method such as the temporally adjacent motion vector reference method in AV1, the motion vectors mf_mv_1 and mf_mv_2 for the current block located at (blk_row, blk_col) within the current frame can be obtained as follows.

[0144]

[0177] As shown in FIG. 21, the motion vector ref_mv for the WxH region located at (ref_blk_row, ref_blk_col) within the specified search area of the reference frame (reference_frame1) is used to find the motion trajectory towards the previous frame (prior_frame). When the trajectory intersects the current block located at (blk_row, blk_col), the motion vectors mf_mv_1 and mf_mv_2 for reference_frame1 and reference_frame2 are given as follows:

[0145]

Equation

[0146]

[0178] The derived motion vectors mf_mv_1 and mf_mv_2, before being used as reference motion vectors in the candidate list,

[0147]

Equation

[0148]

[0179] According to an aspect of the present disclosure, the motion vector candidates of the current block can be classified into two (or more) categories. As shown in FIG. 14, the motion vectors of the spatially adjacent blocks located in the immediately upper row, the immediately left column, and the upper right corner of the current block can be classified as the first category (e.g., category 1) of the motion vector candidates of the current block, while all other candidates are classified as the second category (e.g., category 2). Within each category, the motion vector candidates can be sorted in descending order of the count number in which each candidate appears. That is, the first motion vector candidate that appears more frequently than the second motion vector candidate in the candidate list is placed before the second motion vector candidate in the candidate list. Further, the candidate list of the first category (e.g., category 1) can be concatenated with the candidate list of the second category (e.g., category 2) to form a single candidate list.

[0149]

[0180] In some embodiments, in addition to the category and the count number within each category, the sampling ratio used for the motion vector candidates can be incorporated when constructing the candidate list.

[0150]

[0181] In one embodiment, the motion vector candidates from adjacent blocks having the same sampling rate as the current block can have a higher priority regarding being selected within the candidate list of each category under the same number of appearances.

[0151]

[0182] In one embodiment, two separate candidate lists can be constructed. The first candidate list can include only candidate motion vectors from adjacent blocks having the same sampling ratio. The second candidate list can include candidate motion vectors from adjacent blocks having different sampling ratios. Which candidate list is used can be signaled or inferred in addition to the signaling of the index of the selected reference motion vector within the candidate list.

[0152]

[0183] In one embodiment, the second candidate list can be scanned only if the number of candidates within the first candidate list is less than the specified number of reference motion vectors that are indexed and signaled.

[0153]

[0184] In one embodiment, based on an interleaving scheme, the two separate candidate lists can be merged to form a single list. For example, if the total number of candidates from both lists is greater than or equal to the specified number of reference motion vectors, candidates from the first candidate list and another from the second candidate list can be selected in ascending order of entry position within the list until the specified number of reference motion vectors within the combined list is reached.

[0154]

[0185] FIG. 22 shows an exemplary reference of spatially adjacent motion vectors for affine motion prediction according to an embodiment of the present disclosure.

[0155]

[0186] In one embodiment, based on a motion vector reference method such as the affine motion prediction method in AV1, an affine model that projects a sample at (x, y) within the current block to a predicted sample within a reference block at (x', y') within the reference frame is given as follows:

[0156]

Equation

[0187] Affine parameter {h ij:i = 1, 2 and j = 1, 2} can be obtained as follows. The sample position within the current frame is (a k , b k ) = (x k , y k ) - (x0, y0), where k is the index of an adjacent block having the same reference frame as the current block (k = 0 corresponds to the current block). In the example of FIG. 22, k = 2, 3, 5, 6.

[0157]

[0188] Then, the corresponding sample position within the reference frame is given as follows:

[0158]

Equation

[0159]

[0189] The least-squares solution can be obtained as follows:

[0160]

Equation

[0161]

Equation

[0190] In the above description, it is assumed that the affine parameters (h 13 , h 23 ) correspond to the translational motion vector at full resolution. When a reduced-resolution prediction using the sampling ratio D is applied to a block, (h 13 , h23 ) accordingly

[0162]

Number

[0163]

[0191] According to an aspect of the present disclosure, when scaling adjacent motion vectors to code a block using a reduced-resolution coding mode, the horizontal and vertical components of the scaled motion vector can be derived using one of the following approaches.

[0164]

[0192] In the first approach, when the scaling factor is a power of 2 (for example, 2 N ), the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded to obtain the scaled motion vector.

[0165]

[0193] In the second approach, when the scaling factor is a power of 2 (for example, 2 N ), first, the motion vector is added with a rounding factor (for example, 2 N-1 ), and then the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded to obtain the scaled motion vector.

[0166]

[0194] In the third approach, when the scaling factor is not a power of 2, a look-up table can be used to derive the values of the horizontal and vertical components of the scaled motion vector.

[0167]

[0195] VI. Flowchart

[0196] FIG. 23 shows a flowchart illustrating an exemplary process (2300) according to an embodiment of the present disclosure. In various embodiments, process (2300) is executed by a processing circuit such as the processing circuits of terminal devices (210), (220), (230), (240), the processing circuit that executes the functions of video encoder (303), the processing circuit that executes the functions of video decoder (310), the processing circuit that executes the functions of video decoder (410), the processing circuit that executes the functions of intra prediction module (452), the processing circuit that executes the functions of video encoder (503), the processing circuit that executes the functions of predictor (535), the processing circuit that executes the functions of intra encoder (622), and the processing circuit that executes the functions of intra decoder (772). In some embodiments, process (2300) is implemented by software instructions, and when the processing circuit executes the software instructions, the processing circuit executes process (2300).

[0168]

[0197] Process (2300) can generally start at step (S2310), where process (2300) decodes a video bitstream to obtain a low-resolution residual block for the current block. Then, process (2300) proceeds to step (S2320).

[0169]

[0198] In step (S2330), process (2300) determines that the block-level flag is set to a predefined value. The predefined value indicates that the current block is coded in reduced-resolution coding. Then, process (2300) proceeds to step (S2330).

[0170]

[0199] In step (2330), process (2300) generates a reduced-resolution prediction block for the current block by downsampling the full-resolution reference block of the current block. Then, process (2300) proceeds to step (S2340).

[0171]

[0200] In step (S2340), process (2300) generates a reduced-resolution reconstruction block for the current block based on the reduced-resolution prediction block and the reduced-resolution residual block. Then, process (2300) proceeds to step (S2350).

[0172]

[0201] In step (2350), process (2300) generates a full-resolution reconstruction block for the current block by upsampling the reduced-resolution reconstruction block. Then, process (2300) ends.

[0173]

[0202] In an embodiment, process (2300) determines the size of the reduced-resolution prediction block based on the size of the full-resolution reference block and the downsampling rate of the current block.

[0174]

[0203] In an embodiment, process (2300) decodes a block-level flag for the current block from the video bitstream. The block-level flag indicates that the current block is coded with reduced-resolution coding.

[0175]

[0204] In an embodiment, process (2300) decodes one of a filter coefficient or an index of the filter coefficient from the video bitstream. The filter coefficient is used when upsampling the reduced-resolution reconstruction block.

[0176]

[0205] In an embodiment, process (2300) scales the motion vector of the first adjacent block of the current block based on a scaling factor that is the ratio of the downsampling rates of the current block and the first adjacent block. Process (2300) constructs a first motion vector candidate list for the current block. The first motion vector candidate list includes the scaled motion vectors of the first adjacent blocks.

[0177]

[0206] In an embodiment, process (2300) determines the scaled motion vector based on a shift operation in response to the scaling factor being a power of two. In one example, when the scaling factor is 2 N , the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded to obtain the scaled motion vector. In another example, when the scaling factor is 2 N , first, the motion vector is added with a rounding factor (e.g., 2 N-1 ), and then the lower N bits of the horizontal component of the motion vector and the lower N bits of the vertical component of the motion vector are discarded to obtain the scaled motion vector.

[0178]

[0207] In an embodiment, process (2300) determines the scaled motion vector based on a look-up table in response to the scaling factor not being a power of two.

[0179]

[0208] In an embodiment, process (2300) determines the priority of the scaled motion vector in the first motion vector candidate list based on the downsampling rates of the current block and the first adjacent block.

[0180]

[0209] In an embodiment, process (2300) constructs a second motion vector candidate list for a current block based on one or more second adjacent blocks of the current block. Each of the one or more second adjacent blocks has the same downsampling rate as the current block. Process (2300) constructs a third motion vector candidate list for the current block based on one or more third adjacent blocks of the current block. Each of the one or more third adjacent blocks has a downsampling rate different from that of the current block.

[0181]

[0210] In an embodiment, process (2300) scans the third motion vector candidate list based on the number of motion vector candidates in the second motion vector candidate list being less than a specified number.

[0182]

[0211] In an embodiment, process (2300) determines a fourth motion vector candidate list for the current block by merging the second motion vector candidate list and the third motion vector candidate list in an interleaved manner.

[0183]

[0212] In an embodiment, process (2300) determines the affine parameters of the current block based on the downsampling rate of the current block.

[0184]

[0213] VII. COMPUTER SYSTEM

[0214] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 24 shows a computer system (2400) suitable for implementing a particular embodiment of the disclosed subject matter.

[0185]

[0215] Computer software can be coded using any suitable machine code or computer language that can be the subject of assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or instructions that go through interpretation or microcode execution, etc.

[0186]

[0216] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.

[0187]

[0217] The components shown in FIG. 24 for the computer system (2400) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software that implements the embodiments of the present disclosure. Also, the component configuration should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (2400).

[0188]

[0218] A computer system (2400) can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movements of a data glove), auditory input (e.g., voice, clapping), visual input (e.g., gestures), and olfactory input (not shown). Also, the human interface device can be used to capture specific media such as audio (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic pictures) that are not necessarily directly related to conscious human input.

[0189]

[0219] The input human interface device can potentially include one or more of a keyboard (2401), a mouse (2402), a trackpad (2403), a touch screen (2910), a data glove (not shown), a joystick (2405), a microphone (2406), a scanner (2407), and a camera (2408) (although only one of each is depicted).

[0190]

[0220] The computer system (2400) can also include a specific human interface output device. Such a human interface output device can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such a human interface output device can be a tactile output device (e.g., tactile feedback by a touch screen (2410), a data glove (not shown), a joystick (2405), although there may also be a tactile feedback device that does not serve as an input device), an auditory output device (e.g., a speaker (2409), headphones (not shown)), a visual output device (e.g., a screen (2410) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each of which may or may not have a touch screen input function, each of which may or may not have a tactile feedback function, and some of which may be capable of outputting three-dimensional or higher-dimensional output by means such as two-dimensional visual output, stereoscopic output; virtual reality glasses (not shown), a holographic display, and a smoke tank (not shown)), and a printer (not shown). These visual output devices (such as the screen (2410)) can be connected to the system bus (2448) via a graphics adapter (2450).

[0191]

[0221] The computer system (2400) can also include optical media such as a CD / DVD ROM / RW (2420) using a medium (2421) such as a CD / DVD, a thumb drive (2422), a removable hard drive or a solid state drive (2423), legacy magnetic media (not shown) such as tapes and floppy disks (not shown), and human-accessible storage devices and their related media such as specialized ROM / ASIC / PLD-based devices such as a security dongle (not shown).

[0192]

[0222] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include a transmission medium, a carrier wave, or other transient signals.

[0193]

[0223] The computer system (2400) can also include an interface to one or more communication networks (2455). The one or more networks (2455) can be, for example, wireless, wired, or optical. The one or more networks (2455) can further be related to local, wide area, metropolitan, vehicle industry, real-time, delay-tolerant, etc. Examples of the one or more networks (2455) include Ethernet, wireless LAN, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), wired or wireless wide area digital networks for TV (including cable TV, satellite TV, and terrestrial broadcast TV), vehicle industry including CANBus, etc. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (2449) (e.g., the USB port of the computer system (2400)); others are generally integrated into the core of the computer system (2400) by attaching to the system bus as described below (e.g., an Ethernet interface is integrated within a PC computer system, and a cellular network interface is integrated within a smartphone computer system). Using any of these networks, the computer system (2400) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CANbus for certain CANbus devices), or two-way, for example, for other computer systems using local or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0194]

[0224] The foregoing human interface device, human accessible storage device, and network interface can be attached to the core (2440) of a computer system (2400).

[0195]

[0225] The core (2440) can include one or more central processing units (CPUs) (2441), a graphics processing unit (GPU) (2442), a special programmable processing device in the form of a field programmable gate array (FPGA) (2443), a hardware accelerator for specific tasks (2444), and the like. These devices can be connected via a system bus (2448) together with a read only memory (ROM) (2445), a random access memory (2446), and an internal mass storage device (e.g., an internal non-user accessible hard drive, SSD, etc.) (2447). In some computer systems, the system bus (2448) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2448) of the core or via a peripheral bus (2449). In one example, a screen (2410) can be connected to a graphics adapter (2450). The architecture of the peripheral bus includes PCI, USB, and the like.

[0196]

[0226] The CPU (2441), GPU (2442), FPGA (2443), and accelerator (2444) can be combined to execute specific instructions capable of constituting the aforementioned computer code. The computer code can be stored in the ROM (2445) or RAM (2446). Temporary data can be stored in the RAM (2446), while persistent data can be stored, for example, in the internal mass storage (2447). Fast storage and retrieval for any memory device may be made possible by using cache memory, which can be closely associated with one or more CPUs (2441), GPUs (2442), mass storage (2447), ROM (2445), RAM (2446), etc.

[0197]

[0227] A computer-readable medium can have thereon computer code for executing various computer-implemented operations. The medium and the computer code can be considered to be specially designed and constructed for the purposes of this disclosure, or they can be considered to be of the kind well-known and available to those of ordinary skill in the computer software art.

[0198]

[0228] By way of illustration and not limitation, a computer system having an architecture (2400), specifically a core (2440), can provide the function of executing software embodied on one or more tangible computer-readable media as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable media can be media related to user-accessible mass storage as described above, similar to specific storage of the core (2440) of a non-transitory nature such as mass storage (2447) inside the core or ROM (2445). The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (2440). The computer-readable media can include one or more memory devices or chips according to specific needs. The software includes defining a data structure stored in RAM (2446) and modifying such a data structure according to a process defined by the software, and causing the core (2440) and in particular the processor (including a CPU, GPU, FPGA, etc.) therein to execute a specific process or a specific part of a specific process described in the present application. Further or alternatively, the computer system can provide a function as a result of logic wired or otherwise incorporated within a circuit (e.g., an accelerator (2444)), and the circuit can execute a specific process or a specific part of a specific process described in the present application instead of or together with software. References to software include logic and, if necessary, vice versa. References to a computer-readable media can include a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both where appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0199]

[0229] Although several exemplary embodiments have been described, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Therefore, those skilled in the art will understand that although not explicitly illustrated or described in this application, it is possible to embody the principles of the present disclosure and thus devise many systems and methods that are within its spirit and scope.

[0200]

[0230] Supplementary Note (Supplementary Note 1) A video decoding method in a decoder, comprising: decoding a video bitstream to obtain a residual block with reduced resolution for the current block; determining that a block-level flag is set to a predefined value, the predefined value indicating that the current block is coded in reduced-resolution coding; generating a prediction block with reduced resolution for the current block by downsampling a reference block with full resolution of the current block based on the block-level flag; generating a reconstructed block with reduced resolution for the current block based on the prediction block with reduced resolution and the residual block with reduced resolution; and generating a reconstructed block with full resolution for the current block by upsampling the reconstructed block with reduced resolution. A method comprising the above steps.

[0201] (Supplementary Note 2) In the method according to Supplementary Note 1, the step of generating the prediction block with reduced resolution comprises: determining the size of the prediction block with reduced resolution based on the size of the reference block with full resolution and the downsampling rate of the current block. A method comprising the above step.

[0202] (Supplementary Note 3) In the method according to Supplementary Note 1, the determining step is: including the step of decoding the block-level flag for the current block from the video bitstream, wherein the block-level flag indicates that the current block is coded in the reduced resolution coding; a method.

[0203] (Supplementary Note 4) In the method according to Supplementary Note 1, the determining step is: further including the step of decoding one of the filter coefficients or the index of the filter coefficient from the video bitstream, wherein the filter coefficient is used when upsampling the reconstructed block of the reduced resolution; a method.

[0204] (Supplementary Note 5) In the method according to Supplementary Note 1: scaling the motion vector of the first adjacent block of the current block based on a scaling coefficient that is the ratio of the downsampling rates of the current block and the first adjacent block; and constructing a first motion vector candidate list for the current block, wherein the first motion vector candidate list includes the scaled motion vector of the first adjacent block; a step. further including; a method.

[0205] (Supplementary Note 6) In the method according to Supplementary Note 5, the scaling step is: determining the scaled motion vector based on a shift operation in response to the scaling coefficient being a power of 2; and determining the scaled motion vector based on a look-up table in response to the scaling coefficient not being a power of 2; including; a method.

[0206] (Supplementary Note 7) In the method according to Supplementary Note 5, a step of determining the priority of the scaled motion vector in the first motion vector candidate list based on the downsampling rates of the current block and the first adjacent block; A method further comprising.

[0207] (Supplementary Note 8) In the method according to Supplementary Note 1: A step of constructing a second motion vector candidate list for the current block based on one or more second adjacent blocks of the current block, each of the one or more second adjacent blocks having the same downsampling rate as the current block; and A step of constructing a third motion vector candidate list for the current block based on one or more third adjacent blocks of the current block, each of the one or more third adjacent blocks having a downsampling rate different from that of the current block; A method further comprising.

[0208] (Supplementary Note 9) In the method according to Supplementary Note 8: A step of scanning the third motion vector candidate list based on the number of motion vector candidates in the second motion vector candidate list that is less than a specified number; A method further comprising.

[0209] (Supplementary Note 10) In the method according to Supplementary Note 8: A step of determining a fourth motion vector candidate list for the current block by merging the second motion vector candidate list and the third motion vector candidate list in an interleaved manner; A method further comprising.

[0210] (Supplementary Note 11) In the method according to Supplementary Note 1: A step of determining the affine parameters of the current block based on the downsampling rate of the current block; A method further comprising

[0211] (Appendix 12) An apparatus comprising a processing circuit, wherein the processing circuit: Decoding a video bitstream to obtain a reduced-resolution residual block for a current block; Determining that a block-level flag is set to a predefined value, the predefined value indicating that the current block is coded in reduced-resolution coding; Generating a reduced-resolution prediction block for the current block by downsampling a full-resolution reference block of the current block based on the block-level flag; Generating a reduced-resolution reconstruction block for the current block based on the reduced-resolution prediction block and the reduced-resolution residual block; and Generating a full-resolution reconstruction block for the current block by upsampling the reduced-resolution reconstruction block; An apparatus configured to perform

[0212] (Appendix 13) In the apparatus according to Appendix 12, the processing circuit: Determining the size of the reduced-resolution prediction block based on the size of the full-resolution reference block and the downsampling rate of the current block; An apparatus further configured to perform

[0213] (Appendix 14) In the apparatus according to Appendix 12, the processing circuit: The apparatus is further configured to decode the block-level flag for the current block from the video bitstream, the block-level flag indicating that the current block is coded in the reduced-resolution coding.

[0214] (Appendix 15) In the apparatus according to Appendix 12, the processing circuit: is further configured to decode one of the filter coefficients or the index of the filter coefficient from the video bit stream, and the filter coefficient is used when upsampling the reconstructed block of the reduced resolution, apparatus.

[0215] (Appendix 16) In the apparatus according to Appendix 12, the processing circuit: scaling the motion vector of the first adjacent block of the current block based on a scaling coefficient that is a ratio of the downsampling rate of the current block and the first adjacent block; and constructing a first motion vector candidate list for the current block, the first motion vector candidate list including the scaled motion vector of the first adjacent block, step; is further configured to perform, apparatus.

[0216] (Appendix 17) In the apparatus according to Appendix 16, the processing circuit: determining the scaled motion vector based on a shift operation in response to the scaling coefficient being a power of 2; and determining the scaled motion vector based on a look-up table in response to the scaling coefficient not being a power of 2; is further configured to perform, apparatus.

[0217] (Appendix 18) In the apparatus according to Appendix 16, the processing circuit: is further configured to determine the priority of the scaled motion vector in the first motion vector candidate list based on the downsampling rate of the current block and the first adjacent block, apparatus.

[0218] (Appendix 19) In the apparatus according to Supplementary Note 12, the processing circuit: Constructing a second motion vector candidate list for the current block based on one or more second adjacent blocks of the current block, each of the one or more second adjacent blocks having the same downsampling rate as the current block; and Constructing a third motion vector candidate list for the current block based on one or more third adjacent blocks of the current block, each of the one or more third adjacent blocks having a downsampling rate different from that of the current block; An apparatus further configured to perform.

[0219] (Supplementary Note 20) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to: Decoding a video bitstream to obtain a reduced-resolution residual block for a current block; Determining that a block-level flag is set to a predefined value, the predefined value indicating that the current block is coded in reduced-resolution coding; Generating a reduced-resolution prediction block for the current block by downsampling a full-resolution reference block of the current block based on the block-level flag; Generating a reduced-resolution reconstruction block for the current block based on the reduced-resolution prediction block and the reduced-resolution residual block; and Generating a full-resolution reconstruction block for the current block by upsampling the reduced-resolution reconstruction block; A storage medium that causes the above to be executed.

Explanation of Signs

[0220]

[0231] Appendix A: Acronyms ALF: Adaptive Loop Filter (Adaptive Multiple Transform) AMVP: Advanced Motion Vector Prediction APS: Adaptation Parameter Set ASIC: Application-Specific Integrated Circuit ATMVP: Alternative / Advanced Temporal Motion Vector Prediction AV1: AOMedia Video 1 AV2: AOMedia Video 2 BMS: Benchmark Set BV: Block Vector CANBus: Controller Area Network Bus CB: Coding Block CC-ALF: Cross-Component Adaptive Loop Filter CD: Compact Disc CDEF: Constrained Directional Enhancement Filter CPR: Current Picture Referencing CPU: Central Processing Unit CRT: Cathode Ray Tube (Cathode Ray Tube) CTB: Coding Tree Block (Coding Tree Block) CTU: Coding Tree Unit (Coding Tree Unit) CU: Coding Unit (Coding Unit) DPB: Decoder Picture Buffer (Decoder Picture Buffer) DPCM: Differential Pulse - Code Modulation (Differential Pulse - Code Modulation) DPS: Decoding Parameter Set (Decoding Parameter Set) DVD: Digital Video Disc (Digital Video Disc) FPGA: Field Programmable Gate Area (Field Programmable Gate Area) JCCR: Joint CbCr Residual Coding (Joint CbCr Residual Coding) JVET: Joint Video Exploration Team (Joint Video Exploration Team) GOP: Groups of Pictures (Pictures Group) GPU: Graphics Processing Unit (Graphics Processing Unit) GSM: Global System for Mobile communications (Global System for Mobile Communications) HDR: High Dynamic Range (High Dynamic Range) HEVC: High Efficiency Video Coding (High Efficiency Video Coding) HRD: Hypothetical Reference Decoder (Hypothetical Reference Decoder) IBC: Intra Block Copy (Intra Block Copy) IC: Integrated Circuit (Integrated Circuit) ISP: Intra Sub-Partitions (Intra Sub-Partition) JEM: Joint Exploration Model (Joint Video Exploration Team) LAN: Local Area Network (Local Area Network) LCD: Liquid-Crystal Display (Liquid Crystal Display) LR: Loop Restoration Filter (Loop Restoration Filter) LRU: Loop Restoration Unit (Loop Restoration Unit) LTE: Long-Term Evolution (Long-Term Evolution) MPM: Most Probable Mode (Most Probable Mode) MV: Motion Vector (Motion Vector) OLED: Organic Light-Emitting Diode (Organic Light-Emitting Diode) PBs: Prediction Blocks (Prediction Blocks) PCI: Peripheral Component Interconnect (Peripheral Component Interconnect) PDPC: Position Dependent Prediction Combination (Position Dependent Prediction Combination) PLD: Programmable Logic Device (Programmable Logic Device) PPS: Picture Parameter Set (Picture Parameter Set) PU: Prediction Unit (Prediction Unit) RAM: Random Access Memory (Random Access Memory) ROM: Read-Only Memory (Read-Only Memory) SAO: Sample Adaptive Offset(Sample Adaptive Offset) SCC: Screen Content Coding(Screen Content Coding) SDR: Standard Dynamic Range(Standard Dynamic Range) SEI: Supplementary Enhancement Information(Supplementary Enhancement Information) SNR: Signal Noise Ratio(Signal Noise Ratio) SPS: Sequence Parameter Set(Sequence Parameter Set) SSD: Solid-state Drive(Solid-state Drive) TU: Transform Unit(Transform Unit) USB: Universal Serial Bus(Universal Serial Bus) VPS: Video Parameter Set(Video Parameter Set) VUI: Video Usability Information(Video Usability Information) VVC: Versatile Video Coding(Versatile Video Coding) WAIP: Wide-Angle Intra Prediction(Wide-Angle Intra Prediction)

Claims

1. 1. A video decoding method in a decoder, comprising: decoding the video bitstream to obtain a reduced-resolution residual block for a current block; generating a reduced resolution prediction block for the current block by down-sampling a full resolution reference block of the current block; generating a reduced-resolution reconstructed block for the current block based on the reduced-resolution prediction block and the reduced-resolution residual block; and generating a full resolution reconstructed block for the current block by up-sampling the reduced resolution reconstructed block; The method comprises: scaling motion vectors of neighboring blocks of the current block based on a scaling factor that is a ratio of down-sampling rates of the current block and the neighboring blocks, and constructing a motion vector candidate list including the scaled motion vectors of the neighboring blocks; wherein a priority of the scaled motion vector in the motion vector candidate list is determined such that, when a motion vector is selected from the motion vector candidate list for reconstruction of the current block, a motion vector candidate from the neighboring block that has the same down-sampling ratio as the current block has a higher priority.

2. 2. The method of claim 1, wherein the motion vector candidates are classified into a plurality of categories, and motion vector candidates from immediately adjacent blocks above, to the left, or to the upper right of the current block have a higher priority than motion vector candidates from other blocks.

3. 2. The method of claim 1, wherein generating the reduced resolution prediction block comprises: determining a size of the reduced resolution prediction block based on a size of the full resolution reference block and a down-sampling ratio of the current block; A method comprising:

4. The method of claim 1, 23. The method of claim 22, further comprising: decoding a block level flag for the current block from the video bitstream, the block level flag indicating that the current block is coded at the reduced resolution.

5. The method of claim 1, The method of claim 1, further comprising the step of decoding one of filter coefficients or indices of the filter coefficients from the video bitstream, the filter coefficients being used in upsampling the reduced resolution reconstructed block.

6. 2. The method of claim 1, wherein the scaling step comprises: determining the scaled motion vector based on a shift operation in response to the scaling factor being a power of two; and determining the scaled motion vector based on a look-up table in response to the scaling factor not being a power of two; A method comprising:

7. 13. The method of claim 1 : constructing a second motion vector candidate list for the current block based on one or more second neighboring blocks of the current block, each of the one or more second neighboring blocks having the same down-sampling rate as the current block; and constructing a third motion vector candidate list for the current block based on one or more third neighboring blocks of the current block, each of the one or more third neighboring blocks having a different down-sampling rate than the current block; The method further comprises:

8. 8. The method of claim 7, scanning the third motion vector candidate list based on a number of motion vector candidates in the second motion vector candidate list being less than a specified number; The method further comprises:

9. 8. The method of claim 7, determining a fourth motion vector candidate list for the current block by merging the second motion vector candidate list and the third motion vector candidate list in an interleaved manner; The method further comprises:

10. 13. The method of claim 1 : determining affine parameters for the current block based on a down-sampling rate of the current block; The method further comprises:

11. 11. Apparatus comprising processing circuitry configured to carry out the method of any one of claims 1 to 10.

12. A computer program product configured to cause a computer processor to carry out a method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Sampling-based super-resolution video coding and decoding method and apparatus

    JP2013518463A

  • Spatio-temporal motion vector prediction patterns for video coding

    US20200186825A1

  • Block-level super-resolution based video coding

    WO2019197674A1