Video coding method and apparatus using history-based motion vector prediction

The history-based motion vector prediction system with dedicated tables for each CTU row and wavefront parallel processing addresses the challenge of high-definition video data encoding, improving efficiency by reducing motion vector data and maintaining image quality.

JP7706584B2Active Publication Date: 2025-07-11BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024009209
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-07-18
Filing Date
2024-01-25
Publication Date
2025-07-11
Estimated Expiration
2039-07-16

AI Technical Summary

Technical Problem

The increasing amount of video data with high-definition formats such as 4K×2K or 8K×4K poses a challenge in efficiently encoding and decoding video data while maintaining image quality, as the data required to represent motion vectors substantially increases.

Method used

Implementing a history-based motion vector prediction (HMVP) system that utilizes a dedicated HMVP table for each CTU row, allowing parallel processing of video data through wavefront parallel processing (WPP), and constructing motion vector candidate lists using spatially and temporally adjacent CUs, reducing the need to encode actual motion vectors in the bitstream.

Benefits of technology

This approach significantly reduces the data required to represent motion information in the video bitstream, enhancing encoding efficiency and maintaining image quality by reusing motion vectors within the same CTU row, even in parallel processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706584000001
    Figure 0007706584000001
  • Figure 0007706584000002
    Figure 0007706584000002
  • Figure 0007706584000003
    Figure 0007706584000003
Patent Text Reader

Abstract

To provide a video coding method and system using history-based motion vector prediction.SOLUTION: Processing by a video coder includes obtaining an encoded video bitstream including data associated with a plurality of encoded pictures, resetting a history-based motion vector predictor (HMVP) table for the current CTU row, maintaining a plurality of motion vector predictors in the HMVP table while decoding the current CTU row, extracting the prediction mode from the video bitstream, configuring a motion vector candidate list based at least in part on the plurality of motion vector predictors in the HMVP table according to a prediction mode, selecting a motion vector predictor for the current CU from the motion vector candidate list, determining a motion vector based at least in part on the selected motion vector predictor and the prediction mode, and updating the HMVP table based on the determined motion vector.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] This application generally relates to the encoding and decoding of video data, and more particularly, to a video coding method and system using history-based motion vector prediction.

Background Art

[0002]

[0002] Digital video is supported by various electronic devices such as digital TVs, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game terminals, smartphones, video conferencing devices, and video streaming devices. Such electronic devices perform transmission, reception, encoding, decoding, and / or storage of digital video data by implementing video compression and expansion standards defined by standards such as MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding). Generally, video compression includes procedures that perform spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. In block-based video coding, a video frame is divided into one or more slices, each slice has a plurality of video blocks, and the video blocks can also be called coding tree units (CTUs). Each CTU may include one coding unit (CU), or may be recursively divided into a plurality of smaller CUs until a predetermined minimum CU size is reached. Each CU (also called a leaf CU) includes one or more transform units (TUs), and each CU also includes one or more prediction units (PUs). Each CU can be encoded in either an intra mode, an inter mode, or an IBC mode. A plurality of video blocks in an intra-coded (I) slice in a video frame are encoded using spatial prediction with respect to reference samples of neighboring blocks in the same video frame.Multiple video blocks within an inter-coded (P or B) slice in a video frame may use spatial prediction for reference samples of neighboring blocks in the same video frame, or may use temporal prediction for reference samples of another reference video frame in the past and / or future.

[0003]

[0003] If spatial prediction or temporal prediction based on a reference block encoded in the past (e.g., a neighboring block) is performed, a prediction block for the current video block to be encoded can be obtained. The process of searching for the reference block may be performed by a block matching algorithm. Residual data indicating the difference in pixels between the current block to be encoded and the prediction block is called a residual block or prediction error. An inter-coded block is encoded according to a motion vector indicating the reference frame forming the prediction block and the residual block. The process of determining the motion vector is generally called motion estimation. An intra-coded block is encoded according to an intra prediction mode and a residual block. For further compression, residual conversion coefficients are obtained by converting the residual block from the pixel domain to a transform domain (e.g., the frequency domain), and then these are quantized. The quantized transform coefficients are first arranged in a two-dimensional layout, scanned to generate a one-dimensional vector consisting of a plurality of transform coefficients, and then entropy coding may be performed to achieve further compression to obtain a video stream.

[0004]

[0004] Next, the encoded video stream is recorded on a computer-readable recording medium (such as a flash memory) that can be accessed by other electronic devices having digital video capabilities, or is directly transmitted to the electronic device, either wired or wirelessly. The electronic device then analyzes the encoded bitstream, for example, to obtain syntax elements from the bitstream, and based on at least a part of the syntax elements obtained from the bitstream, performs video decompression by reconstructing the digital video data from the encoded bitstream into the original format (a process opposite to the above video compression), and then draws the reconstructed digital video data on the display of the electronic device.

[0005]

[0005] As the digital video quality shifts from high definition to 4K×2K or 8K×4K, the amount of video data to be encoded / decoded increases exponentially. There is an ongoing effort in terms of how to encode / decrypt video data more efficiently while maintaining the image quality of the decoded video data.

Summary of the Invention

[0006]

[0006] This application relates to encoding and decoding of video data, and particularly to embodiments of a system and method for parallel processing of video data in video encoding and decoding using history-based motion vector prediction.

[0007]

[0007] The method for decoding video data according to the first aspect of the present application is implemented on an arithmetic unit having one or more processors and a memory for storing a plurality of programs executed by the one or more processors. After acquiring a video bitstream, the arithmetic unit begins by extracting data related to a plurality of encoded pictures from the video bitstream. Each picture includes a plurality of Coding Tree Unit (CTU) rows, and each CTU includes one or more Coding Units (CUs). Before starting to decode the CU at the beginning of the current CTU row to be decoded, the arithmetic unit resets the History-based Motion Vector Predictor (HMVP) table. Then, while decoding the current CTU row, the arithmetic unit prepares a plurality of motion vector predictors in the HMVP table. Each motion vector predictor is one that has been used to decode at least one CU. For the current CU of the current CTU row to be decoded, the arithmetic unit extracts a prediction mode from the video bitstream and constructs a motion vector candidate list based on at least a part of the motion vector predictors in the HMVP table according to the prediction mode. After selecting a motion vector predictor from the motion vector candidate list, the arithmetic unit determines a motion vector based on at least a part of the prediction mode and the selected motion vector predictor, decodes the current CU using the determined motion vector, and updates the HMVP table based on the determined motion vector.

[0008]

[0008] The arithmetic unit according to the second aspect of the present application includes one or more processors, a memory, and a plurality of programs stored in the memory. When the program is executed by one or more processors, it causes the arithmetic unit to execute the above-described processing.

[0009]

[0009] The non-transitory computer-readable recording medium according to the third aspect of the present application stores a plurality of programs for execution by an arithmetic unit having one or more processors. When the program is executed by one or more processors, it causes the arithmetic unit to execute the above-described processing.

[0010]

[0010] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated herein, forming a part of this specification, illustrating the embodiments described and serving to explain the principles underlying them. Corresponding components are denoted by the same reference numerals.

Brief Description of the Drawings

[0011]

Figure 1

[0011] FIG. 1 is a block diagram showing an example of a video encoding / decoding system according to some embodiments of the present disclosure.

Figure 2

[0012] FIG. 2 is a block diagram showing an example of a video encoder according to some embodiments of the present disclosure.

Figure 3

[0013] FIG. 3 is a block diagram showing an example of a video decoder according to some embodiments of the present disclosure.

Figure 4A

[0014] FIG. 4A is a block diagram showing a method of recursively quad-tree splitting a frame into a plurality of video blocks having a plurality of different sizes according to some embodiments of the present disclosure.

Figure 4B

Figure 4C

Figure 4D

Figure 5A

[0015] FIG. 5A is a block diagram showing block positions that are spatially adjacent and temporally connected to a current CU to be encoded according to some embodiments of the present disclosure.

Figure 5B

[0016] FIG. 5B is a block diagram showing multi-threaded encoding of a plurality of CTU rows in a picture using wavefront parallel processing according to some embodiments of the present disclosure.

Figure 6

[0017] FIG. 6 is a flowchart showing an example of processing by a video coder that implements a technique for constructing a candidate list of motion vector predictors according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0012]

[0018] Specific embodiments will be referred to in detail below, and examples of such embodiments are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to those skilled in the art that various modifications can be made without departing from the scope of the claims, and that the subject matter can be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented in various electronic devices having digital video capabilities.

[0013]

[0019] FIG. 1 is a block diagram showing an exemplary system 10 for parallel encoding and decoding of video blocks according to some embodiments of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data that will be decoded at a destination device 14 at a later time. The source device 12 and the destination device 14 may comprise a wide variety of any electronic devices including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game terminals, or video streaming devices. In some embodiments, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.

[0014]

[0020] In some embodiments, the destination device 14 may receive the encoded video data to be decoded via the link 16. The link 16 may comprise any type of communication medium or device that transfers the encoded video data from the source device 12 to the destination device 14. As one example, the link 16 may comprise a communication medium that enables the source device 12 to transmit the encoded video data directly to the destination device 14 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium may comprise any wireless or wired communication medium such as a radio frequency (RF) or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may also include routers, switches, base stations, or other devices used to enable communication from the source device 12 to the destination device 14.

[0015]

[0021] In some other embodiments, the encoded video data may be transmitted from the output interface 22 to the recording device 32. Thereafter, the encoded video data within the recording device 32 may be accessed by the destination device 14 via the input interface 28. The recording device 32 may include any of various data recording media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital recording medium, which may be decentralized or locally accessible. As yet another example, the recording device 32 may correspond to a file server or other intermediate recording device that can hold the encoded video data generated at the information source device 12. The destination device 14 may access the recorded video data through streaming or downloading from the recording device 32. The file server may be any kind of computer having a function of storing the encoded video data and transmitting the encoded video data to the destination device 14. Examples of the file server include a web server (for example, for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data through any standard data connection suitable for accessing the encoded video data recorded on the file server, including a wireless channel (for example, Wi-Fi connection), a wired connection (for example, DSL or cable media, etc.), or a combination of both. The transmission of the encoded video data from the recording device 32 may be a streaming transmission, a download transmission, or a combination of both.

[0016]

[0022] As shown in FIG. 1, the information source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 may include, for example, a video camera, a video archive that stores previously acquired video, a video supply interface that receives video from a video content provider, and / or a video capture device such as a computer graphics system that generates computer graphics as the video of the information source (source). As an example, when the video source 18 is a video camera of a security monitoring system, the information source device 12 and the destination device 14 may constitute a camera phone or a video phone. However, the embodiments described in this application are generally applicable to video coding and may be applied to wireless and / or wired applications.

[0017]

[0023] The video thus acquired or previously acquired, or generated by a computer, may be encoded by the video encoder 20. The encoded video data may be directly transmitted to the destination device 14 through the output interface 22 of the information source device 12. Also (or alternatively), the encoded video data may be recorded in the recording device 32 for later access from the destination device 14 or other devices, for decoding and / or playback. Further, the output interface 22 may include a modem and / or a transmitter.

[0018]

[0024] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 may include a receiver and / or a modem and receive the encoded video data via the link 16. The encoded video data communicated via the link 16, or the encoded video data supplied by the recording device 32, may include various syntax elements generated by the video encoder 20 for use in decoding the video data by the video decoder 30. Such syntax elements may be incorporated into the encoded video data transmitted over the communication medium, recorded on the recording medium, or recorded on the file server.

[0019]

[0025] In some embodiments, the destination device 14 may include a display device 34, which can be an integrated display device or an external display device configured to communicate with the destination device 14. The display device 34 is for displaying the decoded video data to the user, and may be equipped with various arbitrary display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.

[0020]

[0026] The video encoder 20 and the video decoder 30 may operate based on intellectual property or industry standards such as VVC, HEVC, MPEG-4 Part10 AVC (Advanced Video Coding), or extended versions of these standards. It should be understood that this application is not limited to specific video encoding / decoding standards and may be applicable to other video encoding / decoding standards. It is usually considered that the video encoder 20 of the source device 12 may be configured to encode video data according to any of the current or future standards. Similarly, it is usually considered that the video decoder 30 of the destination device 14 may be configured to decode video data according to any of the current or future standards.

[0021]

[0027] The video encoder 20 and the video decoder 30 may each be implemented as any of a variety of suitable encoding circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or combinations thereof. When implemented partially in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding processes disclosed in this disclosure. Each of the video encoder 20 and the video decoder 30 may be incorporated into one or more encoders or decoders, and either the encoder or the decoder may be integrated into each device as part of a combined encoder / decoder (CODEC).

[0022]

[0028] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 can perform intra prediction encoding and inter prediction encoding of video blocks within a video frame. Intra prediction encoding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter prediction encoding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures in a video sequence.

[0023]

[0029] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a conversion processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some embodiments, the video encoder 20 further includes an inverse quantization unit 58, an inverse conversion processing unit 60, and an adder 62 for reconstructing video blocks. A deblocking filter (not shown) may be arranged between the adder 62 and the DPB 64 to remove block distortion from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) may be used to filter the output of the adder 62. The video encoder 20 may be in the form of a non-changeable or programmable hardware unit, or may be divided among one or more non-changeable or programmable hardware units as shown.

[0024]

[0030] The video data memory 40 may store the video data to be encoded by the components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores the reference video data used when encoding video data by the video encoder 20. The video data memory 40 and the DPB 64 may be composed of any various memory devices. In various examples, the video data memory 40 may be implemented on a chip together with other components of the video encoder 20, or may be implemented outside the chip with respect to other components.

[0025]

[0031] As shown in FIG. 2, after receiving the video data, the splitting unit 45 in the prediction processing unit 41 splits the video data into a plurality of video blocks. This splitting may include splitting the video frame into a plurality of slices or a plurality of larger coding units (CUs) according to a predetermined splitting structure such as a quadtree structure associated with the video data. The video frame may be split into a plurality of video blocks (or a plurality of sets of video blocks referred to as tiles). The prediction processing unit 41 may select one of a plurality of selectable prediction coding modes, such as one of a plurality of intra-prediction coding modes or one of a plurality of inter-prediction coding modes, for the current video block based on error results (e.g., coding rate and distortion level). The prediction processing unit 41 supplies the obtained intra-prediction coded block or inter-prediction coded block to the adder 50 to generate a residual block, and also supplies it to the adder 62 to reconstruct a coded block that will be used as part of the reference frame later. Further, the prediction processing unit 41 supplies syntax elements such as motion vectors, intra-mode predictors, splitting information, and other syntax information to the entropy coding unit 56.

[0026]

[0032] To select an appropriate intra-prediction coding mode for the current video block, the intra-prediction processing unit 46 in the prediction processing unit 41 may perform intra-prediction coding of the current video block for one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter-prediction coding of the current video block for one or more prediction blocks in one or more reference frames to provide temporal prediction. The video encoder 20 may execute a plurality of coding procedures, for example, to select an appropriate coding mode for each block of the video data.

[0027]

[0033] In some embodiments, the motion estimation unit 42 determines the inter-prediction coding mode for the current video frame by generating a motion vector indicating the displacement of the position of the prediction unit (PU) of the video block in the current video frame with respect to the prediction block in the reference frame according to a predetermined pattern in the sequence of video frames. Motion estimation is performed by the motion estimation unit 42 and is a process of generating a motion vector for estimating the motion of a video block. The motion vector may indicate, for example, the displacement of the position of the PU of the video block in the current video frame or picture with respect to the prediction block in the reference frame (or in another coding unit) with respect to the current block to be coded in the current frame (or in another coding unit). The above-mentioned predetermined pattern may be any pattern that designates a plurality of video frames in the sequence as P-frames or B-frames. The intra BC unit 48 may determine a vector such as a block vector for intra BC coding in the same manner as the determination of the motion vector by the motion estimation unit 42 for inter-prediction, or may determine the block vector using the motion estimation unit 42.

[0028]

[0034] The prediction block is a block in the reference frame that is considered to closely correspond to the PU of the video block to be coded from the perspective of pixel difference, and may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference reference quantities. In some embodiments, the video encoder 20 may calculate the values for the sub-integer pixel positions of the reference frame stored in the DPB 64. For example, the video encoder 20 may interpolate the values at the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference frame. Therefore, the motion estimation unit 42 may perform a motion search for the overall pixel position and fractional pixel position and output a motion vector with fractional pixel accuracy.

[0029]

[0035] The motion estimation unit 42 calculates a motion vector by comparing the position of the prediction block in the reference frame selected from the first reference frame list (list 0) or the second reference frame list (list 1) with the position of the PU of the video block in the inter-predicted coded frame. Here, the first reference frame list or the second reference frame list identifies one or more reference frames stored in the DPB 64, respectively. The motion estimation unit 42 transfers the calculated motion vector to the motion compensation unit 44 and then to the entropy encoding unit 56.

[0030]

[0036] Motion compensation is performed by the motion compensation unit 44 and may include fetching or generating a prediction block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, the motion compensation unit 44 finds the prediction block indicated by the motion vector in one of the plurality of reference frame lists, acquires the prediction block, and sends the prediction block to the adder 50. Then, the adder 50 forms a residual video block consisting of pixel difference values by subtracting the pixel values of the prediction block supplied from the motion compensation unit 44 from the pixel values of the current video block to be encoded. The pixel difference values constituting the residual video block may include a luma difference component, a chroma difference component, or both. Also, the motion compensation unit 44 may generate syntax elements associated with the video blocks of the video frame for use when decoding the video blocks of the video frame by the video decoder 30. The syntax elements may include, for example, a prediction vector used to identify the prediction block, some flag indicating the prediction mode, or any other syntax element defining some other syntax information described herein. It should be noted that the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are separated and illustrated for conceptual purposes.

[0031]

[0037] In some embodiments, the intra BC unit 48 may generate vectors and fetch prediction blocks in the same manner as described above with respect to the motion estimation unit 42 and the motion compensation unit 44. Here, the prediction block exists within the same frame as the current block to be encoded, and the vector is regarded as a block vector in the opposite direction to the motion vector. In particular, the intra BC unit 48 may determine an intra prediction mode for use in encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, in separate encoding procedures, and may evaluate their performance by rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode from among the various evaluated intra prediction modes to generate an intra mode predictor to be used. For example, the intra BC unit 48 may calculate rate-distortion values using rate-distortion analysis for the various evaluated intra prediction modes, and may select, as the appropriate intra prediction mode to be used, the intra prediction mode having the best rate-distortion characteristics from among the evaluated modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original block before encoding that was encoded to generate the encoded block, along with the bit rate (i.e., a large number of bits) used to generate these encoded blocks. The intra BC unit 48 calculates a ratio from the distortion and calculates the rates for the various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0032]

[0038] In other examples, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such a function for intra BC prediction according to the embodiments described herein. In any case, for intra block copy, the prediction block may be a block that is considered to closely correspond to the block to be coded from the perspective of pixel difference, and may be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference reference quantities. And for the identification of the prediction block, it is sufficient if the calculation of the values regarding the sub-integer pixel positions is included.

[0033]

[0039] Whether the prediction block is obtained from the same frame according to intra prediction or from different frames according to inter prediction, the video encoder 20 may form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block to be coded to form pixel difference values. The pixel difference values forming the residual video block may include the difference of the luminance component and the difference of the chroma component.

[0034]

[0040] The intra prediction processing unit 46 may perform intra prediction of the current video block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, or the intra block copy prediction performed by the intra BC unit 48. In particular, the intra prediction processing unit 46 may determine the intra prediction mode to be used for coding the current block. To do so, the intra prediction processing unit 46 may code the current block using various intra prediction modes, for example, in separate coding procedures, and the intra prediction processing unit 46 (or, in some examples, the mode selection unit) may select an appropriate intra prediction mode to be used from among the evaluated intra prediction modes. The intra prediction processing unit 46 may supply information indicating the intra prediction mode selected for the block to the entropy coding unit 56. The entropy coding unit 56 may code the information indicating the selected intra prediction mode into the bitstream.

[0035]

[0041] After the prediction processing unit 41 determines a prediction block for the current video block by either inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be incorporated into one or more transform units (TUs) and supplied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0036]

[0042] The transform processing unit 52 transfers the residual transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. In that quantization process, the bit depth for some or all of the coefficients may be reduced. The quantization degree may be changed by adjusting the quantization parameter. In some examples, next, the quantization unit 54 may perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform that scan.

[0037]

[0043] Subsequent to quantization, entropy encoding unit 56 performs entropy encoding on the quantized transform coefficients using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy coding (PIPE), or other entropy coding methodologies or techniques to obtain a video bitstream. The encoded bitstream may then be transmitted to video decoder 30 or recorded in recording device 32 so as to be transmitted or retrieved by video decoder 30 at a later time. Entropy encoding unit 56 may perform entropy encoding on the motion vector and other syntax elements related to the current video frame to be encoded.

[0038]

[0044] Inverse quantization unit 58 and inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual video block in the pixel domain to generate a reference block for prediction of other video blocks. As described above, motion compensation unit 44 may generate a motion-compensated prediction block from one or more reference blocks of the frames stored in DPB 64. Motion compensation unit 44 may apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values used for motion compensation.

[0039]

[0045] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reference block and stores this in the DPB 64. Next, the reference block may be used by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 as a prediction block for performing inter prediction on other video blocks within subsequent video frames.

[0040]

[0046] FIG. 3 is a block diagram illustrating a video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoder 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. Further, the prediction processing unit 81 includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process generally inverse to the encoding process described above with respect to the video encoder 20 related to FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on the motion vector acquired from the entropy decoder 80, while the intra prediction processing unit 84 may generate prediction data based on the intra prediction mode indicator acquired from the entropy decoder 80.

[0041]

[0047] In some examples, one unit within the video decoder 30 may receive an assignment of tasks for implementing the embodiments of the present application. In some examples, the embodiments of the present disclosure may be divided among one or more units within the video decoder 30. For example, the intra BC unit 85 may implement the embodiments of the present application alone or in cooperation with other units such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoder 80. In some examples, the video decoder 30 may not have an intra BC unit 85, and the functionality of the intra BC unit 85 may be performed by other components such as the motion compensation unit 82 within the prediction processing unit 81.

[0042]

[0048] The video data memory 79 may store video data such as an encoded video bitstream to be decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the recording device 32, or from a local video source such as a camera by wired network communication or wireless network communication of video data, or from a physical data recording medium (for example, a flash drive or a hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 in the video decoder 30 stores reference video data used when the video decoder 30 decodes video data (for example, in an intra prediction coding mode or an inter prediction coding mode). The video data memory 79 and the DPB 92 may be constituted by various arbitrary memory devices such as a synchronous DRAM (SDRAM), a magneto-resistive RAM (MRAM), a resistive RAM (RRAM), or other types of memory devices, such as a dynamic random access memory (DRAM). For the purpose of illustration, the video data memory 79 and the DPB 92 are shown as two different components within the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 may be implemented on the chip together with other components within the video decoder 30, or may be implemented outside the chip with respect to other components.

[0043]

[0049] During the decoding process, the video decoder 30 receives an encoded bitstream indicating a plurality of video blocks of the encoded video frame and the syntax elements associated therewith. The video decoder 30 may receive the syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 in the video decoder 30 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Then, the entropy decoding unit 80 sends the motion vectors and other syntax elements to the prediction processing unit 81.

[0044]

[0050] When the video frame is encoded as an intra prediction encoded (I) frame or for an intra encoded prediction block within another type of frame, the intra prediction processing unit 84 in the prediction processing unit 81 may generate prediction data for the video blocks in the current video frame based on the signaled intra prediction mode and reference data from past decoded blocks in the current frame.

[0045]

[0051] When the video frame is encoded as an inter prediction encoded (e.g., B or P) frame, the motion compensation unit 82 in the prediction processing unit 81 generates one or more prediction blocks for the video blocks in the current video frame based on the motion vectors and other syntax elements obtained from the entropy decoding unit 80. Each of the prediction blocks may be generated from a reference frame within one of a plurality of reference frame lists. The video decoder 30 may configure reference frame lists 0 and 1 using an initial configured technique based on the reference frames stored in the DPB 92.

[0046]

[0052] In some examples, when a video block is encoded according to the intra BC mode described in this specification, the intra BC unit 85 in the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements obtained from the entropy decoding unit 80. The prediction block may be present in the reconstruction area within the same picture as the current video block defined by the video encoder 20.

[0047]

[0053] The motion compensation unit 82 and / or the intra BC unit 85 analyzes the motion vector and other syntax elements to determine prediction information for a video block in the current video frame, and then uses the prediction information to generate a prediction block for the current video block to be decoded. For example, the motion compensation unit 82 uses some of the obtained syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) used for encoding the video block in the video frame, the type of inter prediction frame (e.g., B or P), the configuration information for one or more reference frame lists for the frame, the motion vector for each of the inter prediction encoded video blocks in the frame, the inter prediction state for each of the inter prediction encoded video blocks in the frame, and other information in order to decode the video block in the current video frame.

[0048]

[0054] Similarly, the intra BC unit 85 uses some of the obtained syntax elements (e.g., flags) to determine that the current video block is predicted using the intra BC mode, the configuration information indicating which video blocks in the frame are present in the reconstruction area and should be stored in the DPB 92, the block vector for each of the intra BC prediction video blocks in the frame, the intra BC prediction state for each of the intra BC prediction video blocks in the frame, and other information in order to decode the video block in the current video frame.

[0049]

[0055] The motion compensation unit 82 may perform interpolation using an interpolation filter such as that used by the video encoder 20 during the encoding of the video block in order to calculate the interpolation value for the sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the acquired syntax element, and use the interpolation filter to generate the prediction block.

[0050]

[0056] The inverse quantization unit 86 performs inverse quantization on the quantized transform coefficients given in the bitstream and entropy-decoded by the entropy decoder 80, using the same quantization parameter calculated by the video encoder 20 for each video block in the video frame to determine the quantization degree. The inverse transform processing unit 88 applies an inverse transform, such as, for example, an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0051]

[0057] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vector and other syntax elements, the adder 90 adds the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85 to reconstruct the decoded video block for the current video block. A loop filter (not shown) may be arranged between the adder 90 and the DPB 92 to further process the decoded video block. Next, the decoded video blocks within a specific frame are stored in the DPB 92. The DPB 92 stores the reference frames used for subsequent motion compensation of the next video block. For subsequent display on a display device such as the display device 34 of FIG. 1, the DPB 92 or a memory device separate from the DPB 92 may store the decoded video.

[0052]

[0058] In general video decoding processing, a video sequence generally includes a set of a plurality of frames or pictures arranged. Each frame may include three sample arrays indicated by SL, SCb, and Scr. SL is a two-dimensional array consisting of a plurality of luma samples. SCb is a two-dimensional array consisting of a plurality of Cb chroma samples. SCr is a two-dimensional array consisting of a plurality of Cr chroma samples. In other examples, a frame may be monochromatic, thereby having only a single two-dimensional array consisting of a plurality of luma samples.

[0053]

[0059] As shown in FIG. 4A, the video encoder 20 (or, more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). The video frame may simply include an integer number of CTUs that are continuously ordered in raster scan order from top to bottom and from left to right. Each CTET is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set so that all CTUs in the video sequence have the same size as one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As shown in FIG. 4B, each CTU may have one coding tree block (CTB) consisting of a plurality of luminance samples, two coding tree blocks consisting of a plurality of chrominance samples corresponding thereto, and syntax elements used to code the samples of those coding tree blocks. The syntax elements describe the characteristics of different types of units of the coded pixel block and the way in which the video sequence can be reconstructed by the video decoder 30, and include inter prediction or intra prediction, intra prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture having three different color planes, the CTU may simply have a single coding tree block and syntax elements used to code the samples of the coding tree block. The coding tree block may be an N×N block consisting of a plurality of samples.

[0054]

[0060] To achieve better performance, the video encoder 20 may recursively perform a tree split, such as a binary-tree split, a quad-tree split, or a combination of both, on the coding tree blocks of the CTU to split the CTU into a plurality of smaller coding units (CUs). As shown in FIG. 4C, a 64×64 CTU 400 is first split into four smaller CUs, each CU having a block size of 32×32. Of these four smaller CUs, each of CU410 and CU420 is split into four CUs with a block size of 16×16. The two 16×16 CUs 430, 440 are each further split into four CUs with a block size of 8×8. FIG. 4D is a diagram of a quad-tree data structure representing the final result obtained by subjecting the CTU 400 shown in FIG. 4C to a splitting process. Each leaf node of the quad-tree corresponds to one CU having each size within the range from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU has a coding block (CB) consisting of a plurality of luminance samples in a frame of the same size and two coding blocks consisting of a corresponding plurality of chrominance samples, and syntax elements used to encode the samples of these coding blocks. For a monochrome picture or a picture with three different color planes, the CU may have a single coding block and a syntax structure used to encode the samples of the coding block.

[0055]

[0061] In some embodiments, the video encoder 20 may further divide the encoding block of the CU into one or more N×N prediction blocks (PBs). The prediction block is a rectangular (square or non-square) block consisting of a plurality of samples to which the same prediction, i.e., inter or intra, is applied. The prediction unit (PU) of the CU may have a prediction block consisting of a plurality of luminance samples, two prediction blocks consisting of a plurality of corresponding chrominance samples, and the syntax elements used to predict those prediction blocks. In a monochrome picture or a picture having three different color planes, the PU may have a single prediction block and the syntax structure used to predict the prediction block. The video encoder 20 may generate a predicted luminance block, a predicted Cb block, and a predicted Cr block for the luminance prediction block, the Cb prediction block, and the Cr prediction block in each PU of the CU.

[0056]

[0062] The video encoder 20 may generate those prediction blocks for the PUs using intra prediction or inter prediction. When the video encoder 20 uses intra prediction to generate the prediction blocks of the PUs, the video encoder 20 may generate the prediction blocks of the PUs based on the decoded samples of the frame related to the PUs. When the video encoder 20 uses inter prediction to generate the prediction blocks of the PUs, the video encoder 20 may generate the prediction blocks of the PUs based on the decoded samples of one or more frames other than the frame related to the PUs.

[0057]

[0063] After the video encoder 20 generates a predicted luminance block, a predicted Cb block, and a predicted Cr block for one or more PUs of a CU, the video encoder 20 may generate a luminance residual block for the CU by subtracting the predicted luminance block of the CU from its original luminance encoding block, where each sample of the luminance residual block of the CU indicates the difference between the luminance sample of one block among the plurality of predicted luminance blocks of the CU and the corresponding sample of the original luminance encoding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block, respectively, where each sample of the Cb residual block of the CU indicates the difference between the Cb sample of one block among the plurality of predicted Cb blocks of the CU and the corresponding sample of the original Cb encoding block of the CU, and each sample of the Cr residual block of the CU may indicate the difference between the Cr sample of one block among the plurality of predicted Cr blocks of the CU and the corresponding sample of the original Cr encoding block of the CU.

[0058]

[0064] Furthermore, as shown in FIG. 4C, the video encoder 20 may use quadtree partitioning to decompose the luminance residual block, Cb residual block, and Cr residual block in the CU into one or more luminance transform blocks, Cb transform blocks, and Cr transform blocks. The transform block is a rectangular (square or non-square) block consisting of a plurality of samples to which the same transform is applied. The transform unit (TU) of the CU may have a transform block consisting of a plurality of luminance samples, two transform blocks consisting of a plurality of corresponding luminance samples, and syntax elements used to transform the samples of those transform blocks. Thus, each TU of the CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with the TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. For a monochrome picture or a picture having three different color planes, the TU may have a single transform block and a syntax structure used to transform the samples of the transform block.

[0059]

[0065] The video encoder 20 may apply one or more transforms to the luminance transform block of the TU to generate a luminance coefficient block for the TU. The coefficient block may be a two-dimensional array of a plurality of transform coefficients. The transform coefficient may be a scalar quantity. The video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block for the TU. The video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block for the TU.

[0060]

[0066] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Usually, quantization refers to the process of quantizing the transform coefficients so as to be able to reduce the total amount of data used to represent the transform coefficients, and it gives further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy code the syntax element indicating the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax element indicating the quantized transform coefficients. Finally, the video encoder 20 may output a bitstream indicating the bit sequence that constitutes the representation of the encoded frame and related data, which is stored in the recording device 32 or transmitted to the destination device 14.

[0061]

[0067] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may analyze the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data based on at least a part of the syntax elements obtained from the bitstream. The process of reconstructing video data is usually a process inverse to the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on the coefficient block related to the TU of the current CU to reconstruct the residual block related to the TU of the current CU. The video decoder 30 may also reconstruct the encoded block of the current CU by adding the samples of the prediction block for the PU in the current CU to the corresponding samples of the transform block of the TU in the current CU. After reconstructing the encoded block for each CU of the frame, the video decoder 30 may reconstruct the frame.

[0062]

[0068] As described above, video encoding realizes video compression using two main modes, namely, intra-frame prediction (or intra prediction) and inter-frame prediction (or inter prediction). It should be noted that IBC can be regarded as either intra-frame prediction or a third mode. Among these two modes, inter-frame prediction contributes more to encoding efficiency than intra-frame prediction because it uses motion vectors to predict the current video block from reference video blocks.

[0063]

[0069] However, with the continuous improvement of video data capture technology and the finer video block size to retain detailed information of video data, the amount of data required to represent motion vectors for the current frame has also substantially increased. One way to overcome this problem is to take advantage of the fact that not only does a set of neighboring CUs have similar video data for prediction in both the spatial and temporal domains, but also the motion vectors between these neighboring CUs are similar. Therefore, by detecting spatial and temporal correlations, it is possible to use the motion information of spatially neighboring CUs and / or temporally connected CUs as approximate motion information (e.g., motion vectors) of the current CU. This is also called the "motion vector predictor (MVP)" of the current CU.

[0064]

[0070] Instead of encoding the actual motion vector of the current CU determined by the motion estimation unit 42 as described above with respect to FIG. 2 into the video bitstream, in order to generate a motion vector difference (MVD) for the current CU, the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU. By doing so, it is not necessary to encode the motion vector determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent the motion information in the video bitstream can be significantly reduced.

[0065]

[0071] Similar to the process of selecting a prediction block in the reference frame during inter-frame prediction of a coded block, a set of rules needs to be applied in both the video encoder 20 and the video decoder 30 to construct a motion vector candidate list for the current CU using candidate motion vectors that can be generated in relation to spatially neighboring CUs and / or temporally adjacent CUs for the current CU, and then to select one element from that motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to transmit the motion vector candidate list itself between the video encoder 20 and the video decoder 30, and the index of the selected motion vector predictor in the motion vector candidate list is sufficient for the video encoder 20 and the video decoder 30 to use the same motion vector predictor in the motion vector candidate list for encoding and decoding the current CU.

[0066]

[0072] In some embodiments, each inter-prediction CU has three motion vector prediction modes, including inter (hereinafter also referred to as "advanced motion vector prediction (AMVP)"), skip, and merge, for constructing a motion vector candidate list. Under each mode, one or more motion vector candidates can be added to the motion vector candidate list according to the algorithms described below. Finally, one candidate is selected from the candidate list as the best motion vector predictor for the inter-prediction CU to be encoded into the video bitstream by the video encoder 20 or decoded from the video bitstream by the video decoder 30. To find the best motion vector predictor from the candidate list, a motion vector competition (MVC) method is introduced, which selects a motion vector from a candidate set consisting of a plurality of given motion vectors, that is, a motion vector candidate list including spatial and temporal motion vector candidates.

[0067]

[0073] In addition to extracting candidates for motion vector predictors from spatially neighboring or temporally connected CUs, candidates for the motion vector predictors can also be extracted from a so-called "history-based motion vector prediction (HMVP)" table. The HMVP table stores a predetermined number of motion vector predictors, and each motion vector predictor is used to encode / decode a specific CU in the same CTU row (or sometimes the same CTU). Due to the spatial / temporal proximity of these CUs, there is a high probability that one motion vector predictor in the HMVP table will be reused to encode / decode different CUs within the same CTU row. Therefore, by incorporating the HMVP table into the process of constructing the motion vector candidate list, higher coding efficiency can be achieved.

[0068]

[0074] In some embodiments, the HMVP table has a fixed length (e.g., 5) and is managed in a quasi-FIFO (quasi-First-In-First-Out) manner. For example, a motion vector is reconstructed when decoding one of the inter-coded blocks of a CU for that CU. The HMVP table is updated on-the-fly with the reconstructed motion vector because such a motion vector can be the motion vector predictor for the next CU. When updating the HMVP table, there are two scenarios as follows: (i) the reconstructed motion vector is different from other existing motion vectors in the HMVP table, or (ii) the reconstructed motion vector is the same as one of the existing motion vectors in the HMVP table. For the first scenario, if the HMVP table is not full, the reconstructed motion vector is added to the HMVP table as the latest motion vector. If the HMVP table is already full, the oldest motion vector in the HMVP table is deleted before the reconstructed motion vector is added as the latest motion vector. In other words, in this case, the HMVP table is similar to a FIFO buffer, and the motion vector information located at the head of this FIFO buffer and related to other past inter-coded blocks is shifted out of the buffer, so that the reconstructed motion vector is added to the rear of the FIFO buffer as the latest element in the HMVP table. For the second scenario, before the reconstructed motion vector is added as the latest motion vector, the existing motion vector in the HMVP table that is substantially the same as the reconstructed motion vector is deleted from the HMVP table. Also, if the HMVP table is maintained in the form of a FIFO buffer, the motion vector predictors after the same motion vector in the HMVP table are shifted forward by one element to occupy the space remaining after the deleted motion vector, and then the reconstructed motion vector is added to the rear of the FIFO buffer as the latest element in the HMVP table.

[0069]

[0075] The motion vectors in the HMVP table can be added to the motion vector candidate list under multiple different prediction modes such as AMVP, merge, and skip. It has been found that the motion information of past inter-coded blocks that are not even adjacent to the current block and are stored in the HMVP table can be used for more efficient motion vector prediction.

[0070]

[0076] After one of the MVP candidates in a given candidate set consisting of multiple motion vectors is selected for the current CU, the video encoder 20 generates one or more syntax elements for the corresponding MVP candidate and encodes these into the video bitstream so that the video decoder 30 can obtain the MVP candidate from the video bitstream using the syntax elements. Depending on the specific mode used to construct the motion vector candidate set, multiple different modes (e.g., AMVP, merge, skip, etc.) have multiple different sets of syntax elements. For the AMVP mode, the syntax elements include an inter-prediction indicator (list 0, list 1, or bi-directional prediction), a reference index, a motion vector candidate index, a motion vector prediction residual signal, etc. For the skip mode and the merge mode, only the merge index is encoded in the bitstream. This is because the current CU inherits other syntax elements including the inter-prediction indicator, reference index, and motion vector from a neighboring CU referenced by the encoded merge index. In the case of a skip-encoded CU, the motion vector prediction residual signal is also ignored.

[0071]

[0077] FIG. 5A is a block diagram showing block positions that are spatially adjacent and temporally connected to the current CU to be encoded / decoded according to some embodiments of the present disclosure. For a given mode, first, the availability of motion vectors associated with block positions that are spatially adjacent to the left and above and the availability of motion vectors associated with block positions that are temporally connected are checked, and then a motion vector prediction (MVP) candidate list is constructed by checking the motion vectors in the FDMVP table. During the process of constructing the MVP candidate list, some overlapping MVP candidates are removed from the candidate list, and if necessary, zero-valued motion vectors are added to create a candidate list with a fixed length (note that multiple different modes may have different fixed lengths). After constructing the MVP candidate list, the video encoder 20 can select the best motion vector predictor from the candidate list and encode the corresponding index indicating the selected candidate into the video bitstream.

[0072]

[0078] Taking FIG. 5A as an example, assuming that the candidate list has a fixed length of 2, the motion vector predictor (MVP) candidate list for the current CU may be constructed by sequentially executing the following steps under the AMVP mode. 1) Selection of MVP candidates from a plurality of spatially adjacent CUs a) Extracting one unscaled MVP candidate from one of two CUs that are spatially adjacent to the left, starting from A0 and ending at A1. b) When there is no available unscaled MVP candidate from the left in the previous step, extracting one scaled MVP candidate from one of two CUs that are spatially adjacent to the left, starting from A0 and ending at A1. c) Extracting one unscaled MVP candidate from one of three CUs that are spatially adjacent to the above, starting from B0, passing through B1, and ending at B2. d) When neither A0 nor A1 is available, or when they are encoded in the intra mode, extract one scaled MVP candidate from one of the three spatially adjacent CUs above, starting from B0, passing through B1, and ending at B2. 2) When two MVP candidates are found in the previous step and they are the same, delete one of the two candidates from the MVP candidate list. 3) Selection of MVP candidates from multiple temporally concatenated CUs a) When the MVP candidate list after the previous step does not contain two MVP candidates, extract one MVP candidate from the temporally concatenated CUs. 4) Selection of MVP candidates from the HMVP table a) When the MVP candidate list after the previous step does not contain two MVP candidates, extract two history-based MVPs from the HMVP table. 5) When the MVP candidate list after the previous step does not contain two MVP candidates, add two zero-value MVPs to the MVP candidate list.

[0073]

[0079] Since there are only two candidates in the AMVP mode MVP candidate list configured as described above, in order to indicate which of the two MVP candidates in the candidate list is used for decoding the current CU, a related syntax element such as a two-value flag is encoded in the bitstream.

[0074]

[0080] In some embodiments, in skip mode or merge mode, the MVP candidate list may be configured for the current CU by sequentially executing a similar set of steps as described above. For skip mode or merge mode, it should be noted that a special type of merge candidate called "pair-wise merge candidate" is also incorporated into the MVP candidate list. The pair-wise merge candidate is generated by averaging a plurality of MVs in two previously extracted merge mode motion vector candidates. The size of the merge MVP candidate list (e.g., 1 to 6) is signaled in the slice header of the current CU. In merge mode, for each CU, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first bin of the merge index is encoded in context, and bypass encoding is used for the other bins.

[0075]

[0081] As described above, the history-based MVP can be added to either the AMVP mode MVP candidate list or the merge MVP candidate list after the spatial MVP or the temporal MVP. The motion information of past inter-coded CUs is stored in the HMVP table and used as an MVP candidate for the current CU. The HMVP table is maintained during the encoding / decoding process. Whenever there is a non-sub-block inter-coded CU (when the HMVP table is already full and there is no duplication of the same associated motion vector information), the associated motion vector information is added as a new candidate to the last entry of the HMVP table, while the motion vector information stored in the first entry of the HMVP table is deleted from there. Alternatively, before the associated motion vector information is added to the last entry of the HMVP table, the duplication of the same associated motion vector information is removed from the table.

[0076]

[0082] As described above, Intra Block Copy (IBC) can significantly improve the encoding efficiency of display content materials. Since the IBC mode is implemented as a block-level encoding mode, block matching (BM) is performed in the video encoder 20 to find the optimal block vector for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, and it is already reconstructed within the current picture. The IBC-encoded CU is treated as a third prediction mode other than the intra prediction mode or the inter prediction mode.

[0077]

[0083] At the CU level, the IBC mode can be signaled as the following IBC·AMVP mode or IBC·skip / merge mode. - IBC·AMVP mode: The block vector difference (BVD) between the actual block vector of the CU and the block vector predictor of the CU selected from the block vector candidates of the CU is encoded in the same way as when encoding the motion vector difference in the above AMVP mode. The block vector prediction method uses two block vector candidates as predictors, one obtained from the left neighbor and the other from the upper neighbor. When neither neighbor is available, the default block vector is used as the block vector predictor. A binary flag is signaled to indicate the index of the block vector predictor. The IBC·AMVP candidate list is composed of spatial HMVP candidates. - IBC·skip / merge mode: A merge candidate index is used to indicate which of the block vector candidates in the merge candidate list from neighboring IBC-encoded blocks is used for predicting the block vector for the current block. The IBC merge candidate list is composed of spatial, HMVP, and pairwise candidates.

[0078]

[0084] Another approach to improving the coding efficiency to adapt to conventional coding standards is to introduce parallel processing into video encoding / decoding processes, for example, by using a multi-core processor. For example, wavefront parallel processing (WPP) has already been introduced into HEVC as a function that uses multiple threads to encode or decode multiple CTU rows in parallel.

[0079]

[0085] FIG. 5B is a block diagram showing multi-threaded encoding of multiple CTU rows within a picture using wavefront parallel processing (WPP) according to some embodiments of the present disclosure. When WPP is enabled, it is possible to process multiple CTU rows in parallel by the wavefront method. Here, there may be a delay of two CTU rows between the heads of two adjacent wavefronts. For example, to encode picture 500 using WPP, a video coder such as video encoder 20 and video decoder 30 divides the coding tree units (CTUs) of picture 500 into multiple wavefronts, and each wavefront corresponds to each CTU row within the picture. The video coder may start coding the leading wavefront, for example, using a first coder core or thread. After the video coder has coded two or more CTUs of the leading wavefront, the video coder may start coding the second wavefront from the leading wavefront in parallel with the coding of the leading wavefront, for example, using a second parallel coder core or thread. After the video coder has coded two or more CTUs of the second wavefront from the leading wavefront, the video coder may start coding the third wavefront from the leading wavefront in parallel with the coding of higher-order wavefronts, for example, using a third parallel coder core or thread. This pattern may be continued for multiple wavefronts within picture 500. In the present disclosure, a set of CTUs for which the video coder performs coding simultaneously in parallel is referred to as a "CTU group". Thus, when the video coder encodes a picture using WPP, each CTU of the CTU group may belong to only one wavefront within the picture, and the CTU may be offset by at least two CTU columns within the picture from the CTUs in the wavefront one level above.

[0080]

[0086] The video coder may initialize the context for the current wavefront and execute context adaptive binary arithmetic coding (CABAC) for the current wavefront based on the data of the first two blocks of the upper wavefront and one or more elements in the slice header for the slice including the first coded block of the current wavefront. The video coder may perform CABAC initialization for the next wavefront (or CTU row) using the context state after coding two CTUs in the CTU row above the next CTU row. In other words, before starting to code the current wavefront, the video coder (or, more specifically, the thread of that video coder) may assume that the current wavefront is not the first CTU row in the picture and execute coding of at least two blocks in the wavefront above the current wavefront. Then the video coder may initialize the CABAC context for the current wavefront after coding at least two blocks in the wavefront above the current wavefront. In this example, each CTU row of picture 500 is a separate partition and has associated threads (WPP thread 1, WPP thread 2, …) such that multiple CTU rows in the picture are coded in parallel.

[0081]

[0087] Since the current embodiment of the HMVP table stores the motion vectors reconstructed in the past using a global motion vector (MV) buffer, this HMVP table cannot be realized in the WPP-enabled parallel coding method described above with respect to FIG. 5B. In particular, the fact that the global MV buffer is shared by all threads of the video coder's encoding / decoding process prevents the subsequent WPP threads from starting after the first WPP thread (e.g., WPP thread 1). This is because these WPP threads have to wait until the update of the HMVP table is completed from the last CTU (e.g., the rightmost CTU) of the first WPP thread (e.g., the first CTU row).

[0082]

[0088] To overcome such problems, it is proposed to replace the global MV buffer shared by the WPP thread with a buffer dedicated to multi-CTU rows, so that when the WPP of the video coder is effective, each wavefront of the CTU rows has its own buffer to store the HMVP table corresponding to the CTU rows processed by the corresponding WPP thread. It should be noted that each CTU row having its own HMVP table is equivalent to resetting the HMVP table before coding the first CU of the CTU row. Resetting the HMVP table means erasing all motion vectors in the HMVP table obtained from the coding of other CTU rows. In one embodiment, the reset process is to set the size of the available motion vector predictors in the HMVP table to zero. In other embodiments, the reset process may also be to set the reference index of all items in the HMVP table to an invalid value such as -1. By doing so, the MVP candidate list for the current CTU within a specific wavefront is configured according to the HMVP table associated with the WPP thread that processes that specific wavefront, regardless of which of the three modes of AMVP, merge, and skip it is. Except for the delay of the two CTUs described above, there is no interdependence between different wavefronts, and multiple motion vector candidate lists associated with multiple different wavefronts can be constructed in parallel as in the WPP process shown in FIG. 5B. In other words, when starting to process a specific wavefront, the HMVP table is reset to an empty state without being affected by the coding of other CTU wavefronts by other WPP threads. In some cases, the HMVP table can be reset to an empty state before coding each CTU. In this case, the motion vectors in the HMVP table are limited to a specific CTU, and probably the motion vectors in the HMVP table are more likely to be selected as the motion vectors of the current CU within the specific CTU.

[0083]

[0089] FIG. 6 is a flowchart showing an example of processing by a video coder such as a video encoder 20 or a video decoder 30 that implements a technique for constructing a candidate list of motion vector predictors using at least an HMVP table according to some embodiments of the present disclosure. For illustration purposes, the flowchart shows video decoding processing. First, the video decoder 30 obtains an encoded video bitstream including data related to a plurality of encoded pictures (610). As shown in FIGS. 4A and 4C, each picture includes a plurality of coding tree units (CTUs) in a row, and each CTU includes one or more coding units (CUs). The video decoder 30 extracts a plurality of different types of information such as syntax elements and pixel values from the video bitstream and reconstructs the picture on a row-by-row basis.

[0084]

[0090] Before decoding the current CTU row, the video decoder 30 first resets the history-based motion vector predictor (HMVP) table for the current CTU row (620). As described above, resetting the HMVP table ensures that the video decoder 30 can decode a plurality of CTU rows in the current picture in parallel, for example, using multi-threaded processing, and that one thread has its own HMVP table for each CTU row or for each multi-core processor, or that one core has its own HMVP table for each CTU row, or both. In some further embodiments, before decoding the current CTU, the video decoder 30 first resets the history-based motion vector predictor (HMVP) table for the current CTU (620). As described above, resetting the HMVP table ensures that the video decoder 30 can decode a plurality of CTUs in the current picture in parallel, for example, using multi-threaded processing, and that one thread has its own HMVP table for each CTU or for each multi-core processor, or that one core has its own HMVP table for each CTU, or both.

[0085]

[0091] While decoding the current CTU row (630), the video decoder 30 prepares a plurality of motion vector predictors in the HMVP table (630-1). As described above, each motion vector predictor stored in the HMVP table has been used to decode at least other CUs within the current CTU row. In fact, the reason that there are motion vector predictors in the HMVP table is that when the HMVP table is involved in the process of constructing the motion vector candidate list as described above, the motion vector predictor may be reused to predict other CUs within the current CTU row.

[0086]

[0092]

[0093] For the current CU within the current CTU row, the video decoder 30 extracts a prediction mode from the video bitstream (630-3). As described above, the CU may have a plurality of prediction modes including an advanced motion vector prediction (AMVP) mode, a merge mode, a skip mode, an IBC·AMVP mode, and an IBC merge mode. When the video encoder 20 selects an appropriate prediction mode for the CU, the selected prediction mode is signaled in the bitstream. As described above, there are various groups of steps executed in various orders to construct the motion vector candidate list. Here, the video decoder 30 constructs a motion vector candidate list based on at least a part of the plurality of motion vector predictors in the HMVP table according to the prediction mode (630-5). The motion vector candidate list from other information sources includes motion vector predictors (when the prediction mode is one of the AMVP mode, the IBC·AMVP mode, and the IBC merge mode) derived from CUs that are spatially adjacent and / or temporally connected to the current CU, and optionally includes pairwise motion vector predictors (when the prediction mode is one of the merge mode and the skip mode). Optionally, when the motion vector candidate list does not reach a predetermined length, one or more zero-value motion vector predictors may be added to the motion vector candidate list.

[0087]

[0094] Next, the video decoder 30 selects a motion vector predictor for the current CU from the motion vector candidate list (630-7), and determines a motion vector based on at least a part of the selected motion vector predictor and the prediction mode (630-9). As described above, depending on whether the prediction mode is the AMVP mode or not, the selected motion vector predictor may or may not be the estimated motion vector for the current CU. For example, if the prediction mode is the AMVP mode, the estimated motion vector is determined by adding the motion vector difference reproduced from the bitstream to the selected motion vector predictor, and then the current CU is decoded using at least partially the estimated motion vector and the corresponding CU in the reference picture. However, if the prediction mode is the merge mode or the skip mode, the selected motion vector predictor is already the estimated motion vector and can be used when decoding the current CU together with the corresponding CU in the reference picture. Finally, the video decoder 30 updates the HMVP table based on the determined motion vector (630-11). As described above, all elements in the HMVP table have been previously used at least for decoding other CUs, and are held in the HMVP table to form the motion vector candidate list until they are deleted from the HMVP table by insertion of the motion vector used for decoding the next other CU in the current CTU row or by table reset.

[0088]

[0095] In some embodiments, there are two scenarios for inserting a motion vector into the HMVP table, based on the comparison result between the motion vector determined for the current CU and multiple motion vector predictors in the HMVP table. When none of the multiple motion vector predictors in the HMVP table is the same as the determined motion vector, if the HMVP table is full, the earliest or oldest motion vector predictor is deleted from the HMVP table, and the motion vector is added to the table as the latest one. When one of the multiple motion vector predictors in the HMVP table is the same as the motion vector, the same one motion vector predictor is deleted from the HMVP table, and all the other motion vector predictors after the deleted motion vector predictor are moved forward in the HMVP table, so that the motion vector is added to the rear of the HMVP table as the latest one.

[0089]

[0096] As described above, two or more of the multiple CTU rows may be encoded / decoded in parallel, for example, using WPP, and each CTU row has an associated HMVP table for storing a plurality of history-based motion vector predictors used to encode / decrypt the corresponding CTU row. For example, a thread is assigned to decrypt a specific CTU row in the currently decoded picture, so that multiple different CTU rows can be decoded with multiple different associated threads as described above with respect to FIG. 5B. In some examples, the video decoder 30 identifies one or more motion vector predictors in the motion vector candidate list as redundant and deletes it from the motion vector candidate list to further improve the encoding efficiency.

[0090]

[0097] In one or more examples, the above functions may be implemented in hardware, software, or a combination thereof. When the above functions are implemented in software, they may be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable recording medium corresponding to a tangible medium such as a recording medium, or a communication medium having any medium that facilitates the transfer of a computer program from one location to another, for example, according to a communication protocol. Thus, the computer-readable medium may generally correspond to either (1) a tangible non-transitory computer-readable recording medium or (2) a communication medium such as a signal or a carrier wave. The data recording medium may be any available medium accessible by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0091]

[0098] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. The singular forms "a", "an", and "the" used in the description of the embodiments and the appended claims are intended to include the plural forms as well, unless the context clearly dictates otherwise. It will also be understood that the expression "and / or" used herein means and encompasses any and all possible combinations of one or more of the listed related items. Furthermore, the expressions "comprises" and / or "comprising" as used herein, when used in this specification, specify the presence of the defined features, components, and / or elements, but do not preclude the presence or addition of one or more other features, elements, parts, and / or combinations thereof.

[0092]

[0099] The terms such as "first (beginning)" and "second" are used herein to describe a variety of elements, but it will also be understood that these elements should not be limited by these terms. These terms are only used to distinguish between elements. For example, without departing from the scope of the embodiments, the first electrode may sometimes be referred to as the second electrode, and similarly, the second electrode may sometimes be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but they are not the same electrode.

[0093]

[0100] The description of the present application is for illustration and explanation purposes and is not intended to be comprehensive or to limit the disclosure to the disclosed embodiments. Many improvements, modifications, and alternative embodiments will be apparent to those skilled in the art who benefit from the teachings shown in the above description and its associated drawings. The above embodiments have been selected and described in order to best explain the principles of the present invention and its practical applications, and to enable those skilled in the art to understand the present invention with respect to various embodiments, and to make the best use of the various embodiments together with the underlying principles and various modifications suitable for the intended specific applications. Therefore, it should be understood that the claims should not be limited to the specific examples of the disclosed embodiments, and that variations and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. A decoding method, comprising: before decoding the first CU in the current CTU row in the current picture to be decoded, resetting a history-based motion vector predictor (HMV P) table; decoding the current CTU row; The step of decoding the current CTU row includes: refreshing a plurality of motion vector predictors in the HMV P table, each motion vector predictor in the HMV P table being used to decode at least one CU in the current CTU row; for the current CU in the current CTU row to be decoded, constructing a motion vector candidate list based on at least a part of the plurality of motion vector predictors in the HMV P table according to a prediction mode; selecting a motion vector predictor from the motion vector candidate list; determining a motion vector based on at least a part of the prediction mode and the selected motion vector predictor to decode the current CU; updating the HMV P table based on the determined motion vector, The step of updating the HMV P table based on the determined motion vector includes: comparing the plurality of motion vector predictors in the HMV P table with the determined motion vector; according to a comparison result that the plurality of motion vector predictors in the HMV P table are the same as the determined motion vector, deleting the one motion vector predictor that is the same from the HMV P table; moving each of the motion vector predictors after the deleted motion vector predictor forward in the HMV P table; adding the determined motion vector to the HMV P table as the latest motion vector; The prediction mode is an inter mode; The motion vector candidate list has a fixed length of 2; The step of constructing the motion vector candidate list includes: when a history-based motion vector predictor is selected from the HMV P table to construct the motion vector candidate list, adding two history-based motion vector predictors from the HMV P table to the motion vector candidate list.

2. The method according to claim 1, wherein the step of updating the HMV P table based on the determined motion vector is ​ In response to a determination that none of the plurality of motion vector predictors in the HMVP table is the same as the determined motion vector, when the HMVP table is full, delete the earliest motion vector predictor from the HMVP table, and add the determined motion vector as the latest motion vector to the HMVP table The method further includes.

3. The method according to claim 1, The step of constructing the motion vector candidate list includes: adding zero or more motion vector predictors derived from a temporally adjacent CU and / or a temporally connected CU to the current CU to the motion vector candidate list; when the current length of the motion vector candidate list is shorter than a first predetermined threshold, adding zero or more history-based motion vector predictors derived from the HMVP table to the motion vector candidate list until the current length of the motion vector candidate list equals the first predetermined threshold The method further includes.

4. The method according to claim 3, The step of constructing the motion vector candidate list includes: when the current length of the motion vector candidate list is shorter than a first predetermined threshold, adding zero or more history-based motion vector predictors with zero values to the motion vector candidate list until the current length of the motion vector candidate list equals the first predetermined threshold The method further includes.

5. The method according to claim 3, The step of adding zero or more motion vector predictors from a spatially adjacent CU and / or a temporally connected CU to the current CU to the motion vector candidate list includes: adding zero or more motion vector predictors from a spatially adjacent CU to the current CU to the motion vector candidate list; when the current length of the motion vector candidate list is shorter than the first predetermined threshold, adding zero or more motion vector predictors from a temporally connected CU to the current CU to the motion vector candidate list The method includes.

6. The method according to claim 5, wherein the step of adding zero or more motion vector predictors from the CU that is temporally concatenated to the current CU to the motion vector candidate list includes the step of extracting a candidate for one motion vector predictor from the CU that is temporally concatenated.

7. The method according to claim 1, wherein the step of resetting the HMVP table includes the step of setting the size of the available motion vector predictors in the HMVP table to zero.

8. The method according to claim 1, wherein the current picture to be decoded is obtained from a video bitstream, and the step of decoding the current CTU row includes the step of extracting the prediction mode from the video bitstream.

9. An arithmetic device, one or more processors, a memory connected to the one or more processors, and a plurality of programs stored in the memory and configured to, when the plurality of programs are executed by the one or more processors, cause the arithmetic device to execute the method according to any one of claims 1 to 8.

10. A non-transitory computer-readable recording medium storing a plurality of programs and a bitstream for execution by an arithmetic device having one or more processors, wherein when the plurality of programs are executed by the one or more processors, the arithmetic device is caused to execute the method according to any one of claims 1 to 8 to decode the bitstream.

11. A computer program for execution by an arithmetic device having one or more processors, wherein when the computer program is executed by the one or more processors, the arithmetic device is caused to execute the method according to any one of claims 1 to 8.

12. A method of storing a bitstream, including the step of performing an encoding method to generate a bitstream and the step of storing the bitstream, wherein the encoding method includes the step of dividing a current picture into a plurality of encoded tree units (CTU) rows, each CTU in the plurality of CTU rows including one or more encoded units (CU). ​ Before processing the first CU in the current CTU row within the current picture, a step of resetting a history-based motion vector predictor (HMVP) table; A step of processing the current CTU row; Comprising; The step of processing the current CTU row is; A step of arranging a plurality of motion vector predictors in the HMVP table, wherein each motion vector predictor in the HMVP table is used to process at least one CU in the current CTU row; For the current CU of the current CTU row to be processed, According to the prediction mode, a motion vector candidate list is configured based on at least a part of the plurality of motion vector predictors in the HMVP table; Select a motion vector predictor from the motion vector candidate list; To process the current CU, a motion vector is determined based on at least a part of the prediction mode and the selected motion vector predictor; Including a step of updating the HMVP table based on the determined motion vector; The step of updating the HMVP table based on the determined motion vector is; A step of comparing the plurality of motion vector predictors in the HMVP table with the determined motion vector; According to the comparison result that the plurality of motion vector predictors in the HMVP table are the same as the determined motion vector, Deleting the one motion vector predictor that is the same from the HMVP table; Moving each of the motion vector predictors after the deleted motion vector predictor forward in the HMVP table; Including a step of adding the determined motion vector as the latest motion vector to the HMVP table; The prediction mode is the inter mode; The motion vector candidate list has a fixed length of 2; The step of configuring the motion vector candidate list is; When a history-based motion vector predictor is selected from the HMVP table to configure the motion vector candidate list, including adding two history-based motion vector predictors from the HMVP table to the motion vector candidate list. Method.

Citation Information

Patent Citations

  • Multiple History-Based Non-Adjacent MVP for Wavefront Processing in Video Coding

    JP2021530904A

  • Image encoding method, image decoding method, image encoding device, image decoding device and program

    WO2018123317A1