Method and apparatus for video coding
The method addresses the challenges of video coding by using history-based motion vector prediction candidates with weight parameters to enhance encoding and decoding efficiency and compression ratios.
Patent Information
- Application Number
- JP2024120779
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-14
- Filing Date
- 2024-07-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-12-04
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding video data, particularly in reducing redundancy and improving compression ratios while maintaining acceptable video quality.
The proposed method and apparatus for video encoding and decoding involve obtaining prediction information from an encoded video bitstream, generating reconstructed samples, and storing motion information candidates as history-based motion vector prediction (HMVP) candidates. These candidates include motion information and weight parameters for bidirectional or unidirectional prediction, allowing for efficient encoding and decoding processes.
This approach enhances video encoding and decoding efficiency by reducing data redundancy and improving compression ratios, while maintaining acceptable video quality through effective use of motion information and weight parameters.
Smart Images

Figure 0007690659000001 
Figure 0007690659000002 
Figure 0007690659000003
Abstract
Description
Technical Field
[0001] [Related Applications] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 777,593, "Methods of GBi Index Inheritance and Constraints", filed on December 10, 2018, and U.S. Patent Application No. 16 / 441,879, "Method and Apparatus for Video Coding", filed on June 14, 2019. The entire disclosure of the foregoing applications is hereby incorporated by reference in its entirety.
[0002] [Technical Field] This disclosure generally describes embodiments related to video coding.
Background Art
[0003] The background description provided herein is for the purpose of presenting an overview of the context of the present disclosure. The research of the presently named inventors, to the extent it is within the scope of the research described in this background chapter, is not to be regarded as prior art, whether explicitly or implicitly, to the present disclosure, any more than is the case for aspects of the description that may not be regarded as prior art at the time of filing.
[0004] Video encoding and decoding can be performed using inter-picture prediction with motion compensation. Uncompressed digital video can include a series of pictures, each picture having a spatial dimension of, for example, 1920×1080 luminance samples and associated chrominance samples. A series of pictures can have a fixed or variable picture rate, for example, 60 pictures per second or 60 Hz (also known as the frame rate in short form). Uncompressed video has significant bitrate requirements. For example, 8-bit / sample 1080p60 4:2:0 video (1920×1080 luminance sample resolution at 60 Hz frame rate) requires a bandwidth of 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.
[0005] One purpose of video encoding and decoding can be the reduction of redundancy in an input video signal through compression. Compression can, in some cases, help reduce the aforementioned bandwidth or storage space requirements by more than an order of magnitude. Both lossy and lossless compression, and combinations thereof, are available. Lossless compression represents techniques where an exact copy of the original signal can be reconstructed from the compressed original signal. When using lossy compression, the reconstructed signal is not identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to produce a useful reconstructed signal for the intended application. In the case of video, lossy compression is widely used. The amount of tolerable distortion depends on the application, and users of certain consumer streaming applications can tolerate more distortion than users of television distribution applications. The achievable compression ratio can reflect that the higher the acceptable / tolerable distortion, the higher the compression ratio that can be achieved.
[0006] Motion compensation is a lossy compression technique and can be related to techniques where a block of sample data from a previously reconstructed picture or a portion thereof (reference picture) is spatially shifted in a direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or picture portion. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions X and Y, or three dimensions where the third dimension is an indication of the reference picture in use (the latter can be indirectly the temporal dimension).
[0007] In some video compression techniques, the motion vectors (MVs) applicable to a particular region of sample data can be predicted from other MVs, for example, from MVs related to another region of sample data that is spatially adjacent to the region being reconstructed and that precedes the said MV in the decoding order. Doing so can, as a result, reduce the amount of data required to encode the MVs, thereby removing redundancy and improving compression. MV prediction can, for example, be statistically possible when encoding an input video signal obtained from a camera (known as natural video), as regions larger than the region to which a single MV is applicable move in a similar direction, and thus, in some cases, can be predicted using similar motion vectors derived from the MVs of neighboring regions. This results in an MV found for a given region that is similar to or the same as the MV predicted from the surrounding MVs. Also, this can be presented in fewer bits than would be used if the MVs were directly encoded after entropy coding. In some cases, MV prediction can be an example of lossless compression of the signals (i.e., MVs) obtained from the original signal (i.e., the sample stream). In other cases, MV prediction itself can result in loss, for example, when rounding errors when calculating predictors from some of the surrounding MVs.
[0008] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). One of the many MV prediction mechanisms provided by H.265 described herein is a technique hereinafter referred to as "spatial merge".
[0009] Referring to FIG. 1, the current block (101) includes samples found by the encoder as being predictable from a previous block of the same size that has been spatially shifted during motion search processing. Instead of directly encoding the MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the nearest reference picture (in decoding order), using an MV associated with any one of five surrounding samples A0, A1, and B0, B1, B2 (102-106 respectively). In H.265, MV prediction can use predictors from the same reference picture used by neighboring blocks.
SUMMARY OF THE INVENTION
[0010] Disclosed aspects provide a method and apparatus for video encoding and decoding. In some examples, the apparatus includes a processing circuit that obtains prediction information for a first block in a picture from an encoded video bitstream and generates reconstructed samples for the first block for output according to one of bidirectional prediction and unidirectional prediction and the prediction information. When it is determined that motion information candidates should be stored according to the prediction information of the first block and should be stored as history-based motion vector prediction (HMVP) candidates, the processing circuit stores the motion information candidates, and the motion information candidates include at least: when the first block is encoded according to bidirectional prediction, a first motion information, and a first weight parameter indicating a first weight for performing bidirectional prediction for the first block; and when the first block is encoded according to unidirectional prediction, the first motion information and a prescribed weight parameter indicating a prescribed weight are stored. When it is determined that a second block in the picture should be decoded based on the motion information candidates, the processing circuit generates reconstructed samples for the second block for output according to the motion information candidates.
[0011] In some embodiments, when motion information candidates are stored as normal spatial merge candidates and the second block is encoded according to bidirectional prediction, the processing circuit, when the first block is spatially adjacent to the second block, sets a second weight for performing bidirectional prediction for the second block according to a first weight parameter stored in the motion information candidates, and when the first block is not spatially adjacent to the second block, sets the second weight for performing bidirectional prediction for the second block to a specified weight.
[0012] In some embodiments, when motion information candidates are stored as candidates that are neither normal spatial merge candidates nor HMVP candidates and the second block is encoded according to bidirectional prediction, the processing circuit is further configured to set a second weight for performing bidirectional prediction for the second block to a specified weight.
[0013] In some embodiments, when the first block is in a CTU row different from the current coding tree unit (CTU) row in which the second block is included, the motion information candidates are stored as normal merge candidates or affine merge candidates, and the second block is encoded according to bidirectional prediction, the processing circuit sets a second weight for performing bidirectional prediction for the second block to a specified weight.
[0014] In some embodiments, when the first block is outside the current CTU in which the second block is included, the motion information candidates are stored as translational merge candidates or inherited affine merge candidates, and the second block is encoded according to bidirectional prediction, the processing circuit sets a second weight for performing bidirectional prediction for the second block to a specified weight.
[0015] In some embodiments, the first block is encoded according to one of bidirectional prediction and unidirectional prediction with the picture as a reference picture.
[0016] In some embodiments, the first block is encoded according to bidirectional prediction, and both the first weight corresponding to the first reference picture in the first list and the second weight corresponding to the second reference picture in the second list derived from the first weight are positive when the first and second reference pictures are the same reference picture. In some embodiments, the first block is encoded according to bidirectional prediction, and one of the first weight corresponding to the first reference picture in the first list and the second weight corresponding to the second reference picture in the second list derived from the first weight is negative when the first and second reference pictures are different reference pictures.
[0017] In some embodiments, the defined weight is 1 / 2.
[0018] In some embodiments, the first block is encoded according to bidirectional prediction, and the first weight w 1 is determined according to w 1 = w / F, and another weight w 0 corresponding to the second reference picture in the second list is determined according to w 0 = 1 - w 1 where w and F are integers, w represents the first weight parameter, and F represents the accuracy coefficient. In some embodiments, F is 8.
[0019] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to execute a method for video decoding.
Brief Description of the Drawings
[0020] Further features, characteristics, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.
[0021]
Figure 1
[0022]
Figure 2
[0023]
Figure 3
[0024]
Figure 4
[0025]
Figure 5
[0026]
Figure 6
[0027]
Figure 7
[0028]
Figure 8
[0029]
Figure 9
[0030]
Figure 10A
[0031]
Figure 10B
[0032]
Figure 11A
[0033]
Figure 11B
[0034]
Figure 12
[0035]
Figure 13
[0036]
Figure 14
[0037]
Figure 15
[0038]
Figure 16
MODE FOR CARRYING OUT THE INVENTION
[0039] Figure 2 shows a simplified block diagram of a communication system (200) according to an embodiment of the present invention. The communication system (200) includes a plurality of terminal devices that can communicate with each other via, for example, a network (250). For example, the communication system (200) includes a first pair of terminal devices (210) and (220) interconnected via a network (250). In the example of Figure 2, the first pair of terminal devices (210) and (220) perform unidirectional data transmission. For example, the terminal device (210) encodes video data (a stream of video pictures captured by the terminal device (210)) for transmission to another terminal device (220) via the network (250). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (220) receives the encoded video data from the network (250), decodes the encoded video data to restore the video pictures, and may display the video pictures according to the restored video data. Unidirectional data transmission may be common in media serving applications and the like.
[0040] In another example, the communication system (200) includes a second pair of terminal devices (230) and (240) that perform bidirectional transmission of encoded video data that may occur, for example, during a video conference. In bidirectional data transmission, the terminal devices (230) and (240) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the terminal devices (230) and (240) via the network (250). Each of the terminal devices (230) and (240) may receive the encoded video data transmitted by the other of the terminal devices (230) and (240), may decode the encoded video data to restore the video pictures, and may display the video pictures on an accessible display device according to the restored video data.
[0041] In the example of FIG. 2, the terminal devices (210), (220), (230), and (240) may be shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure have applications in laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (250) represents any number of networks that carry encoded video data among the terminal devices (210), (220), (230), and (240), including, for example, wired (wired) and / or wireless communication networks. The communication network 250 may exchange data over a circuit-switched and / or packet-switched channel. Representative networks include electronic communication networks, local area networks, wide area networks, and / or the Internet. For the purposes of the discussion of the present invention, the architecture and topology of the network (250) may not be important for the operation of the present disclosure, unless otherwise specifically noted below.
[0042] FIG. 3 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter is equally applicable to, for example, video conferencing, digital TV, CD, DVD, memory stick, and other video-enabled applications, such as storage of compressed video on digital media.
[0043] A streaming system may include a capture subsystem (313) that can include, for example, a video source (301) that generates an uncompressed video picture stream (302). In one example, the video picture stream (302) includes samples captured by a digital camera. The video picture stream (302) is shown in bold lines to emphasize its high data capacity when compared to the encoded video data (304) (or encoded video bitstream), and can be processed by an electronic device (320) that includes a video encoder (303) coupled to the video source (301). The video encoder (303) includes hardware, software, or a combination thereof and can enable or implement aspects of the disclosed subject matter as detailed below. The encoded video data (304) (or video bitstream (304)) is shown in thin lines to emphasize its low data capacity when compared to the video picture stream (302), and can be stored in a streaming server for future use. One or more streaming client subsystems, such as the client subsystems (306) and (308) of FIG. 3, can access the streaming server (305) to read copies (307) and (309) of the encoded video data (304). The client subsystem (306) can include, for example, a video decoder (310) within an electronic device (330). The video decoder (310) decodes an input copy (307) of the encoded video data and generates an output video picture stream (311) that can be rendered on a display (312) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (304), (307), and (309) (e.g., the video bitstream) can be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video encoding standard under development is known informally as VVC (Versatile Video Coding). The disclosed subject matter may be used in the context of VVC.
[0044] Note that the electronic devices (320) and (330) may include other components (not shown). For example, the electronic device (320) can include a video decoder (not shown), and the electronic device (330) can also include a video encoder (not shown).
[0045] FIG. 4 shows a block diagram of a video decoder (410) according to an embodiment of the present disclosure. The video decoder (410) may be included in an electronic device (430). The electronic device (430) may include a receiver (431) (e.g., a receiving circuit). The video decoder (410) can be used instead of the video decoder (310) in the example of FIG. 3.
[0046] The receiver (431) may receive one or more encoded video sequences to be encoded by the video decoder (410), or in the same or another embodiment, one encoded video sequence at a time. Here, the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel 401 which may be a hardware / software link to a storage device storing the encoded video data. Receiver 431 may receive the encoded video data along with other data, such as encoded audio data and / or an accompanying data stream that may be transferred to respective using entities (not shown). Receiver 431 may separate the encoded video sequence from the other data. To remove network jitter, a buffer memory (415) may be coupled between the receiver (431) and the entropy decoder / parser (420) (hereinafter, “parser (420)”). In certain applications, buffer memory (415) is part of video decoder (410). Alternatively, it may be external to video decoder (410) (not shown). Still alternatively, for example, to remove network jitter, in addition to another buffer memory (415) that may be external to video decoder (410), for example, to handle playout timing, there may be a buffer memory (not shown) internal to video decoder (410). When receiver (431) is receiving data controllably from a sufficient bandwidth storage / transfer device or from an isosynchronous network, buffer memory (415) may not be necessary or may be made small. For use in a best effort packet network such as the Internet, buffer memory (415) may be required, may be relatively large, advantageously be of an appropriate size, and may be implemented at least partially external to the operating system or similar elements (not shown) external to video decoder (410).
[0047] The video decoder (410) may include a parser (420) to reconstruct symbols (421) from an encoded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (410) and, in some cases, information for controlling a rendering device (412) (e.g., a display screen) that is not an integral part of the electronic device (430) but can be coupled to the electronic device (430) as shown in FIG. 4. The control information for the rendering device may be in the form of an SEI (Supplemental Enhancement Information) message or a VUI (Video Usability Information) parameter set fragment (not shown). The parser (420) may parse / entropy decode the received coded video sequence. The encoding of the encoded video sequence can follow various principles, including arithmetic coding, etc., which may or may not be dependent on following a video encoding technology or standard. The parser (420) may extract a set of subgroup parameters from the encoded video sequence based on at least one parameter corresponding to at least one subgroup of pixels in the video decoder, where the subgroup may include GOP (Groups of Picture), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (420) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence.
[0048] The parser (420) may perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (415) to generate symbols (421).
[0049] The reconstruction of symbol 421 may include multiple different units depending on the type of the encoded video picture or a portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. How the units are included can be controlled by subgroup control information parsed from the encoded video sequence by parser 420. Such a flow of subgroup control information between parser 420 and the following multiple units is not shown for clarity.
[0050] Beyond the function blocks already mentioned, video decoder (410) can be conceptually subdivided into a number of functional units, as will be described later. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0051] The first unit is the scaler / inverse transform unit 451. The scaler / inverse transform unit (451) receives, as symbols (421) from parser (420), quantized transform coefficients and control information including which transform may be used, block size, quantization coefficients, quantization scaling matrix, etc. The scaler / inverse transform unit (451) can output a block including sample values that can be input to aggregator (455).
[0052] In some examples, the output samples of the scaler / inverse transform unit (451) can belong to an intra-coded block, i.e., a block that does not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit 452. In some cases, the intra-picture prediction unit (452) generates a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current picture buffer (458). The current picture buffer (458) buffers, for example, the reconstructed current picture partially and / or the reconstructed current picture completely. The aggregator (455) adds, in some cases, sample by sample, the prediction information generated by the intra prediction unit (452) to the output sample information provided by the scaler / inverse transform unit (451).
[0053] In other cases, the output samples of the scaler / inverse transform unit (451) can be related to inter-coded, and possibly motion-compensated, blocks. In such cases, the motion compensation prediction unit (453) can access the reference picture memory (457) to fetch the samples used for prediction. After motion compensating the samples fetched according to the symbols (421) associated with the block, these samples can be added by the aggregator (455) to the output of the scaler / inverse transform unit (451) to generate the output sample information (in this case, called residual samples or a residual signal). The address in the reference picture memory (457) from which the motion compensation prediction unit (453) fetches the prediction samples can be controlled by the available motion vectors of the motion compensation prediction unit (453) in a format of, for example, symbols (421) having X, Y, and reference picture components. Motion compensation can include interpolation of the sample values fetched from the reference picture memory (457) when an exact sub-sample motion vector is in use, a motion vector prediction mechanism, etc.
[0054] The output samples of the aggregator (455) may undergo various loop filtering techniques in the loop filter unit (456). The video compression technology is controlled by parameters included in the encoded video sequence (also called the encoded video bitstream) and made available to the loop filter unit (456) as symbols (421) from the parser (420), but also responds to meta information obtained during the composition of the previous part of the encoded picture or encoded video sequence (in composite order), and may include in-loop filtering techniques that can also respond to previously reconstructed and loop-filtered sample values.
[0055] The output of the loop filter unit (456) can be an output to the renderer device (412) and a sample stream that can be stored in the reference picture memory (457) for use in future inter-picture prediction.
[0056] Once a particular encoded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (420)), the current picture buffer (458) can become part of the reference picture memory (457), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent encoded pictures.
[0057] The video decoder (410) may perform a decoding operation according to a predetermined video compression technique of a standard such as ITU-T Rec. H.265. In the sense that the encoded video sequence conforms to both the video compression technique or standard and the profile documented in the video compression technique or standard, the encoded video sequence may conform to the syntax specified by the video compression technique or standard in use. Specifically, the profile can select specific tools from all the tools available in the video compression technique or standard as tools that can be used only under the profile. Also, what may be required for compliance is that the complexity of the encoded video sequence is within the limits determined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured, for example, in megasamples per second), the maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted through the HRD (Hypothetical Reference Decoder) specification and the metadata for HDR buffer management signaled in the encoded video sequence.
[0058] In one embodiment, the receiver 431 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder 410 to correctly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0059] FIG. 5 shows a block diagram of a video encoder (503) according to an embodiment of the present disclosure. The video encoder (503) is included in an electronic device (520). The electronic device (520) includes a transmitter (540) (e.g., a transmission circuit). The video encoder (503) can be used instead of the video encoder (303) in the example of FIG. 3.
[0060] The video encoder (503) may receive video samples from a video source (501) (not part of the electronic device (520) in the example of FIG. 5) that can capture a video image to be encoded by the video encoder (503). In another example, the video source (501) is part of the electronic device (520).
[0061] The video source (501) may provide a source video sequence to be encoded by the video encoder (503) in the form of a digital video sample stream of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit,...), any color space (e.g., BT.601 YCrCb, RGB,...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media providing system, the video source 501 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 501 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that give movement when viewed subsequently. The pictures themselves may be organized as a spatial array of pixels. Each pixel may contain one or more samples depending on the sampling structure, color space, etc. in use. One of ordinary skill in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.
[0062] According to one embodiment, the video encoder (503) may encode and compress pictures of a source video sequence into an encoded video sequence (543) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is one function of the control unit (550). In some embodiments, the control unit (550) controls other functional units described below and is functionally coupled to the other functional units. The coupling is not shown for clarity. The parameters set by the control unit (550) may include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique,...), picture size, GOP (group of pictures) layout, maximum motion vector search range, etc. The control unit (550) may be configured to have other appropriate functions related to the video encoder (503) optimized for a particular system design.
[0063] In some embodiments, the video encoder (503) is configured to operate within an encoding loop. As a very simplified explanation, in one example, the encoding loop may include a source coder (530) (which is responsible for generating symbols such as a symbol stream based on the input picture and reference pictures to be encoded), and a (local) decoder (533) built into the video encoder (503). The decoder (533) reconstructs the symbols to generate sample data in the same way as a (remote) decoder would when any compression between the symbols and the encoded bitstream is lossless in the video compression techniques contemplated in the subject disclosure. The reconstructed sample stream (sample data) is input into the reference picture memory (534). When the decoding of the symbol stream yields bit-exact results independent of the decoder location (local or remote), the contents of the reference picture memory (534) are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the decoder would "see" as reference picture samples when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used similarly in some related arts.
[0064] The operation of the "local" decoder (533) can be the same as that of a "remote" decoder such as the video decoder (410) detailed above in relation to FIG. 4. Briefly referring to FIG. 4 for a moment, however, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy coder (545) and the parser (420) can be lossless, the entropy decoding part of the video decoder (410) including the buffer memory (415), and the parser (420) need not be fully implemented in the local decoder (533).
[0065] The consideration made in this regard is that any decoder technology other than the parse / entropy decoding existing in the decoder needs to exist in substantially the same functional form as that in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted because they are the reverse of the decoder technologies that are comprehensively described. More detailed descriptions are necessary only in specific areas and are provided below.
[0066] During operation, in some examples, the source coder (530) may perform motion-compensated predictive coding. This predictively encodes the input picture by referring to one or more previously encoded pictures from the video sequence designated as the "reference picture". In this method, the encoding engine (532) encodes the difference between the pixel block of the input picture and the pixel block of the reference picture that may be selected as the prediction criterion for the input picture.
[0067] The local video decoder (533) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbols generated by the source coder (530). The operation of the encoding engine 532 may advantageously be a lossy process. When the encoded video data can be decoded in a video decoder (not shown in FIG. 5), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (533) replicates the decoding process that may be performed by the video decoder for the reference picture and may generate a reconstructed reference picture to be stored in the reference picture cache (534). In this way, the video encoder (503) may store a copy of the reconstructed reference picture having the same content as the reconstructed reference picture obtained by the remote video decoder (if there are no transmission errors).
[0068] The predictor (535) may perform a prediction search for the encoding engine (532). That is, for a new picture to be encoded, the predictor (535) may search the reference picture memory (534) for sample data (such as a candidate reference pixel block) or specific metadata such as reference picture motion vectors, block shapes, etc. that can function as an appropriate prediction criterion for the new picture. The predictor (535) may operate on a sample block - pixel block basis to find an appropriate prediction criterion. In some examples, the input picture may have prediction criteria drawn from a plurality of reference pictures stored in the reference picture memory 534, as determined by the search results obtained by the predictor 535.
[0069] The control unit (550) may manage the encoding operation of the source coder (530), including, for example, setting parameters and subgroup parameters used for the encoding of video data.
[0070] The outputs of all the aforementioned functional units may undergo entropy encoding in the entropy coder (545). The entropy coder (545) converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc.
[0071] The transmitter (540) may buffer the encoded video sequence generated by the entropy coder (545) for transmission via a communication channel (560) that may be a hardware / software link to a storage device capable of storing the encoded video data. The transmitter 540 may merge the encoded video data from the video coder 503 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (not shown source).
[0072] The control unit (550) may manage the operation of the video encoder (503). During encoding, the control unit 550 may assign to each encoded picture a specific encoded picture type that can affect the encoding technique applicable to that picture. For example, a picture may often be assigned as one of the following picture types.
[0073] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, for example, IDR (Independent Decoder Refresh) pictures. Those skilled in the art recognize the variations of I pictures and their individual applications and characteristics.
[0074] A predicted picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, typically using one motion vector and a reference index to predict the sample values of each block.
[0075] A bi - directionally predicted picture (B picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, typically using up to two motion vectors and reference indices to predict the sample values of each block. Similarly, a multi - predicted picture can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0076] The source picture may be spatially subdivided, in common, into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. The blocks may be encoded predictively by reference to other (already encoded) blocks determined by an encoding assignment applied to each picture of the block. For example, blocks of an I picture may be encoded non-predictively, or they may be encoded predictively by reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture may be encoded predictively via spatial prediction or via temporal prediction by reference to one previously encoded reference picture. Blocks of a B picture may be encoded predictively via spatial prediction or via temporal prediction by reference to one or two previously encoded reference pictures.
[0077] The video decoder (503) may perform an encoding operation according to a predetermined video encoding technique or standard such as ITU-T Rec. H.265. In that operation, the video encoder (503) may perform various compression operations including a predictive encoding operation that utilizes temporal and spatial redundancies in the input video sequence. The encoded video data may thus conform to a syntax specified by the video encoding technique or standard being used.
[0078] In one embodiment, the transmitter 540 may transmit additional data together with the encoded video. The source coder (530) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, and the like.
[0079] Video may be captured as a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (which may be abbreviated as intra-prediction) utilizes the spatial correlation within a given picture, and inter-picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture during encoding / decoding is called the current picture and is partitioned into blocks. When a block in the current picture is similar to a reference block in a reference picture that has been encoded previously in the video and is still buffered, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture when multiple reference pictures are in use.
[0080] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to the bi-prediction technique, two reference pictures such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order respectively) are used. A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by the combination of the first reference block and the second reference block.
[0081] Furthermore, in order to improve the encoding efficiency, merge mode techniques can be used in inter-picture prediction.
[0082] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed within a unit of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are partitioned into coding tree units (CTUs) for compression. CTUs within a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Usually, a CTU includes three coding tree blocks (CTBs), that is, one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Usually, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in encoding (encoding / decoding) are performed within a unit of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0083] FIG. 6 shows a diagram of a video encoder (603) according to another embodiment of the present disclosure. The video encoder (603) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a video picture sequence and encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (603) is used in place of the video encoder (303) in the example of FIG. 3.
[0084] In an example of HEVC, a video encoder (603) receives a matrix of sample values of a processing block, such as a prediction block of 8×8 samples. The video encoder (603) determines, for example using rate distortion optimization, whether the processing block is optimally encoded using an intra mode, an inter mode, or a bi-prediction mode. When the processing block is encoded in the intra mode, the video encoder (603) may use intra prediction techniques to encode the processing block into an encoded picture. When the processing block is encoded in the inter mode or the bi-prediction mode, the video encoder (603) may use inter prediction or bi-prediction techniques, respectively, to encode the processing block into an encoded picture. In certain video encoding techniques, a merge mode may be an inter-picture prediction sub-mode in which a motion vector is obtained from one or more motion vector predictors without encoded motion vector components of a gear part of the predictor. In certain other video encoding techniques, there may be motion vector components applicable to a target block. In one example, the video encoder (603) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.
[0085] In the example of FIG. 6, the video encoder (603) includes an inter encoder (630), an intra encoder (622), a residual calculator (623), a switch (626), a residual encoder (624), a general-purpose control unit (621), and an entropy encoder (625) together as shown in FIG. 6.
[0086] The inter-encoder (630) is configured to receive samples of a current block (e.g., a block being processed), compare the block with one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generate inter-prediction information (e.g., an explanation of redundant information by an inter-encoding technique, a motion vector, a merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.
[0087] The intra-encoder (622) is configured to receive samples of a current block (e.g., a block being processed), and in some cases, compare the block with already encoded blocks in a sample picture, and generate quantized coefficients after transformation, and in some cases, also generate intra-prediction information (e.g., intra-prediction direction information by one or more intra-encoding techniques). In one example, the intra-encoder (622) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks in the same picture.
[0088] The general control unit (621) is configured to determine general control data and control other components of the video encoder (603) based on the general control data. In one example, the general control unit (621) determines a mode of a block and provides a control signal to a switch (626) based on the mode. For example, when the mode is an intra mode, the general control unit (621) controls the switch (626) to select an intra-mode result for use by the residual calculator (623), and controls the entropy encoder (625) to select the intra-prediction information and include the intra-prediction information in the bitstream. When the mode is an inter mode, the general control unit (621) controls the switch (626) to select an inter-prediction result for use by the residual calculator (623), and controls the entropy encoder (625) to select the inter-prediction information and include the inter-prediction information in the bitstream.
[0089] The residual calculator (623) is configured to calculate the difference (residual data) between the received block and the selected prediction result from the intra encoder (622) or the inter encoder (630). The residual encoder (624) is configured to operate based on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (624) is configured to convert the residual data from the spatial domain to the frequency domain and generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (603) also includes a residual decoder (628). The residual decoder (628) is configured to perform inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra encoder (622) and the inter encoder (630). For example, the inter encoder (630) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (622) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is appropriately processed to generate a decoded picture, which in some examples is buffered in a memory circuit (not shown) and can be used as a reference picture.
[0090] The entropy encoder (625) is configured to format the bitstream to include the encoded block. The entropy encoder (625) is configured to include various information according to an appropriate standard such as the HEVC standard. In one example, the entropy encoder (625) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other appropriate information in the bitstream. Note that there is no residual information when encoding a block in either the inter mode or the dual prediction mode fusion submode according to the disclosed subject matter.
[0091] FIG. 7 shows a diagram of a video encoder (710) according to another embodiment of the present disclosure. The video decoder (710) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (710) is used in place of the video decoder (310) in the example of FIG. 3.
[0092] In the example of FIG. 7, the video decoder (710) includes an entropy decoder (771), an inter decoder (780), a residual decoder (773), a reconstruction module (774), and an intra decoder (772) together as shown in FIG. 7.
[0093] The entropy decoder (771) may be configured to reconstruct from the encoded picture specific symbols representing the generated syntax elements of the encoded picture. Such symbols may include, for example, the encoded mode of a block (e.g., the latter two of intra mode, inter mode, bi - directional mode, merge sub - mode or another sub - mode), prediction information (e.g., intra prediction information or inter prediction information) that can identify specific samples or metadata used for prediction by the intra decoder (772) or the inter decoder (780) respectively, residual information in the form of, for example, quantized transform coefficients, etc. In one example, when the prediction mode is an inter or bi - directional prediction mode, the inter prediction information is provided to the inter decoder (780), and when the prediction type is an intra prediction type, the intra prediction information is provided to the intra decoder (772). The residual information is inverse - quantized and provided to the residual decoder (773).
[0094] The inter decoder (780) is configured to receive the inter prediction information and generate an inter prediction result based on the inter prediction information.
[0095] The intra decoder (772) is configured to receive the intra prediction information and generate a prediction result based on the intra prediction information.
[0096] The residual decoder (773) is configured to perform inverse quantization to extract the inverse quantized transform coefficients, process the inverse quantized transform coefficients, and convert the residual from the frequency domain to the spatial domain. The residual decoder (773) may also require specific control information (for including the Quantizer Parameter: QP). This information may be provided by the entropy decoder (771) (since this is only low-capacity control information, the data path is not shown).
[0097] The reconstruction module (774) is configured to combine, in the spatial domain, the residual as the output by the residual decoder (773) and the prediction result (optionally as the output by the inter or intra prediction module) to form a reconstruction block that can be part of a reconstructed picture and, on the other hand, can be part of a reconstructed video. Other appropriate operations such as a deblocking operation can be performed to improve the visual quality.
[0098] Note that the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using any appropriate technology. In one embodiment, the video encoders (303), (503), and (603), and the video decoders (310), (410), and (710) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (303), (503), and (503), and the video decoders (310), (410), and (710) can be implemented using one or more processors that execute software instructions.
[0099] FIG. 8 is a schematic diagram of the spatial neighboring blocks and temporal neighboring blocks of the current block (801) according to one embodiment. The list of motion information candidates for the current block (801) can be configured according to the normal merge mode based on the motion information of one or more of the spatial neighboring blocks and / or temporal neighboring blocks. The motion information candidates in the list of motion information candidates configured according to the normal merge mode are also referred to as normal merge candidates in the present disclosure.
[0100] For example, in HEVC, a merge mode for inter-picture prediction is introduced. When a merge flag (including a skip flag) carried in an encoded video bitstream is signaled as true, a merge index may also be signaled to indicate which candidate in a list of motion information candidates (also called a merge candidate list) configured using the merge mode should be used to indicate the motion vector of the current block. In a decoder, the merge candidate list is configured based on the motion information (i.e., candidates) of spatial neighboring blocks and / or temporal neighboring blocks of the current block. Such a merge mode is also called a normal merge mode in the present disclosure to be distinguishable from an affine merge mode which will be further described with reference to FIG. 9 only.
[0101] Referring to FIG. 8, the motion vector (for one-direction prediction) or the motion vectors (for two-direction prediction) of the current block (801) can be predicted based on a list of motion information candidates configured using the normal merge mode. The list of motion information candidates can be derived based on the motion information of neighboring blocks, such as spatial neighboring blocks denoted as A0, A1, B0, B1, B2 (802 to 806 respectively), and / or temporal neighboring blocks denoted as C0 and C1 (812 and 813 respectively). In some examples, the spatial neighboring blocks A0, A1, B0, B1, and B2, and the current block (801) belong to the same picture. In some examples, the temporal neighboring blocks C0 and C1 belong to a reference picture. Block C0 corresponds to a position outside the current block (801) and is adjacent to the lower right corner of the current block (801). Block C1 corresponds to a position at the lower right side of the center of block (801) and may be adjacent thereto.
[0102] In some examples, to construct a list of motion information candidates using the normal merge mode, neighboring blocks A1, B1, B0, A0, and B2 are checked in order. If any of the checked blocks contains favorable motion information (e.g., a valid candidate), the valid candidate can be added to the list of motion information candidates. A pruning operation can be performed to avoid including duplicate motion information candidates in the list.
[0103] Temporal neighboring blocks can be checked to add temporal motion information to the list. Temporal neighboring blocks can be checked after spatial neighboring blocks. In some examples, the motion information of block C0 is added to the list as a temporal motion information candidate if available. If block C0 is encoded in inter mode or not available, the motion information of block C1 can be used instead as the temporal motion information candidate.
[0104] In some examples, after checking and / or adding spatial and temporal motion information candidates to the list of motion information candidates, a zero motion vector can be added to the list.
[0105] FIG. 9 is a schematic diagram of spatial neighboring blocks and temporal neighboring blocks of a current block (901) according to an embodiment. The list of motion information candidates for the current block (901) can be constructed using the affine merge mode based on the motion information of one or more of the spatial neighboring blocks and / or temporal neighboring blocks. FIG. 9 shows the current block (901), its spatial neighbors shown as A0, A1, A2, B1, B2, B3 (902, 903, 907, 904, 905, 906, 908 respectively), and the temporal neighboring block shown as C0 (912). In some examples, the spatial neighboring blocks A0, A1, A2, B0, B1, B2, and B3, and the current block (901) belong to the same picture. In some examples, the temporal neighboring block C0 belongs to a reference picture, corresponds to a position outside the current block (901), and is adjacent to the lower right corner of the current block (901).
[0106] In some examples, the motion vector of the current block (901), and / or the sub-blocks of the current block, can be derived using an affine model (e.g., a 6-parameter affine model or a 4-parameter affine model). In some examples, the affine model has 6 parameters (e.g., a 6-parameter affine model) to describe the motion vector of the block. In one example, the 6 parameters of the affine-coded block can be represented by 3 motion vectors (also called 3 control point motion vectors) at 3 different positions of the block (e.g., control points CP0, CP1, and CP2 at the upper left, upper right, and lower left corners of FIG. 9). In another example, a simplified affine model uses 4 parameters to describe the motion information of the affine-coded block. These can be represented by 2 motion vectors (also called 2 control point motion vectors) at 2 different positions of the block (e.g., control points CP0 and CP1 at the upper left and upper right corners of FIG. 9).
[0107] The list of motion information candidates can be constructed using the affine merge mode (also called the affine merge candidate list). In the present disclosure, the motion information candidates in the list of motion information candidates constructed according to the affine merge mode are also called affine merge candidates. In some examples, when the current block has a width and height of 8 samples or more, the affine merge mode can be applied. According to the affine merge mode, the control point motion vector (CPMV) of the current block can be generated based on the motion information of spatially neighboring blocks. In some examples, the list of motion information candidates can include up to 5 CPMV candidates, and an index can be signaled to indicate which CPMV candidate should be used for the current block.
[0108] In some embodiments, the affine merge candidate list can have three types of CPMV candidates, including inherited affine candidates, constructed affine candidates, and zero MVs. The inherited affine candidates can be derived by estimation from the CPMVs of neighboring blocks. The constructed affine candidates can be derived using the translational MVs of neighboring blocks.
[0109] In the example of VTM, there can be at most two inherited affine candidates, which are derived from the corresponding affine motion models of neighboring blocks including one from the left neighboring blocks (A0 and A1) and one from the upper neighboring blocks (B0, B1, B2). For the candidate from the left, the neighboring blocks A0 and A1 are checked in order, and the first available inherited affine candidate from the neighboring blocks A0 and A1 is used as the inherited affine candidate from the left. For the candidate from the top, the neighboring blocks B0, B1, and B2 are checked in order, and the first available inherited affine candidate from the neighboring blocks B0, B1, and B2 is used as the inherited affine candidate from the top. In some examples, no pruning check is performed between the two inherited affine candidates.
[0110] Once the neighboring affine blocks are identified, the corresponding inherited affine candidates to be added to the affine merge list of the current CU can be derived from the control point motion vectors of the neighboring affine blocks. In the example of FIG. 9, when the neighboring block A1 is encoded according to the affine mode, the motion vectors of the upper left corner (control point CP0 A1 ), upper right corner (control point CP1 A1 ), and lower left corner (control point CP2 A1 ) of the block A1 can be obtained. When the block A1 is encoded using the 4-parameter affine model, the two CPMVs as the inherited affine candidates of the current block (901) can be calculated according to the motion vectors of the control point CP0 A1 and the control point CP1 A1 . When the block A1 is encoded using the 6-parameter affine model, the three CPMVs as the inherited affine candidates of the current block (901) are the control point CP0 A1 , the control point CP1 A1and control point CP2 A1 can be calculated according to the motion vector of A1 .
[0111] Furthermore, the configured affine candidates can be derived by combining the neighboring translational motion information of each control point. The motion information of control points CP0, CP1, and CP2 is derived from the specified spatial neighboring blocks A0, A1, A2, B0, B1, B2, and B3.
[0112] For example, CPMV k (k = 1, 2, 3, 4) represents the motion vector of the k-th control point. Here, CPMV 1 corresponds to control point CP0, CPMV 2 corresponds to control point CP1, CPMV 3 corresponds to control point CP2, and CPMV 4 corresponds to the temporal control point based on the temporal neighboring block C0. For CPMV 1 , the neighboring blocks B2, B3, and A2 can be checked in order, and the first available motion vector from the neighboring blocks B2, B3, and A2 is used as CPMV 1 . For CPMV 2 , the neighboring blocks B1 and B0 can be checked in order, and the first available motion vector from the neighboring blocks B1 and B0 is used as CPMV 2 . For CPMV 3 , the neighboring blocks A1 and A0 can be checked in order, and the first available motion vector from the neighboring blocks A1 and A0 is used as CPMV 3 . Furthermore, the motion vector of the temporal neighboring block C0 can be used as CPMV 4 if available.
[0113] After the CPMV 1 , CPMV 2 , CPMV 3 , and CPMV 4 of the four control points CP0, CP1, CP2, and the temporal control point are obtained, the affine merge candidate list can be configured to include the affine merge candidates configured in the following order: {CPMV 1 , CPMV 2,CPMV 3}, {CPMV 1 ,CPMV 2 ,CPMV 4}, {CPMV 1 ,CPMV 3 ,CPMV 4}, {CPMV 2 ,CPMV 3 ,CPMV 4}, {CPMV 1 ,CPMV 2} and {CPMV 1 ,CPMV 3}. Any combination of three CPMVs can form a six-parameter affine merge candidate, and any combination of two CPMVs can form a four-parameter affine merge candidate. In some examples, to avoid motion scaling processing, if the reference indices of a set of control points are different, the corresponding combination of CPMVs can be discarded.
[0114] 10A is a schematic diagram of spatial neighboring blocks that can be used to determine predicted motion information of a current block (1011) using a sub-block-based temporal MV prediction method based on motion information of spatial neighboring blocks according to one embodiment. FIG. 10A shows a current block (1011) and its spatial neighboring blocks denoted as A0, A1, B0, and B1 (1012, 1013, 1014, and 1015, respectively). In some examples, the spatial neighboring blocks A0, A1, B0, and B1 and the current block (1011) belong to the same picture.
[0115] 10B is a schematic diagram of determining motion information for a sub-block of a current block (1011) using a sub-block-based temporal MV prediction method based on a selected spatial neighboring block, such as block A1 in this non-limiting example, according to one embodiment. In this example, the current block (1011) is in a current picture (1010), and a reference block (1061) is in a reference picture (1060) and can be identified based on a motion shift (or displacement) between the current block (1011) and the reference block (1061) indicated by a motion vector (1022).
[0116] In some embodiments, similar to temporal motion vector prediction (TMVP) in HEVC, sub-block based temporal MV prediction (SbTMVP) uses motion information within various reference sub-blocks in a reference picture for a current block within a current picture. In some embodiments, the same reference pictures used by TMVP can be used for SbTVMP. In some embodiments, TMVP predicts motion information at the CU level, while SbTVMP predicts motion at the sub-CU level. In some embodiments, TMVP uses a temporal motion vector from a block at the same position within a reference picture that has a corresponding position adjacent to the lower right corner or the center of the current block, and SbTVMP uses a temporal motion vector from a reference block that can be identified by performing a motion shift based on a motion vector from one of the spatial neighboring blocks of the current block.
[0117] For example, as shown in FIG. 10A, neighboring blocks A1, B1, B0, and A0 can be sequentially checked in the SbTVMP process. For example, as soon as a first spatial neighboring block having a motion vector that uses the reference picture (1060) as its reference picture, such as block A1 having a motion vector (1022) pointing to reference block AR1 within the reference picture (1060), is identified, this motion vector (1022) can be used to perform a motion shift. When such motion vectors are not available from the spatial neighboring blocks A1, B1, B0, and A0, the motion shift is set to (0, 0).
[0118] After determining the motion shift, the reference block (1061) can be identified based on the position of the current block (1011) and the determined motion shift. In FIG. 10B, the reference block (1061) can be further divided into 16 sub-blocks by the reference motion information MRa to MRp. In some examples, the reference motion information of each sub-block within the reference block (1061) can be determined based on the minimum motion grid covering the central sample of such sub-block. The motion information can include a motion vector and a corresponding reference index. The current block (1011) can be further divided into 16 sub-blocks, and the motion information MVa to MVp of the sub-blocks within the current block (1011) can be derived from the reference motion information MRa to MRp in a manner similar to the TMVP process, and in some examples by temporal scaling.
[0119] The sub-block size used in the SbTVMP process can be fixed (or predetermined) or signaled. In some examples, the sub-block size used in the SbTVMP process can be 8×8 samples. In some examples, the SbTVMP process is applicable only to blocks having a width and height of a fixed or signaled size, such as 8 pixels or more.
[0120] In an example of VTM, a merge list based on combined sub-blocks, including SbTVMP candidates and affine merge candidates, is used for signaling the merge mode based on sub-blocks. The SbTVMP mode can be enabled or disabled by a sequence parameter set (SPS) flag. In some examples, when the SbTMVP mode is enabled, the SbTMVP candidate is added as the first entry in the list of merge candidates based on sub-blocks, followed by the affine merge candidates. In some embodiments, the maximum allowable size of the merge list based on sub-blocks is set to 5. However, in other embodiments, other sizes may be utilized.
[0121] In some embodiments, the encoding logic for additional SbTMVP merge candidates is the same as for other merge candidates. That is, for each block in a P or B slice, an additional rate distortion check can be performed to determine whether to use the SbTMVP candidate.
[0122] FIG. 11A is a flowchart showing an overview of a process (1100) of constructing and updating a list of motion information candidates using a history based MV prediction (HMVP) method according to an embodiment. Motion information candidates in the list constructed according to the MHVP method are also referred to as HMVP candidates in the present disclosure.
[0123] In some embodiments, a list of motion information candidates (also referred to as a history list) using the HMVP method can be constructed and updated during the encoding or decoding process. The history list can be emptied when a new slice starts. In some embodiments, regardless of whether there are newly encoded or decoded inter-coded non-affine blocks, the relevant motion information can be added as a new HMVP candidate to the last entry of the history list. Therefore, before processing (encoding or decoding) the current block, a history list with HMVP candidates can be loaded (S1112). The current block can be encoded or decoded using the HMVP candidates in the history list (S1114). Later, the history list can be updated using the motion information to encode or decode the current block (S1116).
[0124] FIG. 11B is a schematic diagram of updating a list of motion information candidates using a history based MV prediction method according to an embodiment. FIG. 11B shows a history list having a size L, and each candidate in the list can be identified by an index in the range of 0 to L - 1. L is an integer greater than or equal to 0. Before encoding or decoding the current block, the history list (1120) has L candidates: HMVP 1 、HMVP 2 、...、HMVP m 、...、HMVP L-2 、and HMVPL-1 including. Here, m is an integer in the range of 0 to L. After encoding or decoding the current block, a new entry HMVP c is added to the history list.
[0125] In the example of VTM, the size of the history list can be set to 6, which indicates that up to 6 HMVP candidates can be added to the history list. When inserting a new motion candidate (e.g., HMVP c ) into the history list, a constrained first-in-first-out (FIFO) rule can be used. Here, a redundancy check is first applied to find out whether there is a redundant HMVP in the history list. If no redundant HMVP is found, the first HMVP candidate (which is HMVP 1 in the example of FIG. 11B and has index = 0) is removed from the list, and all other HMVP candidates are then moved forward, for example, the index is decreased by only 1. The new HMVP c candidate can be added to the last entry of the resulting list (as shown in the resulting list (1130), for example, having index = L - 1 in FIG. 11B). On the other hand, if a redundant HMVP (such as HMVP 2 in FIG. 11B) is found, the redundant HMVP in the history list is removed from the list, and all subsequent HMVP candidates are moved forward, for example, the index is decreased by only 1. The new HMVP c candidate can be added to the last entry of the resulting list (as shown in the resulting list (1140), for example, having index = L - 1 in FIG. 11B).
[0126] In some examples, the HMVP candidates can be used in the merge candidate list construction process. For example, the latest HMVP candidates in the list can be checked in order and inserted into the candidate list after the TMVP candidates. In some embodiments, pruning can be applied to the HMVP candidates for spatial or temporal merge candidates instead of sub-block motion candidates (i.e., SbTMVP candidates).
[0127] In some embodiments, in order to reduce the number of trimming operations, one or more of the following rules may be possible next: (a) The number of HMPV candidates to be checked indicated by M is set as follows: M = (N <= 4)? L : (8 - N). Here, N indicates the number of available sub-block merge candidates, and L indicates the number of available HMVP candidates in the history list. (b) Further, when the total number of available merge candidates is one less than the maximum number of merge candidates signaled, the merge candidate list construction process from the HMVP list may be terminated. (c) Further, the number of pairs for deriving combined two-way prediction merge candidates can be reduced from 12 to 6.
[0128] In some embodiments, HMVP candidates may be used in the AMVP candidate list construction process. The motion vectors of the latest K HMVP candidates in the history list can be added to the AMVP candidates after the TMVP candidates. In some examples, only HMVP having the same reference picture as the AMVP target reference picture should be added to the AMVP candidate list. Trimming can be applied to HMVP candidates. In some examples, K is set to 4 and the AMVP list size is kept invariant, for example, equal to 2.
[0129] In some examples, the list of motion information candidates can be constructed by the various list construction processes described above and / or any other applicable list construction processes. After the list of motion information candidates is constructed, an additional average reference motion vector can be determined by averaging the already determined reference motion information for each pair according to a predetermined pairing.
[0130] In some examples, there may be up to four reference motion vectors determined to be included in the reference motion candidate list for encoding the current block. These four reference motion vectors may be arranged in the list in association with an index (e.g., 0, 1, 2, and 3) indicating the order of the reference motion vectors in the list. Additional average reference motion vectors may be determined respectively based on pairs defined, for example, as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}. In some examples, if both reference motion vectors of a given pair are available in the list of reference motion vectors, these two reference motion vectors can still be averaged even when they point to different reference pictures.
[0131] Furthermore, current picture referencing (CPR) is sometimes referred to as intra block copy (IBC). Here, the motion vector of the current block represents samples already reconstructed within the current picture. CPR is supported in the HEVC screen content coding extension (HEVC SCC). The CPR encoded block can be signaled as an inter encoded block. The luma motion (or block) vector of the CPR encoded block can be provided with integer precision. Similarly, the chroma motion vector can be clipped with integer precision. When combined with AMVR, the CPR mode can switch between 1 pixel (pel) and 4 pixel (pel) motion vector precision. The current picture can be placed at the end of the reference picture list L0. To reduce memory consumption and decoder complexity, CPR in an example of VTM may allow reference only to the reconstructed part of the current CTU. This reduction enables implementing the CPR mode using local on-chip memory for hardware implementation.
[0132] On the encoder side, hash-based motion estimation is performed for CPR. The encoder can perform rate distortion checks for blocks that no longer have a width or height greater than 16 luma samples. In non-merge mode, block vector search can first be performed using hash-based search. If the hash search does not return valid candidates, a local search based on block matching can be performed.
[0133] In some examples, when performing a hash-based search, the hash key comparison (32-bit CRC) between the current block and the reference block can be extended to all allowable block sizes. The hash key calculation for all positions within the current picture can be based on 4×4 sub-blocks. For larger-sized current blocks, when all the hash keys of the 4×4 sub-blocks match the hash keys at the corresponding reference positions, the hash key can be determined to match the hash key of the reference block. In some examples, if it is found that the hash keys of multiple reference blocks match the hash key of the current block, the cost of each matching reference block vector is calculated, and the one with the minimum cost can be selected.
[0134] In some examples, when performing a block matching search, the search range can be set to R samples to the left above the current block within the current CTU. At the start of the CTU, the value of R can be initially set to 128 if there is no temporal reference picture, and can be initially set to 64 if at least one temporal reference picture exists. The hash hit rate can be defined as the ratio of samples within the CTU that had a match using hash-based search. In some embodiments, during encoding the current CTU, if the hash hit rate is lower than 5%, R can be reduced by half.
[0135] For some applications, for blocks reconstructed using two - direction prediction, a generalized bi - prediction (GBi) index can be used and included in the motion information of the block. In some examples, the GBi index indicates a first weight applicable to the first reference picture from list 1, and a second weight applicable to the second reference picture from list 0 can be derived from the first weight. In some examples, the GBi index can represent the derived weight parameter w, and the first weight w 1 is, w 1 =w / F, where F represents a precision factor.
[0136] Also, the second weight w 0 is, w 0 =1 - w 1 can be determined according to.
[0137] In some examples, when the precision factor F is set to 8 (i.e., having a sample precision of 1 / 8), the bi - prediction P bi-pred of the current block is, =((8 - w)*P 0 +w*P 1 +4)>>3. Here, P 0 and P 1 are reference samples from the reference pictures in list 0 and list 1 respectively.
[0138] In some examples, with the precision factor F set to 8, the GBi index can be assigned to represent one of the 5 weights {-2 / 8, 3 / 8, 4 / 8, 5 / 8, 10 / 8} available for low - delay pictures and 3 weights {3 / 8, 4 / 8, 5 / 8} for non - low - delay pictures. The weight to be signaled using the GBi index can be determined by rate - distortion cost analysis. In some examples, the weight check order can be {4 / 8, -2 / 8, 10 / 8, 3 / 8, 5 / 8} at the encoder.
[0139] In some applications that use the advanced motion vector prediction (AMVP) mode, when a particular CU is coded by dual prediction, the weight parameter selection can be explicitly signaled at the CU level using the GBi index. If dual prediction is used and the CU area is smaller than 256 luma samples, the GBi index signaling can be discarded.
[0140] In some applications, in the interlaced mode from spatial candidates and the inherited affine merge mode, the weight parameter selection can be inherited from the merge candidate selected based on the GBi index. For other merge types, such as temporal merge candidates, HMVP candidates, SbTMVP candidates, configured affine merge candidates, pairwise averaging candidates, etc., the weights may be set to represent a default weight (e.g., 1 / 2 or 4 / 8), and the GBi index from the selected candidate is not inherited.
[0141] The inheritance of the GBi index can sometimes be advantageous in some merge modes but not in others, and the disadvantage due to the additional memory space required to record the GBi index from the previously coded block can sometimes be significant. Thus, in addition to the above examples, in some embodiments, the GBi index inheritance function can be further configured for the application of various scenarios as further described below. The GBi index inheritance configurations as described below can be used separately or combined in any order. In some examples, not all of the following GBi index inheritance configurations are implemented simultaneously.
[0142] According to a first GBi index inheritance configuration among some examples, the GBi index inheritance can be enabled for HMVP merge candidates. In some embodiments, the GBi index of the dual-predicted coded block can be stored in each entry of the HMVP candidate list. Thus, according to the first GBi index inheritance configuration, the management of HMVP candidates can be coordinated with the management of motion information candidates configured or obtained according to other modes such as the normal merge mode.
[0143] In one embodiment, when any entry in the HMVP candidate list is updated, the motion information (e.g., including MV, reference list, and reference index) and the corresponding GBi index from the coded block can be updated as well. For example, after obtaining the prediction information of a block, if it is determined that the motion information candidate should be stored in the HMVP candidate list according to the prediction information of the block, the motion information candidate can be stored to include at least the motion information of the block and the weight parameter for performing the bidirectional prediction of the first block when the first block is coded according to bidirectional prediction. Also, the motion information candidate can be stored to include at least the motion information of the block and the specified weight parameter when the block is coded according to unidirectional prediction.
[0144] The normal interleave candidate from the HMVP candidate list is selected as the MV predictor of the current coded block in the decoding process. When the candidate is coded using bidirectional prediction (also called dual prediction), the GBi index of the current block can be set to the value of the GBi index value from the candidate. Alternatively, when the HMVP merge candidate is coded using unidirectional prediction (also called single prediction), the GBi index of the current block can be set to the specified GBi index corresponding to a weight of 1 / 2 in some examples.
[0145] According to the second GBi index inheritance configuration in some examples, GBi index inheritance can be enabled only for spatial interleave candidates. In some embodiments, GBi index inheritance can be enabled only for normal spatial interleave candidates. Thus, according to the second GBi index inheritance configuration, the implementation of the GBi index inheritance function can be simplified to reduce the computational complexity.
[0146] In some embodiments, after obtaining the prediction information of a block, when it is determined that the motion information candidate should be stored in a normal merge candidate list according to the prediction information of the block, the motion information candidate can be stored to include at least the motion information of the block and the weight parameter for performing the bidirectional prediction of the first block when the first block is encoded according to bidirectional prediction. Also, the motion information candidate can be stored to include at least the motion information of the block and a specified weight parameter when the block is encoded according to unidirectional prediction.
[0147] In one embodiment, when a spatial merge candidate is selected in the decoding process of the current block and the candidate is bi-predicted, the GBi index of the current block can be set to the GBi index value from the candidate block. Alternatively, when the spatial merge candidate is uni-predicted, the GBi index of the current block can be set to a specified GBi index corresponding to a weight of 1 / 2 in some examples. In some examples, when the merge candidate selected for the current block is not a normal spatial merge candidate such as a temporal merge candidate, an affine model inheritance candidate, an SbTMVP candidate, an HMVP candidate, or a per-pair averaging candidate, etc., the GBi index of the current block can be set to a specified GBi index, for example, to reduce the memory load.
[0148] According to a third GBi index inheritance configuration among some examples, the GBi index inheritance can be disabled for normal spatial merge candidates from the motion data line buffer. Thus, according to the third GBi index inheritance configuration, the computational complexity can be simplified and the memory space for storing the motion information candidates can be reduced.
[0149] Specifically, when the current block is located in the top row of the current CTU, the spatial merge candidate from the position above the current block can be stored in the motion data line buffer. The motion data line buffer is used to store the motion information of the last row of the smallest allowable size of the inter-block at the bottom of the CTU in the previous CTU row.
[0150] In one embodiment, the inheritance of the GBi index of the normal spatial merge candidates from the motion data line buffer can be disabled to save the memory storage of the motion data line buffer. For example, only the normal motion information including the motion vector values and reference index values of list 0 and / or list 1 may be stored in the motion data line buffer. The motion data line buffer can avoid storing the GBi index values corresponding to the blocks in the previous CTU row. Thus, in some examples, when the selected spatial merge candidate for the current block is from a block located in the previous CTU row above the CTU boundary, the motion information can be loaded from the motion data line buffer, and the GBi index of the current block can be set to the predefined GBi index corresponding to a weight of 1 / 2 in some examples.
[0151] FIG. 12 is a schematic diagram of a current block (801) at a CTU boundary (1210) and its spatial neighboring blocks according to one embodiment. In FIG. 12, components that are the same as or similar to those shown in FIG. 8 are given the same reference numerals or labels, and their detailed descriptions are provided above with respect to FIG. 8.
[0152] For example, as shown in FIG. 12, for the current block (801), if the selected spatial merge candidate corresponds to the neighboring blocks B0, B1, or B2, their motion information can be stored in the motion data line buffer. Thus, in some examples, according to the third GBi index inheritance configuration, the GBi index of the current block (801) may be set to the predefined GBi index without referring to the GBi index value of the selected spatial merge candidate.
[0153] In some examples, according to the fourth GBi index inheritance configuration, GBi index inheritance can be disabled for normal affine merge candidates from the motion data line buffer. Specifically, when the current block is located at the top row of the current CTU, the affine motion information of the spatial neighboring blocks from the position above the current block for deriving the inherited affine merge candidates of the current block may be stored in the motion data line buffer. The motion data line buffer is used to store the motion information of the last row of the smallest allowable size of the inter-block at the bottom of the CTU in the previous CTU row.
[0154] In one embodiment, the GBi index inheritance of the affine merge candidates from the motion data line buffer can be disabled to save the memory storage of the motion data line buffer. For example, the GBi index value may not need to be stored in the motion data line buffer. When the selected inherited affine merge candidate for the current block is derived from an affine encoded block located above the CTU boundary, the affine motion information is loaded from the motion data line buffer. The GBi index of the current block can be set to a predefined GBi index corresponding to a weight of 1 / 2 in some examples.
[0155] FIG. 13 is a schematic diagram of a current block (901) at the CTU boundary (1310) and its spatial affine neighboring blocks according to one embodiment. In FIG. 13, components that are the same or similar to those shown in FIG. 9 are given the same reference numerals or labels, and their detailed descriptions are provided above with respect to FIG. 9.
[0156] For example, as shown in FIG. 13, for the current block (901), when the selected inherited affine merge candidate corresponds to the neighboring blocks B0, B1, or B2, their affine motion information can be stored in the motion data line buffer. Thus, in some examples, according to the fourth GBi index inheritance configuration, the GBi index of the current block (901) may be set to the predefined GBi index without referring to the GBi index value of the selected affine merge candidate.
[0157] In some examples, according to the fifth GBi index inheritance configuration, the GBi index inheritance may be limited to the reference blocks within the current CTU. In some embodiments, the GBi index information can only be stored for the blocks within the current CTU. When using spatial MV prediction, if the predictor is from a block outside the current CTU, the GBi index information of the reference block may not need to be stored for the reference block. Therefore, in some examples, the GBi index of the current block can be set to a specified GBi index corresponding to a weight of 1 / 2.
[0158] In one embodiment, the above limitation according to the fifth GBi index inheritance configuration can only be applied to the spatial interleave of translational motion information.
[0159] In another embodiment, the above limitation according to the fifth GBi index inheritance configuration can only be applied to the inherited affine merge candidates.
[0160] In another embodiment, the above limitation according to the fifth GBi index inheritance configuration can only be applied to both the spatial interleave of translational motion information and the inherited affine merge candidates.
[0161] In some examples, according to the sixth GBi index inheritance configuration, the use of negative weights as indicated by the GBi index may be restricted such that the negative weight can only be used when the two reference pictures used for the block belong to the same picture. In some examples, according to the sixth GBi index inheritance configuration, the use of positive weights as indicated by the GBi index may be restricted such that the positive weight can only be used when the two reference pictures used for the block belong to different pictures. In some examples, according to the sixth GBi index inheritance configuration, the coding efficiency can be improved in some cases.
[0162] In some examples, pictures having the same picture order count (POC) value are determined as the same picture. In some examples, the signaling of GBi information and the reference picture of the current block can be adjusted according to the above constraints.
[0163] In some examples, according to the seventh GBi index inheritance configuration, GBi index inheritance can be enabled for current picture referencing (CPR).
[0164] In examples of VVC and VTM, under the CPR mode, the spatial merge candidate selected for the current block can be from the same picture as the current block. In some embodiments, the CPR implementation can only enable the encoding of blocks using one - direction prediction. However, in other embodiments, the CPR implementation can enable the encoding of blocks using two - direction prediction. When the current block is encoded using two - direction prediction by one reference block selected from the current picture, the GBi index can be used for the current block.
[0165] For example, in one embodiment, when the current block is encoded in merge mode, the selected merge candidate is bi - predicted by the GBi index, and both reference blocks are from the current picture (CPR mode), the GBi index of the current block can be set to be the same as the GBi index of the candidate. When the selected merge candidate is uni - predicted using the CPR mode, the GBi index of the current block can be set to a predefined GBi index corresponding to a weight of 1 / 2 in some examples. In some embodiments, when the current block is a bi - prediction block that references two different reference pictures, the GBi index of the current block can be set to the predefined GBi index, and the GBi index is not signaled.
[0166] In some embodiments, a block is encoded according to one of bi - directional prediction and uni - directional prediction (CPR mode) with the current picture as the reference picture. In some examples, after obtaining the prediction information of the block, when it is determined that the motion information candidate should be stored in the normal merge candidate list according to the prediction information of the block, the motion information candidate can be stored to include at least the motion information of the block and the weight parameter for performing the bi - directional prediction of the first block when the first block is encoded according to bi - directional prediction. Also, the motion information candidate can be stored to include at least the motion information of the block and the specified weight parameter when the block is encoded according to uni - directional prediction.
[0167] FIG. 14 shows a flowchart illustrating an overview of a decoding process (1400) according to an embodiment of the present disclosure. The process (1400) can be used when constructing a list of motion information candidates to reconstruct a block of a picture (i.e., the current block). In some embodiments, one or more operations are performed before or after the process (1400), and some of the operations shown in FIG. 14 may be reordered or omitted.
[0168] In various embodiments, the process (1400) is executed by a processing circuit such as a processing circuit in the terminal devices (210), (220), (230), and (240), or a processing circuit that performs the functions of the video decoders (310), (410), or (710). In some embodiments, the process (1400) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1400). The process starts at (S1401) and proceeds to (S1410).
[0169] (S1410) In this step, the prediction information of the first block in the picture is obtained from the encoded video bitstream. The prediction information indicates a mode that constitutes a list of motion information candidates from one or more previous decoded blocks. In some embodiments, the mode that constitutes the list of motion information candidates includes various modes as described above, such as the HEVC merge mode, the affine merge mode, temporal MV prediction based on sub-blocks, MV prediction based on history, average MV candidates per pair, CPR, and / or other applicable processes, or combinations thereof. In some examples, the prediction information can be obtained using the systems or decoders shown in FIGS. 3, 4, and 7.
[0170] (S1420) In this step, the reconstructed samples of the first block are generated according to one of bidirectional prediction and unidirectional prediction (e.g., for output) and the prediction information. In some examples, the reconstructed samples of the first block can be generated using the systems or decoders shown in FIGS. 3, 4, and 7.
[0171] (S1430) In this step, when it is determined that the motion information candidate should be stored in the list of motion information candidates according to the prediction information of the first block, the motion information candidate can be stored.
[0172] In some embodiments, GBi index inheritance from HMVP merge candidates is enabled. Therefore, the motion information candidate can be stored as an HMVP candidate that includes at least the first information and the first weight parameter indicating the first weight for performing the bidirectional prediction of the first block when the first block is encoded according to bidirectional prediction. Also, the motion information candidate as an HMVP candidate may include the first motion information and the specified weight parameter indicating the specified weight when the first block is encoded according to unidirectional prediction. In some examples, the specified weight is 1 / 2.
[0173] In some embodiments, GBi index inheritance can be enabled for CPR. Thus, the motion information candidate can be stored to include at least, when the first block is encoded according to bidirectional prediction, the first information and the first weight parameter for performing the bidirectional prediction of the first block. Also, the motion information candidate can include the first motion information and the specified weight parameter indicating the specified weight when the first block is encoded according to unidirectional prediction.
[0174] In some embodiments, GBi index inheritance can be enabled for spatial interleave candidates. The motion information candidate can be stored to include at least, when the first block is encoded according to bidirectional prediction, the first information and the first weight parameter for performing the bidirectional prediction of the first block. Also, the motion information candidate can include the first motion information and the specified weight parameter indicating the specified weight when the first block is encoded according to unidirectional prediction.
[0175] In some examples, the motion information candidate can be stored using the systems or decoders shown in FIGS. 3, 4, and 7.
[0176] In some embodiments, the first block is encoded according to bidirectional prediction based on the first reference picture in the first list (e.g., list 1) and the second reference picture in the second list (e.g., list 0). Here, the weight w applicable to the first reference picture 1 can be determined according to w 1 = w / F.
[0177] Also, the weight w applicable to the second reference picture 0 can be determined according to w 0 = 1 - w 1 Here, w and F are integers, w represents the derived weight parameter (e.g., the weight indicated by the GBi index) to be included in the stored motion information candidate, and F represents the accuracy coefficient. In some examples, the accuracy coefficient F is 8.
[0178] wherein the bidirectional prediction is performed.
[0179] In some examples, the reconstructed samples of the current block can be generated using the systems or decoders shown in FIGS. 3, 4, and 7.
[0180] (S1460), when it is determined that the second block in the picture should be decoded based on the motion information candidate, the reconstructed samples of the second block can be generated for output according to the motion information candidate. In some examples, the reconstructed samples of the second block can be generated using the systems or decoders shown in FIGS. 3, 4, and 7.
[0181] In some embodiments, in (S1460), when the motion information candidate is stored as a normal spatial merge candidate and the second block is encoded according to bidirectional prediction, the second weight for performing the bidirectional prediction of the second block can be set according to the first weight parameter stored in the motion information candidate when the first block is spatially adjacent to the second block. Also, when the first block is not spatially adjacent to the second block, the second weight for performing the bidirectional prediction of the second block can be set to a default weight.
[0182] In some embodiments, in (S1460), when the motion information candidate is stored as a candidate that is neither a normal spatial merge candidate nor an HMVP candidate and the second block is encoded according to bidirectional prediction, the second weight for performing the bidirectional prediction for the second block can be a default weight.
[0183] In some embodiments, when the first block is in a CTU row different from the current CTU row in which the second block is included, the motion information candidate is stored as a normal merge candidate or an affine merge candidate, and the second block is encoded according to bidirectional prediction, the second weight for performing the bidirectional prediction for the second block can be set to a default weight.
[0184] In some embodiments, the first block is outside the current CTU containing the second block, the motion information candidate is stored as a translational merge candidate or an inherited affine merge candidate, and when the second block is encoded according to bidirectional prediction, the second weight for performing bidirectional prediction on the second block can be set to a specified weight.
[0185] In some embodiments, the first block is encoded according to bidirectional prediction, and both the first weight corresponding to the first reference picture in the first list and the second weight derived from the first weight and corresponding to the second reference picture in the second list are positive when the first and second reference pictures are the same reference picture.
[0186] In some embodiments, the first block is encoded according to bidirectional prediction, and one of the first weight corresponding to the first reference picture in the first list and the second weight derived from the first weight and corresponding to the second reference picture in the second list is negative when the first and second reference pictures are different reference pictures.
[0187] After (S1460), the process proceeds to (S1499) and ends.
[0188] FIG. 15 shows a flowchart illustrating an overview of an encoding process (1500) according to an embodiment of the present disclosure. The process (1500) can be used to encode a block of a picture (i.e., the current block) encoded using the inter mode. In some embodiments, one or more operations are performed before or after the process (1500), and some of the operations shown in FIG. 15 may be rearranged or omitted.
[0189] In various embodiments, the process (1500) is performed by a processing circuit such as a processing circuit in the terminal devices (210), (220), (230), and (240), or a processing circuit that performs the functions of the video encoders (303), (503), or (603). In some embodiments, the process (1500) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (1500). The process starts at (S1501) and proceeds to (S1510).
[0190] (S1501), first prediction information for encoding a first block in a picture is obtained. In some embodiments, the first prediction information can be obtained by any applicable motion estimation process that generates an appropriate predictor for the first block. In some examples, the determination can be performed using the systems or encoders shown in FIGS. 3, 5, and 6.
[0191] (S1530), if it is determined that a motion information candidate should be stored in a list of motion information candidates according to the prediction information of the first block, the motion information candidate can be stored. In some examples, (S1530) can be performed in a manner similar to (S1430). In some examples, storing the motion information candidate can be performed using the systems or encoders shown in FIGS. 3, 5, and 6.
[0192] (S1560), if it is determined that a second block in the picture should be decoded based on the motion information candidate, second prediction information for encoding the second block can be obtained according to the motion information candidate. In some examples, (S1560) can include a step of determining a weight for performing bidirectional prediction for the second block in a manner similar to (S1460). In some examples, the determination of the second prediction information can be performed using the systems or encoders shown in FIGS. 3, 5, and 6.
[0193] After (S1560), the process proceeds to (S1599) and ends.
[0194] The above technology can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 16 shows a computer system (1600) suitable for implementing a particular embodiment of the subject matter of the present disclosure.
[0195] The computer software can be processed by mechanisms such as assembly, compilation, linking, etc., to generate code containing instructions executable directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., and can be encoded using any suitable machine code or computer language.
[0196] The instructions can be executed on various computers or their components including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, etc.
[0197] The components shown in FIG. 16 of the computer system (1600) are exemplary in nature and do not imply any limitation to the use or functionality scope of the computer software implementing the embodiments of the present disclosure. Further, the component configuration should not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1600).
[0198] The computer system (1600) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users through, for example, sensory input (e.g., keystrokes, swipes, data grab operations), voice input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by a human, such as voice (e.g., conversation, music, ambient sound), images (e.g., scanned images, photographic images obtained from a digital camera), video (e.g., including 2D video, 3D video, stereoscopic video).
[0199] The input human interface device may include one or more of a keyboard (1601), a mouse (1602), a trackpad (1603), a touch screen (1610), a data grab (not shown), a joystick (1605), a microphone (1606), a scanner (1607), a camera (1608) (only one of which is shown).
[0200] The computer system (1600) may also include a specific human interface output device. Such a human interface output device may stimulate the senses of one or more human users, for example, through sensory output, sound, light, and smell / taste. Such a human interface output device may include a sensory output device (e.g., a touch screen (1610), a data glove (not shown), or a joystick (1605 (sensory feedback, but there may also be sensory feedback devices that do not function as input devices)), a sound output device (e.g., a speaker (1609), headphones (not shown)), a visual output device (e.g., a screen (1610), a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capability, each with or without sensory feedback capability, some of which may output two-dimensional visual output or output of three dimensions or more through means such as, for example, stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), and printers (not shown))).
[0201] The computer system (1600) may also include a human-accessible storage device and related media such as an optical medium like a CD / DVD ROM / RW (1620) with a medium (1621) such as a CD / DVD, a thumb drive (1622), a removable hard drive or a solid-state drive (1623), legacy magnetic media such as tapes and floppy disks (not shown), and devices based on dedicated ROM / ASIC / PLD such as security dongles (not shown).
[0202] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0203] The computer system (1600) may also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LET, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, terrestrial broadcast TV, vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface attached to a particular general-purpose data port or peripheral device bus (1649) (e.g., the USB port of the computer system (1600)). Others are generally integrated into the core of the computer system (1600) by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using these networks, the computer system (1600) can communicate with other entities. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used for each of the networks and network interfaces described above.
[0204] The aforementioned human interface device, human-accessible storage device, and network interface can be attached to the core (1640) of the computer system (1600).
[0205] The core (1640) may include one or more central processing units (CPUs) (1641), a graphics processing unit (GPU) (1642), a dedicated programmable processing unit in the form of a GPGA (1643), a hardware accelerator for specific tasks (1644), etc. These devices may be connected through a system bus (1648) together with a built-in mass storage device (1647) such as a read-only memory (ROM) (1645), a random access memory (1646), an internal hard drive that is not accessible to users, an SSD, etc. In some computer systems, the system bus 1648 is accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus 1648 of the core or through a peripheral device bus 1649. The architecture of the peripheral device bus includes PCI, USB, etc.
[0206] The CPU (1641), GPU (1642), FPGA (1643), and accelerator (1644) can be combined to execute specific instructions that can generate the aforementioned computer code. The computer code can be stored in the ROM (1645) or the RAM (1646). Temporary data can also be stored in the RAM (1646), while permanent data can be stored, for example, in the built-in mass storage device (1647). Fast storage and reading to any of the memory devices can be enabled through the use of a cache memory that can be closely associated with one or more of the CPU (1641), GPU (1642), mass storage device (1647), ROM (1645), RAM (1646), etc.
[0207] A computer-readable medium may have computer code for performing operations implemented by various computers. The medium and the computer code may be specially designed and configured for the purposes of the present disclosure or may be of the types well known and available to those skilled in the field of computer software.
[0208] By way of example and not limitation, a computer system (1600) having an architecture, and specifically a core (1640), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be associated with a specific storage device of the core (1640) with non-transitory characteristics such as an on-core mass storage device (1647) or ROM (1645), and media associated with a user-accessible mass storage device as described above. The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (1640). The computer-readable media can include one or more memory devices or chips as required by a particular need. The software can cause the core (1640) and specifically the processor (including a CPU, GPU, FPGA, etc.) therein to execute a particular process or a particular part of a particular process described herein, including the definition of a data structure stored in a RAM (1646) and changes to the data structure according to the process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of an implementation in logic hardwired or other circuitry (e.g., an accelerator (1644)) operable with or instead of the software to execute a particular process or a particular part of a particular process described herein. References to software include logic and, where appropriate, vice versa. References to computer-readable media can include, where appropriate, a circuit (such as an integrated circuit (IC)) that stores software for execution, a circuit that implements logic for execution, or both. The present disclosure includes any suitable combination of hardware and software. Appendix A: Glossary JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units, PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC:Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit SDR: standard dynamic range HDR: high dynamic range VTM: VVC Test Mode CPMV: Control point motion vector CPMVP: Control point motion vector predictor MVP: Motion Vector Prediction AMVP: Advanced Motion Vector Prediction ATMVP: Advanced Temporal Motion Vector Prediction HMVP: History-based Motion Vector Prediction STMVP: Spatial-temporal Motion Vector Prediction TMVP: Temporal Motion Vector Prediction SbTMVP: subblock-based temporal motion vector prediction GBi: Generalized Bi-prediction HEVC SCC: HEVC screen content coding CPR: Current Picture Referencing AMVR: Adaptive motion vector resolution SPS: sequence parameter set RD: Rate-Distortion
[0209] Although the present disclosure has described several exemplary embodiments, there are alternatives, substitutions, and various equivalents that are encompassed within the scope of the present disclosure. It will be apparent to those skilled in the art that many systems and methods, which are not explicitly shown or described herein, can be devised that implement the principles of the present disclosure and are thus included within the spirit and scope of the present disclosure.
Claims
1. 1. A method for video decoding performed by a decoder, comprising: obtaining prediction information for a first block in a picture from an encoded video bitstream; generating reconstructed samples of the first block for output according to one of bidirectional prediction and unidirectional prediction and according to the prediction information; When it is determined that a motion information candidate should be stored according to the prediction information of the first block, and is stored as a history-based motion vector prediction (HMVP) candidate, Storing the motion information candidates, the motion information candidates including at least first motion information and a weight index: When the first block is encoded according to the one-way prediction, the weight index is a specified value indicating a specified weight; when the first block is coded according to the bidirectional prediction, the weight index is a first value indicating a first weight for performing the bidirectional prediction of the first block, the first value being one of a plurality of values including the specified value; when it is determined that a second block in the picture should be decoded based on the motion information candidate, obtaining prediction information for generating a reconstructed sample of the second block according to the motion information candidate; The method includes:
2. The step of obtaining prediction information for generating the reconstructed samples of the second block comprises: When the motion information candidate is stored as a normal spatial merge candidate and the second block is coded according to the bi-prediction, setting a second weight for performing the bi-directional prediction on the second block according to the first weight parameter stored in the motion information candidate when the first block is spatially adjacent to the second block; setting the second weight for performing the bi-directional prediction on the second block to the predetermined weight when the first block is not spatially adjacent to the second block; The method of claim 1 , comprising:
3. The step of obtaining prediction information for generating the reconstructed samples of the second block comprises: When the motion information candidate is stored as a candidate other than the normal spatial merge candidate or the HMVP candidate, and the second block is coded according to the bi-directional prediction, The method of claim 2 , comprising setting the second weight for performing the bi-directional prediction on the second block to the defined weight.
4. The step of obtaining prediction information for generating the reconstructed samples of the second block comprises: When the first block is in a current coding tree unit (CTU) row different from a CTU row in which the second block is included, the motion information candidates are stored as normal merge candidates or affine merge candidates, and the second block is coded according to the bi-directional prediction, The method of claim 1 , further comprising setting a second weight for performing the bi-directional prediction on the second block to the defined weight.
5. The step of obtaining prediction information for generating the reconstructed samples of the first block comprises: When the first block is outside a current coding tree unit (CTU) including the second block, the motion information candidate is stored as a translational merge candidate or an inherited affine merge candidate, and the second block is coded according to the bi-directional prediction, The method of claim 1 , further comprising setting a second weight for performing the bi-directional prediction on the second block to the defined weight.
6. The method of claim 1 , wherein the first block is coded according to one of the bidirectional prediction and the unidirectional prediction using the picture as a reference picture.
7. The method of claim 1 , wherein the defined weight is 1 / 2.
8. The first block is coded according to the bi-directional prediction, and the first weight w 1 Ha, w 1 = w / F, Another weight w corresponding to the second reference picture in the second list 0 Ha, w 0 = 1 - w 1 is determined in accordance with The method of claim 1 , wherein w and F are integers, w represents the first value indicating the first weight, and F represents a precision factor.
9. The method of claim 8 , wherein the precision factor F is eight.
10. one or more processors; one or more memories storing a computer program; having The computer program causes the one or more processors to carry out a method according to any one of claims 1 to 9. device.
11. A computer program causing a computer to carry out the method according to any one of claims 1 to 9.
12. 1. A method for video encoding performed by an encoder, comprising: obtaining and signaling prediction information for encoding a first block in a picture; encoding the first block according to one of bidirectional prediction and unidirectional prediction and the prediction information; storing the motion information candidate when it is determined that the motion information candidate should be stored according to the prediction information of the first block and is stored as a history-based motion vector prediction (HMVP) candidate, the motion information candidate including at least a first motion information and a weight index: When the first block is encoded according to the one-way prediction, the weight index is a specified value indicating a specified weight; when the first block is coded according to bidirectional prediction, the weight index is a first value indicating a first weight for performing the bidirectional prediction of the first block, the first value being one of a plurality of values including the specified value; when it is determined that a second block in the picture should be coded based on the motion information candidate, obtaining prediction information for coding the second block according to the motion information candidate; The method includes:
13. one or more processors; one or more memories storing a computer program; having The computer program causes the one or more processors to perform the method of claim 12. device.
14. A computer program product causing a computer to carry out the method according to claim 12.
Citation Information
Patent Citations
Improved interpolation of video compressed frames
JP2006513592A
Different weightings for unidirectional and bidirectional prediction in video coding
JP2012533212A
Different weights for uni-directional prediction and bi-directional prediction in video coding
US20110007803A1
Performing residual prediction in video coding
US20140105299A1