Method, apparatus, and computer program for time motion vector prediction based on sub-blocks

The method enhances video decoding by determining motion vector prediction values for sub-blocks based on their encoding modes, even when they are coded in intra modes, thus addressing the limitations of existing sub-block based temporal motion vector prediction modes.

JP7684007B2Active Publication Date: 2025-05-27TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024040066
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-05
Filing Date
2024-03-14
Publication Date
2025-05-27
Estimated Expiration
2039-05-30

AI Technical Summary

Technical Problem

Sub-block based temporal motion vector prediction modes, such as ATMVP and STMVP, cannot handle sub-blocks coded in an intra mode, limiting their applicability in video encoding.

Method used

A method for video decoding that identifies whether a reference picture of a reference block sub-block is the current picture, and if so, determines the encoding mode as an intra-frame mode. For sub-blocks with reference pictures different from the current picture, the method determines the motion vector prediction value based on the encoding mode of the corresponding reference block sub-block.

Benefits of technology

Enables efficient motion vector prediction for sub-blocks, even when they are coded in intra modes, thereby improving the compression efficiency and adaptability of video encoding techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684007000001
    Figure 0007684007000001
  • Figure 0007684007000002
    Figure 0007684007000002
  • Figure 0007684007000003
    Figure 0007684007000003
Patent Text Reader

Abstract

To disclose a technology for sub-block based temporal motion vector prediction.SOLUTION: A method of video decoding includes acquiring a current picture, and identifying, for a current block included in the current picture, a reference block included in a reference picture that is different from the current picture, where the current block is divided into a plurality of sub-blocks (CBSBs), and the reference block has a plurality of sub-blocks (RBSBs). The method includes: determining whether the reference picture for the RBSB is the current picture; and in response to determining that the reference picture for the RBSB is the current picture, determining a coding mode of the RBSB as an intra-frame mode. The method further includes, in response to determining that the reference picture for the RBSB is not the current picture, determining a motion vector predictor for one of the CBSBs based on whether the coding mode of the corresponding RBSB is one of the intra-frame mode and an inter-frame mode.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This disclosure is incorporated herein by reference in its entirety. BASED TEMPORAL MOTION No. 6,399,633, filed on Oct. 13, 2003, and claims priority to "Sub-Block Based Temporal Motion Vector Prediction Method," the entire contents of which are incorporated herein by reference.

[0002] This disclosure describes embodiments that generally relate to video encoding. [Background technology]

[0003] The background description provided in this specification is intended to provide a general background to the present disclosure. In light of the extent of the work described in the background section, the work of the currently signed inventors and aspects not otherwise limited as prior art at the time of filing are not expressly or implicitly admitted as prior art to the present disclosure.

[0004] Since the last decades it has been known to perform video encoding and decoding by inter-picture prediction with motion compensation. An uncompressed digital video comprises a sequence of pictures, each with spatial dimensions, e.g. 1920x1080 luminance samples and associated chrominance samples. The sequence of pictures may have a fixed or variable picture rate (also informally called frame rate), e.g. 60 pictures per second or 60 Hz. Uncompressed video has high bitrate requirements. For example, 1080p60 4:2:0 video (60 A typical video stream of 1080x1920 resolution (1920x1080 luminance samples at 10 Hz frame rate) requires a bandwidth of about 1.5 Gbit / s. One hour of such video would require more than 600 GB of storage space.

[0005] Video encoding and decoding has one objective to reduce redundancy in the video signal input by compression. Compression contributes to reducing the demand for bandwidth or storage space mentioned above, in some circumstances by more than two orders of magnitude. Lossless compression, lossy compression, and combinations thereof can be used. Lossless compression refers to techniques that reconstruct an exact copy of the original signal from the compressed original signal. When lossy compression is used, the reconstructed signal may differ from the original signal, but the distortion between the original and reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, lossy compression is applied to a large extent. The amount of distortion that is tolerated is application dependent, for example, a user of a consumer streaming application will tolerate higher distortion than a user of a television contribution application. The achievable compression ratio reflects that the higher the permitted / tolerable distortion, the higher the compression ratio.

[0006] Motion compensation may be a lossy compression technique, and involves the following technique: blocks of sample data from a previously constructed picture or part thereof (reference picture) are used to predict a newly reconstructed picture or part of a picture after being spatially shifted in a direction indicated by a motion vector (hereafter called MV). In some situations, the reference picture may be the same as the picture currently being reconstructed. The MV may have two dimensions, X and Y, or it may have three dimensions, the third dimension being an indication of the reference picture being used (the latter may indirectly be a temporal dimension).

[0007] In some video compression techniques, the MV to be applied to a region of sample data can be predicted from other MVs, e.g., from those related to other regions of sample data spatially adjacent to the region to be reconstructed, and the MVs are predicted in decoding order. In this way, the amount of data required to encode the MV is significantly reduced, redundancy is eliminated, and compression is increased. MV prediction works effectively because, for example, when encoding a video input signal derived from a camera (called natural video), there is a statistical possibility that regions larger than the region to which a single MV is applicable move in a similar direction and therefore can be predicted in some circumstances by similar motion vectors derived from the MVs of neighboring regions. This allows the MV found for a particular region to be similar or the same as the MV predicted from the surrounding MVs and, after entropy encoding, can be represented by fewer bits than would be used to encode the MV directly. In some circumstances, MV prediction may be an instantiation of a lossless compression of a signal (i.e., an MV) derived from an original signal (i.e., a sample stream). In other situations, the MV prediction itself may be non-lossy, for example due to rounding errors when computing the prediction value from several surrounding MVs.

[0008] H.265 / HEVC (ITU-T Recommendation H.265, “High Efficiency Video Encoding / Decoding (High Various MV prediction mechanisms are described in "H.265: MV Prediction Mechanisms for Video Coding with High Efficiency," published in December 2016. Among the various MV prediction mechanisms provided by H.265, the one described in this application is a technique referred to below as "spatial merging." Summary of the Invention [Problem to be solved by the invention]

[0009] Some forms of inter prediction are performed at the sub-block level, but sub-block based temporal motion vector prediction modes, such as Alternative Temporal Motion Vector Prediction (ATMVP) and Spatial-Temporal Motion Vector Prediction (STMVP), require the corresponding sub-blocks to be coded in an inter mode, but these temporal motion vector prediction modes cannot handle sub-blocks coded in an intra mode, such as the intra block copy mode. [Means for solving the problem]

[0010] An illustrative embodiment of the present disclosure includes a method for video decoding in a decoder, the method including: obtaining a current picture from an encoded video bitstream; the method further including: identifying, for a current block included in the current picture, a reference block included in a reference picture different from the current picture, the current block being divided into a plurality of sub-blocks (CBSBs), the reference block having a plurality of sub-blocks (RBSBs), each of which corresponds to a different CBSB among the plurality of CBSBs; the method further including: determining whether a reference picture of the RBSB is the current picture; and determining an encoding mode of the RBSB as an intra-frame mode in response to determining that the reference picture of the RBSB is the current picture. The method further includes, in response to determining that the reference picture of the RBSB is not the current picture, (i) for one of the CBSBs, determining whether the encoding mode of the corresponding RBSB is an intra-frame mode or an inter-frame mode, and (ii) determining a motion vector prediction value for the one of the CBSBs based on whether the encoding mode of the corresponding RBSB is an intra-frame mode or an inter-frame mode.

[0011] An illustrative embodiment of the present disclosure includes a video decoder for video decoding, the video decoder having a processing circuit configured to obtain a current picture from an encoded video bitstream. The processing circuit is further configured to identify a reference block included in a reference picture different from the current picture for a current block included in the current picture, the current block being divided into a plurality of sub-blocks (CBSBs), and the reference block having a plurality of sub-blocks (RBSBs), each of which corresponds to a different one of the plurality of CBSBs. The processing circuit is further configured to determine whether the reference picture of the RBSBs is the current picture, and in response to determining that the reference picture of the RBSBs is the current picture, determine an encoding mode of the RBSBs as an intra-frame mode. The processing circuit further, in response to determining that the reference picture for the RBSB is not the current picture, (i) for one of the CBSBs, determines whether the encoding mode of the corresponding RBSB is an intraframe mode or an interframe mode, and (ii) determines a motion vector prediction value for the one of the CBSBs based on whether the encoding mode of the corresponding RBSB is an intraframe mode or an interframe mode.

[0012] An exemplary embodiment of the present disclosure includes a non-transitory computer-readable medium having instructions stored thereon, which, when executed by a processor in a video decoder, cause the processor to perform a method. The method includes obtaining a current picture from an encoded video bitstream. The method further includes, for a current block included in the current picture, identifying a reference block included in a reference picture different from the current picture, where the current block is divided into a plurality of sub-blocks (CBSBs), and the reference block has a plurality of sub-blocks (RBSBs), each of which corresponds to a different one of the plurality of CBSBs. The method further includes determining whether the reference picture of the RBSBs is the current picture, and in response to determining that the reference picture of the RBSBs is the current picture, determining an encoding mode of the RBSBs as an intra-frame mode. The method further includes, in response to determining that the reference picture for the RBSB is not the current picture, (i) for one of the CBSBs, determining whether the encoding mode of the corresponding RBSB is one of an intraframe mode or an interframe mode, and (ii) determining a motion vector prediction value for the one of the CBSBs based on whether the encoding mode of the corresponding RBSB is one of the intraframe mode or the interframe mode. [Brief description of the drawings]

[0013] Other features, properties and advantages of the disclosed subject matter will become more apparent from the following detailed description and drawings.

[0014] [Figure 1] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system (100) according to one embodiment. [Diagram 2] FIG. 2 is a schematic diagram of a simplified block diagram of a communication system (200) according to one embodiment. [Diagram 3] FIG. 2 is a schematic diagram of a simplified block diagram of a decoder according to one embodiment. [Figure 4] FIG. 2 is a schematic diagram of a simplified block diagram of an encoder according to one embodiment. [Diagram 5] 4 shows a block diagram of an encoder according to another embodiment. [Figure 6] 4 shows a block diagram of a decoder according to another embodiment. [Figure 7] FIG. 2 is a schematic diagram of intra-frame picture block compensation; [Figure 8] FIG. 2 is a schematic diagram of a current block and surrounding spatial merge candidates for the current block; [Figure 9] 2 is a schematic diagram of sub-blocks of a current block and corresponding sub-blocks of a reference block; FIG. [Figure 10] 1 illustrates an embodiment of a process performed by an encoder or decoder. [Figure 11] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] FIG. 1 illustrates a simplified block diagram of a communication system (100) according to one embodiment of the present disclosure. The communication system (100) includes a plurality of terminal devices that can communicate with each other, for example, via a network (150). For example, the communication system (100) includes a first pair of terminal devices (110) and (120) that are connected to each other via the network (150). In the example of FIG. 1, the first pair of terminal devices (110) and (120) perform unidirectional data transmission. For example, the terminal device (110) transmits video data (e.g., a video picture stream captured by the terminal device (110)) to another terminal device (120) via the network (150). The encoded video data may be transmitted in the form of one or more encoded video bitstreams. The terminal device (120) receives the encoded video data from the network (150), recovers the video pictures by decoding the encoded video data, and displays the video pictures based on the recovered video data. One-way data transmission is common in media service applications, for example.

[0016] In another example, the communication system (100) includes a second pair of terminal devices (130), (140) for performing bidirectional transmission of encoded video data, e.g., generated during a video conference. For the bidirectional data transmission, in the example, each terminal device in the terminal devices (130), (140) transmits video data (e.g., a video picture stream captured by the terminal device) to another terminal device in the terminal devices (130), (140) via the network (150). Each terminal device in the terminal devices (130), (140) can further receive the encoded video data transmitted from the other terminal device in the terminal devices (130), (140), decode the encoded video data to recover the video pictures, and display the video pictures on an accessible display device based on the recovered video data.

[0017] In the example of FIG. 1, terminal devices 110, 120, 130, and 140 are shown as a server, a personal computer, and a smartphone, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure may apply to laptop computers, tablets, media players, and / or specialized video conferencing equipment. Network 150 may represent any number of networks, including, for example, wired and / or wireless communication networks, for transmitting encoded video data between terminal devices 110, 120, 130, and 140. Communication network 150 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, unless otherwise specified, the architecture and topology of network 150 is not important to the operation of the present disclosure.

[0018] As an example application of the disclosed subject matter, Figure 2 shows a video encoder and decoder arrangement in a streaming transmission environment. The disclosed subject matter is equally applicable to other applications with video capabilities, such as video conferencing, digital television, and storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0019] The streaming transmission system includes a capture subsystem (213) including a video source (201), such as a digital camera, for constructing an uncompressed video picture stream (202). In the example, the video picture stream (202) includes samples captured by the digital camera. The video picture stream (202), depicted as a thick line to emphasize the amount of data when compared to the encoded video data (204) (or encoded video bitstream), is processed by an electronic device (220) including a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof to realize or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (204) (or encoded video bitstream (204)), depicted as a thin line to emphasize the amount of data when compared to the video picture stream (202), is stored in a streaming server (205) for later use. One or more streaming client subsystems, such as client subsystems (206), (208) in FIG. 2, can access the streaming server (205) to retrieve copies (207), (209) of the encoded video data (204). The client subsystem (206) includes a video decoder (210), for example in an electronic device (230). The video decoder (210) decodes the incoming copy (207) of the encoded video data and constructs an output video picture stream (211) that is displayed on a display (212) (e.g., a screen) or other display device (not shown). In a streaming transmission system, the encoded video data (204), (207), and (209) (e.g., a video bitstream) can be encoded according to a video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In an example, the video encoding standard under development is informally referred to as Versatile Video Coding, or VVC.The topic of disclosure applies in the context of VVC.

[0020] Additionally, the electronics (220) and (230) may include other components (not shown). For example, the electronics (220) may include a video decoder (not shown), and the electronics (230) may include a video encoder (not shown).

[0021] 3 shows a block diagram of a video decoder (310) according to one embodiment of the present disclosure. The video decoder (310) is included in an electronic device (330). The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may replace the video decoder (210) in the example of FIG. 2.

[0022] The receiver (331) can receive one or more encoded video sequences to be decoded by the video decoder (310), and in the same or other embodiments, can receive one encoded video sequence at a time, with the decoding of each encoded video sequence being independent of the other encoded video sequences. The receiver (331) can receive the encoded video sequences from a channel (301), which can be a hardware / software link to a storage device for storing the encoded video data. The receiver (331) can receive the encoded video data together with other data, e.g., encoded audio data and / or auxiliary data streams, which can be forwarded to a respective utilization entity (not shown). The receiver (331) can separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (315) is coupled between the receiver (331) and the entropy decoder / parser (320), hereinafter referred to as the "parser (320)". In some applications, the buffer memory (315) is part of the video decoder (310). In other applications, the buffer memory (315) may be external to the video decoder (310) (not shown). Additionally, in other applications, a buffer memory (not shown) may be provided external to the video decoder (310) to, for example, prevent network jitter, and another buffer memory (315) may be provided internal to the video decoder (310) to, for example, handle broadcast timing. When the receiver (331) receives data from a store-and-forward device or an isochronous network with sufficient bandwidth and controllability, the buffer memory (315) may not be required, or may be small. For example, when used with a best-effort packet network such as the Internet, the buffer memory (315) may be required and may be significantly larger, advantageously adaptively sized, and implemented, at least in part, in an operator system or similar element (not shown) external to the video decoder (310).

[0023] The video decoder (310) has a parser (320) that reconstructs codes (321) based on the encoded video sequence. These categories of codes include information for managing the operation of the video decoder (310) and latent information for controlling a display device, such as a display device (312) (e.g., a screen), which is not an integral part of the electronic device (330) but is coupled to the electronic device (330), as shown in FIG. 3. The control information used by the display device(s) may be in the form of Supplementary Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set segments (not shown). The parser (320) performs parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence is based on a video encoding technique or standard and may be implemented using variable length codes, Huffman codes, or other coding techniques. The parser (320) may extract subgroup parameter sets for at least one of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. Subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (320) may further extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0024] The parser (320) can construct a code (321) by performing an entropy decoding / parsing operation on the video sequence received from the buffer memory (315).

[0025] Depending on the type of encoded video picture or part thereof (e.g., inter-frame and intra-frame pictures, inter-frame and intra-frame blocks) and other factors, the reconstruction of the code (321) involves several different units. Which units are involved, and how, can be controlled by subgroup control information that the parser (320) parses from the encoded video sequence. For the sake of brevity, such subgroup control information streams between the parser (320) and the following units are not described:

[0026] In addition to the functional blocks already mentioned, the video decoder (310) can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, several of these units may closely interact with each other and may be at least partially integrated with each other. However, for the purpose of explaining the subject matter of the disclosure, it is appropriate to conceptually subdivide the video decoder (310) into the following functional units:

[0027] The first unit is a scalar / inverse transform unit (351), which receives quantized transform coefficients as code(s) (321) from the parser (320) and control information including what transform scheme to use, block size, quantization factor, quantization scaling matrix, etc. The scalar / inverse transform unit (351) can output a block containing sample values ​​that are input to an aggregator (355).

[0028] In some situations, the output samples of the scalar / inverse transform unit (351) may belong to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information is provided by an intra picture prediction unit (352). In some situations, the intra picture prediction unit (352) uses surrounding already reconstructed information fetched from a current picture buffer (358) to generate blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some situations, the aggregator (355) adds the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scalar / inverse transform unit (351) based on each sample.

[0029] In other situations, the output samples of the scalar / inverse transform unit (351) may belong to a block that is inter-frame coded and may have been motion compensated. In such a situation, the motion compensated prediction unit (353) may access the reference picture memory (357) to fetch the samples used for prediction. After performing motion compensation on the fetched samples based on the code (321) for that block, these samples are added by the aggregator (355) to the output of the scalar / inverse transform unit (351) (in this case referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (357) from which the motion compensated prediction unit (353) fetches the prediction samples may be controlled by a motion vector, which is provided to the motion compensated prediction unit (353) in the form of a code (321), which may have, for example, X, Y, and reference picture components. Motion compensation may further include interpolation of sample values ​​fetched from a reference picture memory (357) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0030] The output samples of the aggregator (355) can utilize various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the coded video sequence (also called coded video bitstream) and made available to the loop filter unit (356) as codes (321) from the parser (320), but can also be responsive to meta-information obtained during decoding of a coded picture or previous part of the coded video sequence (in decoding order), as well as to previously reconstructed loop filtered sample values.

[0031] The output of the loop filter unit (356) can be a sample stream that can be output to a display device (312), or can be stored in a reference picture memory (357) for use in later inter-frame picture prediction.

[0032] Once fully reconstructed, a coded picture can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and identified as a reference picture (e.g., by the parser (320)), the current picture buffer (358) becomes part of the reference picture memory (357), and a new current picture buffer is reallocated before reconstructing a subsequent coded picture.

[0033] The video decoder (310) may perform decoding operations based on a given video compression technique in a standard, such as ITU-T Recommendation H.265. The encoded video sequence complies with the grammar specified by the video compression technique or standard in use, in the sense that the encoded video sequence complies with both the grammar of the video compression technique or standard and with a configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select some tools from all available tools in the video compression technique or standard as the only tools available in the configuration file. Compliance requires that the complexity of the encoded video sequence is within the limits defined by the level of the video compression technique or standard. In some situations, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some situations, the limits set by the level are further limited via the specification of a Hypothetical Reference Decoder (HRD) and metadata for HRD buffer management signaled in the encoded video sequence.

[0034] In one embodiment, the receiver (331) can receive additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence(s). The additional data can be utilized by the video decoder (310) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0035] 4 shows a block diagram of a video encoder (403) according to one embodiment of the present disclosure. The video encoder (403) is included in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., a transmission circuit). The video encoder (403) may replace the video encoder (203) in the example of FIG. 2.

[0036] The video encoder (403) can receive video samples from a video source (401) (not part of the electronic device (420) in the example of FIG. 4) that can capture video images to be encoded by the video encoder (403). In other examples, the video source (401) is part of the electronic device (420).

[0037] The video source (401) provides a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (403), the digital video sample stream being of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling configuration (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source (401) may be a storage device for storing previously prepared video. In a video conferencing system, the video source (401) may be a camera for capturing local image information as a video sequence. The video data may be provided as a number of individual pictures that convey motion when viewed in sequence. The picture itself is organized as a spatial pixel array, with each pixel containing one or more samples depending on the sampling configuration, color space, etc. in use. The relationship between pixels and samples is readily understood by those skilled in the art. The following description focuses on samples.

[0038] According to one embodiment, the video encoder (403) encodes and compresses pictures of a source video sequence into an encoded video sequence (443) in real time or under other time constraints required by the application. Enforcing an appropriate encoding rate is one function of the controller (450). In some embodiments, the controller (450) controls and is operatively coupled to other functional units described below, which are not shown for the sake of simplicity. Parameters set by the controller (450) may include parameters related to rate control (picture skip, quantizer, lambda value for rate distortion optimization techniques...), picture size, group of pictures (GOP) placement, maximum motion vector search range, etc. The controller (450) may be configured to have other appropriate functions for optimizing the video encoder (403) for a particular system design.

[0039] In some embodiments, the video encoder (403) is configured to operate in an encoding loop. As a very simple description, in one example, the encoding loop includes a source encoder (430) (e.g., responsible for constructing a code, such as a code stream, based on an input picture to be encoded and reference picture(s)) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs the code and constructs sample data in the same way that a (remote) decoder constructs sample data (because in the video compression techniques considered in the disclosed subject matter, the compression between the code and the encoded video bit stream is both lossless). The reconstructed sample stream (sample data) is input to a reference picture memory (434). The decoding of the code stream produces bit-exact results that are independent of the decoder location (local or remote), so that the contents in the reference picture memory (434) are bit-exact between the local encoder and the remote encoder. In other words, the predicted parts that the encoder "sees" as reference picture samples are exactly the same as the sample values ​​that the decoder "sees" when it tries to use the prediction during decoding. The basic principles of reference picture synchrony (and the drift that occurs when synchrony cannot be maintained, e.g. from channel errors) also apply in related fields.

[0040] The operation of the "local" decoder (433) may be similar to that of the "remote" decoder of the video decoder (310), for example, as described in detail in connection with Figure 3. However, with brief reference to Figure 3, the entropy decoding portion of the video decoder (310), including the buffer memory (315) and the parser (320), may not be implemented entirely in the local decoder (433), since the code is available and the encoding / decoding of the code into an encoded video sequence by the entropy encoder (445) and parser (320) may be lossless.

[0041] In this case, any decoder technique other than analysis / entropy decoding present in the decoder necessarily needs to be present in the corresponding encoder in essentially the same functional form. For this reason, the subject of the disclosure focuses on the operation of the decoder. The description of the encoder technique may be simplified since it is the inverse of the decoder technique described in full. Only in certain areas is more detailed explanation required and is provided below.

[0042] In operation, in some examples, the source encoder (430) can perform motion-compensated predictive encoding, which predictively encodes an input picture with reference to one or more previously encoded pictures from a video sequence, designated as “reference pictures.” In this manner, the encoding engine (432) encodes differences between pixel blocks of the input picture and pixel blocks of the reference picture(s) that may be selected as prediction reference(s) for the input picture.

[0043] The local video decoder (433) can decode the encoded video data of the pictures that can be designated as reference pictures based on the codes constructed by the source encoder (430). The operation of the encoding engine (432) is preferably a lossy process. When the encoded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can be a copy of the source video sequence, which generally has some errors. The local video decoder (433) copies the decoding process that the video decoder performs on the reference pictures and stores the reconstructed reference pictures in a reference picture cache (434). In this manner, the video encoder (403) stores a copy of the reconstructed reference picture locally, which has a common content (no transmission errors) with the reference picture of the reconstruction obtained by the remote video decoder.

[0044] The predictor (435) may perform a prediction search for the coding engine (432). That is, for a new picture to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (435) may find suitable prediction references by operating on a pixel block by pixel block basis based on the sample blocks. In some situations, an input picture may have prediction references obtained from multiple reference pictures stored in the reference picture memory (434), as determined based on search results obtained by the predictor (435).

[0045] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0046] The output of all the functional units mentioned above can be entropy coded in the entropy coder (445), which performs lossless compression on the codes generated by the various functional units and converts the codes into an encoded video sequence based on techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0047] The sender (440) buffers the encoded video sequence(s) constructed by the entropy encoder (445) to prepare them for transmission over a communication channel (460), which may be a hardware / software link to a storage device for storing the encoded video data. The sender (440) merges the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0048] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) assigns each encoded picture a number of encoding picture types that can affect the encoding technique applied to the respective picture. For example, pictures are typically assigned one of the following picture types:

[0049] An Intraframe picture (I-picture) may be a picture that is coded and decoded without using any other picture in a sequence as a source of prediction. Some video codecs allow different types of Intraframe pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variants of I-pictures, as well as their respective uses and characteristics.

[0050] A predicted picture (P-picture) may be a picture that is encoded and decoded using intra-frame or inter-frame prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0051] A bidirectionally predicted picture (B-picture) may be a picture that is encoded and decoded using intra-frame or inter-frame prediction, with at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, a multi-predicted picture may use more than two reference pictures and associated metadata to reconstruct a single block.

[0052] A source picture can generally be spatially subdivided into sample blocks (e.g., blocks of 4x4, 8x8, 4x8 or 16x16 samples) and coded block by block. Blocks can be predictively coded with reference to other (coded) blocks determined by the coding assignment applied to their respective pictures. For example, blocks of an I-picture can be non-predictively coded or predictively coded (spatial or intraframe) with reference to already coded blocks of the same picture. Pixel blocks of a P-picture can be predictively coded via spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture can be predictively coded via spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0053] The video encoder (403) may perform encoding operations based on a given video encoding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (403) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a grammar specified by the video encoding technique or standard used.

[0054] In one embodiment, the transmitter (440) can transmit additional data along with the encoded video. The source encoder (430) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Video Usability Information (VUI) parameter set segments, etc.

[0055] The captured video may be multiple source pictures (video pictures) that represent a time sequence. Intraframe picture prediction (often abbreviated as intraframe prediction) exploits spatial correlation in a particular picture, while interframe picture prediction exploits correlation (temporal or other) between pictures. In the example, a particular picture in the encoding / decoding, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, that block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in the reference picture, and may have a third dimension to identify the reference picture when multiple reference pictures are used.

[0056] In some embodiments, a bidirectional prediction technique is used for inter-frame picture prediction. Based on the bidirectional prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are utilized, both of which are before the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block can be predicted by a combination of the first reference block and the second reference block.

[0057] Also, merged mode techniques can be applied to inter-frame picture prediction to improve coding efficiency.

[0058] According to some embodiments of the present disclosure, prediction such as inter-frame picture prediction and intra-frame picture prediction is performed for each block. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), including one luminance CTB and two chrominance CTBs. Each CTU is recursively quad-tree partitioned into one or more coding units (CUs). For example, a CTU of 64×64 pixels is partitioned into one CU of 64×64 pixels, or four CUs of 32×32 pixels, or 16 CUs of 16×16 pixels. In the example, each CU is analyzed to determine a prediction type for the CU, such as an inter-frame prediction type or an intra-frame prediction type. Depending on the temporal and / or spatial predictability, a CU is divided into one or more prediction units (PUs). Generally, each PU includes a luma prediction block (PB) and two chrominance PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed on a prediction block basis. As an example of a prediction block, a luma prediction block is used, which includes a matrix of pixel values ​​(e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0059] 5 shows a diagram of a video decoder (503) according to another embodiment of the present disclosure. The video encoder (503) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a video picture sequence, and to encode the processed block into an encoded picture as part of an encoded video sequence. In the example, the video encoder (503) is used in place of the video encoder (203) in the example of FIG. 2.

[0060] In an HEVC example, a video encoder (503) receives a processing block, e.g., a matrix of sample values, such as a predictive block of 8x8 samples. The video encoder (503) determines, e.g., by rate-distortion optimization, whether the processing block is best coded using intra mode, inter mode, or bidirectional prediction mode. When coding the processing block in intra mode, the video encoder (503) can code the processing block into a coded picture using intra prediction techniques, and when coding the processing block in inter mode or bidirectional prediction mode, the video encoder (503) can code the processing block into a coded picture using inter prediction or bidirectional prediction techniques, respectively. In some video encoding techniques, the merged mode can be an inter picture prediction submode in which motion vectors are derived from one or more motion vector predictors without the benefit of coded motion vector components outside the predictors. In some other video encoding techniques, there can be motion vector components applied to the theme block. In the example, the video encoder (503) includes other components, such as a mode decision module (not shown) for determining the mode of the processing blocks.

[0061] In the example of FIG. 5, the video encoder (503) includes an interframe encoder (530), an intraframe encoder (522), a residual calculator (523), a switch (526), ​​a residual encoder (524), a general controller (521), and an entropy encoder (525), concatenated as shown in FIG. 5.

[0062] The inter-frame encoder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block to one or more reference blocks in a reference picture (e.g., blocks in previous and subsequent pictures), generate inter-frame prediction information (e.g., a description of redundant information, motion vectors, merge mode information based on an inter-frame coding technique), and calculate an inter-frame prediction result (e.g., a block of predictions) based on the inter-frame prediction information using any suitable technique.

[0063] The intraframe encoder (522) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block with previously coded blocks in the same picture, generate transformed and quantized coefficients, and optionally further generate intraframe prediction information (e.g., intraframe prediction direction information based on one or more intraframe coding techniques).

[0064] The generic controller (521) is configured to determine generic control data and control other components of the video encoder (503) based on the generic control data. In the illustrated example, the generic controller (521) determines the mode of the block and provides a control signal to the switch (526) based on the mode. For example, if the mode is an intra mode, the generic controller (521) controls the switch (526) to select an intra mode result for use by the residual calculator (523) and controls the entropy encoder (525) to select intra prediction information and include the intra prediction information in the bitstream. If the mode is an inter mode, the generic controller (521) controls the switch (526) to select an inter prediction result for use by the residual calculator (523) and controls the entropy encoder (525) to select inter prediction information and include the inter prediction information in the bitstream.

[0065] The residual calculator (523) is configured to calculate a difference (residual data) between a received block and a prediction result selected from the intraframe encoder (522) or the interframe encoder (530). The residual encoder (524) is configured to operate on the residual data to encode the residual data and generate transform coefficients. In an example, the residual encoder (524) is configured to transform the residual data in the frequency domain to generate transform coefficients. The transform coefficients are then quantized to obtain quantized transform coefficients.

[0066] The entropy encoder (525) is configured to format a bitstream to include the encoded block. The entropy encoder (525) is configured to include various information in accordance with a suitable standard, such as the HEVC standard. In an example embodiment, the entropy encoder (525) is configured to include general control data, selected prediction information (e.g., intraframe prediction information or interframe prediction information), residual information, and other suitable information in the bitstream. Note that, in accordance with the disclosed subject matter, the residual information is not present when encoding a block in an interframe mode or a merged sub-mode of the bi-predictive mode.

[0067] 6 shows a diagram of a video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive encoded pictures as part of an encoded video sequence and to decode the encoded pictures to generate reconstructed pictures. In the example, the video decoder (610) is used in place of the video decoder (210) in the example of FIG. 2.

[0068] In the example of FIG. 6, the video decoder (610) includes an entropy decoder (671), an interframe decoder (680), a residual decoder (673), a reconstruction module (674), and an intraframe decoder (672), concatenated as shown in FIG. 6.

[0069] The entropy decoder (671) is configured to reconstruct, based on the coded picture, specific codes indicative of the grammar elements that make up the coded picture. Such codes include, for example, prediction information (e.g., intraframe or interframe prediction information) that can identify the mode for coding the block (e.g., intraframe mode, interframe mode, bidirectional prediction mode, the latter combined submode or other submodes), specific samples or metadata used for prediction by the intraframe decoder (672) or the interframe decoder (680), respectively, residual information in the form of, for example, transform coefficients for quantization, etc. In the example, if the prediction mode is an interframe or bidirectional prediction mode, the interframe prediction information is provided to the interframe decoder (680), and if the prediction type is an intraframe prediction type, the intraframe prediction information is provided to the intraframe decoder (672). The residual information is provided to the residual decoder (673) via inverse quantization.

[0070] The inter-frame decoder (680) is configured to receive the inter-frame prediction information and to generate an inter-frame prediction result based on the inter-frame prediction information.

[0071] The intraframe decoder (672) is configured to receive the intraframe prediction information and to generate a prediction result based on the intraframe prediction information.

[0072] The residual decoder (673) is configured to perform inverse quantization to extract inverse quantized transform coefficients and to transform the residual from the frequency domain to the spatial domain by processing the inverse quantized transform coefficients. The residual decoder (673) may require some control information (to include quantizer parameters (QP)), which is provided by the entropy decoder (671) (data path not shown since this is a small amount of control information).

[0073] The reconstruction module (674) is configured to combine the residual output from the residual decoder (673) and a prediction result (which may be output from an inter-frame prediction module or an intra-frame prediction module) in the spatial domain to form a reconstructed block, which may be part of a reconstructed picture, which may be part of a reconstructed video, and to perform other suitable operations, such as a deblocking operation, to improve visual quality.

[0074] It should be noted that the video encoder (203), video encoder (403), video encoder (503), and video decoder (210), video decoder (310), and video decoder (610) may be implemented using any suitable technology. In one embodiment, the video encoder (203), video encoder (403), video encoder (503), and video decoder (210), video decoder (310), and video decoder (610) may be implemented using one or more integrated circuits. In another embodiment, the video decoders (203), (403), (403), and video decoders (210), (310), and (610) are implemented by one or more processors for executing software instructions.

[0075] Compensation based on blocks from different pictures is called motion compensation. Block compensation may also be based on previously constructed regions in the same picture, which is called intra-picture block compensation or intra-frame block copy. For example, a displacement vector for indicating the offset between a current block and a reference block is called a block vector. According to some embodiments, the block vector points to a reference block that has already been reconstructed and is used for reference. Also, from the perspective of parallel processing, reference regions beyond tile / slice boundaries or wavefront trapezoid boundaries may be excluded from the reference of block vectors. Due to these constraints, the block vector may be different from the motion vector (MV) in motion compensation, and in motion compensation, the motion vector may be any value (positive or negative in x or y direction).

[0076] FIG. 7 illustrates an example of intra-frame picture block compensation (e.g., intra-frame block copy mode). In FIG. 7, a current picture 700 has a set of blocks that have been coded / decoded (i.e., gray blocks) and a set of blocks that have not been coded / decoded (i.e., white blocks). A sub-block 702 of one of the blocks that has not been coded / decoded may be associated with a block vector 704 that points to another sub-block 706 that has been previously coded / decoded. Thus, any motion information associated with the sub-block 706 can be used for coding / decoding the sub-block 702.

[0077] According to some embodiments, the coding of the block vectors is explicit. In other embodiments, the coding of the block vectors is implicit. In the explicit mode, the difference between the block vectors and their predicted values ​​is signaled, and in the implicit mode, the block vectors are recovered from the predicted values ​​of the block vectors in a manner similar to the motion vector prediction in the merged mode. In some embodiments, the resolution of the block vectors is limited to integer positions. In other embodiments, the block vectors point to fractional positions.

[0078] According to some embodiments, the reference index signals the intra-frame picture block compensation using block level (i.e., intra-frame block copy mode), the currently decoded picture is regarded as a reference picture, and the reference picture is placed at the end of the reference picture list, which may further be managed in a decoded picture buffer (DPB) together with other temporal reference pictures.

[0079] According to some embodiments, the reference block is flipped horizontally or vertically (e.g., flipped intraframe block copy) before being used in prediction for the current block. In some embodiments, each compensation unit within an M×N coding block is M×1 or 1×N rows (e.g., row-based intraframe block copy).

[0080] According to some embodiments, block-level motion compensation is performed, and the current block is the processing unit for performing motion compensation with the same motion information. Therefore, when the size of a block is specified, all pixels in the block use the same motion information to form its prediction block. Examples of block-level motion compensation include using spatial merge candidates, temporal candidates, and combination of motion vectors from existing merge candidates in bidirectional prediction.

[0081] Referring to Fig. 8, a current block (801) contains samples that have already been found by the encoder / decoder in the motion search process to be predictable from a previous block of the same size but spatially shifted. In some embodiments, instead of encoding the MV directly, the MV can be derived from metadata associated with one or more reference pictures, e.g., from the nearest reference picture (in decoding order), using the MV associated with any of the five surrounding samples, denoted A0, A1, B0, B1, B2 (corresponding to 802-806, respectively). Blocks A0, A1, B0, B1, B2 are called spatial merge candidates.

[0082] According to some embodiments, pixels at different positions (e.g., sub-blocks) within a motion compensation block may have different motion information. The difference between these block-level motion information is not signaled but derived. This type of motion compensation is called sub-block level motion compensation, which allows the motion compensation of a block to be smaller than the block itself. In this regard, each block may have multiple sub-blocks, where each sub-block may contain different motion information.

[0083] An example of sub-block level motion compensation includes sub-block based temporal motion vector prediction, where sub-blocks of a current block have different motion vectors. Another example of sub-block level motion compensation is ATMVP, which is a technique that allows each coding block to fetch multiple sets of motion information from multiple blocks smaller than the current coding block from collocated reference pictures.

[0084] Another example of subblock-level motion compensation includes spatial / temporal blending with subblock adjustment, which adjusts the motion vector of each subblock in the current block based on the motion vectors of the subblock's spatial / temporal neighbors. In this mode, some subblocks may require motion information from corresponding subblocks in a temporal reference picture.

[0085] Another example of sub-block level motion compensation is affine coding motion compensation block, which first derives motion vectors at the four corners of the current block based on the motion vectors of neighboring blocks, and then derives other motion vectors of the current block (e.g., at the sub-block or pixel level) according to the affine model, so that each sub-block can have a different motion vector from its respective sub-block neighbors.

[0086] Another example of subblock level motion compensation is merge candidate refinement using motion vector derivation at the decoder side. In this mode, after obtaining the motion vector predictor(s) of the current block or the subblock of the current block, the given motion vector predictor(s) can be further refined using methods such as template matching or bilateral matching. The refined motion vector is used to perform motion compensation. By performing the same refinement operation on both the encoder side and the decoder side, the decoder does not need additional information on how the refinement is displaced from the original prediction. Also, skip mode can be regarded as a special merge mode, in which in addition to deriving the motion information of the current block from the neighbors of the current block, the prediction residual of the current block is also zero.

[0087] According to some embodiments, in sub-block temporal motion vector prediction, sub-blocks of a current block may have different motion vector prediction values ​​derived from a temporal reference picture. For example, identify a set of motion information including a motion vector of the current block and an associated reference index. Determine the motion information from a first available spatial merging candidate. Use the motion information to determine a reference block in a reference picture for the current block. The reference block is also divided into sub-blocks. In some embodiments, for each current block sub-block (CBSB) in the current picture, there is a corresponding reference block sub-block (RBSB) in the reference picture.

[0088] In some embodiments, for each CBSB, when the corresponding RBSB is coded in an inter-frame mode using a set of motion information, the motion information is transformed (e.g., using a method such as motion vector scaling in temporal motion vector prediction) and used as a prediction value of the motion vector of the CBSB. In the following, a method for handling RBSB coded in an intra-frame mode (e.g., intra-frame block copy mode) is described in more detail.

[0089] According to some embodiments, when using sub-block-based temporal motion vector prediction mode, each CBSB is not allowed to be coded in an intra-frame mode such as an intra-frame block copy mode. This is achieved by regarding an RBSB coded in intra-frame block copy as an intra-frame mode. In particular, regardless of how the intra-frame block copy mode is regarded (e.g., regarded as an inter-frame mode, an intra-frame mode, or a third mode), for a CBSB, if the corresponding RBSB is coded in an intra-frame block copy mode, the RBSB is regarded as an intra-frame mode in sub-block-based temporal motion vector prediction. Thus, in some embodiments, an RBSB coded in an intra-frame block copy mode is treated according to a default setting in sub-block-based temporal motion vector prediction. For example, when coding the corresponding RBSB of a CBSB in an intra-frame block copy mode, a default motion vector such as a zero motion vector is used as a prediction value for the CBSB. In the example, the reference picture used for the CBSB is not the current picture but a temporal reference picture. For example, a temporal reference picture may be a picture shared by all sub-blocks of the current block, the first reference picture in a reference picture list, a co-located picture for TMVP purposes, etc. In another example, for a CBSB, if the corresponding RBSB is coded in inter mode, but the reference picture is the current picture, a default motion vector, e.g., a zero motion vector, is assigned to the CBSB. In this regard, even if the RBSB is coded in inter mode, the RBSB is treated as if it was coded in intra mode, since the current picture and the reference picture are the same.

[0090] FIG. 9 shows an example of performing sub-block-based temporal motion vector prediction. FIG. 9 shows a current picture 900 having nine blocks, including a current block 900A. The current block 900A is divided into four sub-blocks 1-4. The current picture 900 can be associated with a reference picture 902, which has nine previously coded / decoded blocks. Also, as shown in FIG. 9, the current block 900A has a motion vector 904 that points to a reference block 902A. The motion vector 904 can be determined using the motion vectors of one or more neighboring blocks (e.g., spatial merging candidates) of the current block 900A. The reference block 902A is divided into four sub-blocks 1-4. The sub-blocks 1-4 of the reference block 902A correspond to the sub-blocks 1-4 of the current block 900A, respectively. If the reference picture 902 and the current picture 900 are the same, each RBSB in the block 902A is treated as if these blocks were coded in an intraframe mode. In this regard, for example, if subblock 1 of block 902A is coded in interframe mode, but the reference picture 902 and the current picture 900 are the same, then subblock 1 of block 902A is treated as if it were coded in intraframe mode, and a default motion vector is assigned to subblock 1 of block 900A.

[0091] If the reference picture 902 and the current picture 900 are different, the sub-blocks 1-4 of the reference block 902A are used to perform sub-block based temporal motion vector prediction for the sub-blocks 1-4 of the block 900A, respectively. For example, a motion vector prediction value for the sub-block 1 of the current block 900A is determined based on whether the sub-block 1 of the reference block 902A is coded in an inter-frame mode or an intra-frame mode (e.g., an intra-frame block copy mode). If the sub-block 1 of the reference block 902A is coded in an inter-frame mode, the motion vector of the sub-block 1 of the reference block 902A is used to determine the motion vector for the sub-block 1 of the current block 900A. If the sub-block 1 of the reference block 902A is coded in an intra-frame mode, the motion vector of the sub-block 1 of the current block 900A is set to a zero motion vector.

[0092] FIG. 10 shows an example of a process performed in an encoder or decoder, such as intraframe encoder 522 or intraframe decoder 672. The process begins at step S1000, where a current picture is obtained from an encoded video bitstream. For example, see FIG. 9, where a current picture 900 is obtained from an encoded video bitstream. The process proceeds to step S1002, where a reference block from a reference picture is identified for a current block in the current picture. For example, see FIG. 9, where a reference picture 902 is retrieved from a reference picture list associated with a current block 900A. When performing sub-block temporal motion vector prediction for the current block 900A, a motion vector 904 may be used to identify a reference block 902A in a reference picture 902.

[0093] The process proceeds to step S1004, where it is determined whether the reference picture is the same as the current picture. If the reference picture is different from the current picture, the process proceeds to step S1006, where it determines the coding mode of the RBSB corresponding to the CBSB. For example, referring to FIG. 9, the CBSB of the current block 900A is The process proceeds to step S1008, where it is determined whether the coding mode of the RBSB is an inter-frame mode. If the coding mode of the RBSB is an inter-frame mode, the process proceeds to step S1010, where it determines a motion vector prediction value of the CBSB based on the motion vector prediction value of the RBSB. For example, if the coding mode of the RBSB of the reference block 902A is an inter-frame mode, the process proceeds to step S1010, where it determines a motion vector prediction value of the CBSB based on the motion vector prediction value of the RBSB. If the coding mode of 1 is the interframe mode, the RBSB of the reference block 902A For example, the motion vector prediction value of RBSB 1 of the reference block 902A is transformed (e.g., using a method such as motion vector scaling in temporal motion vector prediction) to determine the motion vector prediction value of CBSB 1 of the current block 900A. It is used as a motion vector predictor of 1.

[0094] Returning to step S1008, if the coding mode of RBSB is not the inter-frame mode (e.g., the coding mode of RBSB is the intra-frame mode), the process proceeds to step S1012, where the motion vector prediction value of CBSB is set to the default motion vector. If 1 is coded in intraframe mode, the CBSB of the current block 900A A motion vector predictor value of one is set to a default motion vector, such as the zero motion vector.

[0095] Returning to step S1004, if the reference picture and the current picture are the same, the process proceeds to step S1012, where the motion vector predictor of CBSB is set to the default motion vector. In this regard, if the reference picture and the current picture are the same, the coding mode of RBSB is determined to be intra mode, and the motion vector predictor for the corresponding CBSB is set to the default motion vector. In this regard, even if RBSB is in inter mode, by setting the motion vector predictor of CBSB to the default motion vector, RBSB can be treated as if it was coded in intra mode. Steps S1004-S1012 may be repeated for each sub-block in the current block 900A.

[0096] The techniques may be implemented as computer software by computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 11 illustrates a computer system (1100) for implementing some embodiments of the disclosed subject matter.

[0097] Computer software may be encoded in any suitable machine code or computer language, which may be edited, compiled, linked, or otherwise constructed to produce code containing instructions that are executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or the like, or may be executed by interpretation, microcode execution, or the like.

[0098] The instructions may be executed by various types of computers or components thereof, including personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0099] The components of computer system (1100) shown in Figure 11 are exemplary in nature and not limiting on the scope or functionality of use of the example computer software for implementing the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components in the illustrated example of computer system (1100).

[0100] The computer system (1100) may include several human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users, for example, by tactile input (e.g., keystrokes, slides, data glove movements), audio input (e.g., voice, claps), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface devices may also capture certain media that are not necessarily directly related to the conscious input of a human being, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image capture devices), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0101] The input man-machine interface devices may include one or more of a keyboard (1101), a mouse (1102), a touchpad (1103), a touch panel (1110), a data glove (not shown), a joystick (1105), a microphone (1106), a scanner (1107), and an imaging device (1108) (only one of each listed).

[0102] The computer system (1100) may further include man-machine interface output devices. Such man-machine interface output devices may stimulate one or more of the senses of the human user, for example, through haptic output, sound, light, and smell / taste. Such man-machine interface output devices may include haptic output devices (e.g., haptic feedback via a touch panel (1110), data gloves (not shown), or joystick (1105), although there are also haptic feedback devices that are not used as input devices), audio output devices (e.g., speakers (1109), headphones (not shown)), visual output devices (e.g., a screen (1110), including a CRT screen, LCD screen, plasma screen, OLED screen, each of which may or may not have touch panel input and haptic feedback capabilities, some of which may provide two-dimensional visual output or three or more dimensional output, such as by means of stereoscopic image output, including virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0103] The computer system (1100) may further include human-accessible storage devices and associated media, such as a CD / DVD storage device (1121). These include optical media including ROM / RW (1120), thumb drives (1122), removable hard drives or solid state drives (1123), traditional magnetic media such as magnetic tape and floppy disks (not shown), dedicated ROM / ASIC / PLD based devices such as dongles (not shown), etc.

[0104] Those skilled in the art will appreciate that in conjunction with the presently disclosed subject matter, the term "computer-readable medium" as used does not include transmission media, carrier waves or other transient signals.

[0105] The computer system (1100) may further include an interface for one or more communication networks. The network may be, for example, wireless, wired, or optical. The network may further be local, wide area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial television, vehicular and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter coupled to a general-purpose data port or peripheral bus (1149) (e.g., a USB port of the computer system (1100)), while other networks are typically integrated into the core of the computer system (1100) by coupling to a system bus described below (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system). Through any of these networks, the computer system (1100) can communicate with other entities. Such communications may be one-way and receive only (e.g., broadcast television), one-way and transmit only (e.g., a CANbus to a CANbus device), or two-way (e.g., to another computer system over a local or wide area digital network). Specific protocols and protocol stacks may be utilized for each of these networks and network interfaces described above.

[0106] The man-machine interface devices, human-accessible storage devices and network interfaces may be coupled to a core (1140) of the computer system (1100).

[0107] The core (1140) includes one or more central processing units (CPU) (1141), graphics processing units (GPU) (1142), specialized programmable processing units in the form of field programmable gate arrays (FPGA) (1143), hardware accelerators for certain tasks (1144), etc. These devices are connected via a system bus (1148), along with read only memory (ROM) (1145), random access memory (1146), internal mass storage devices (1147) such as hard disk drives, SSDs, etc. that are not accessible to the internal user. In some computer systems, the system bus (1148) can be expanded by additional CPUs, GPUs, etc., by accessing the system bus (1148) in the form of one or more physical plugs. Peripheral devices are coupled to the core's system bus (1148) directly or through a peripheral bus (1149). Peripheral bus architectures include PCI, USB, etc.

[0108] The CPU (1141), GPU (1142), FPGA (1143) and accelerator (1144) can execute a number of instructions, which, when combined, constitute the computer code referred to above. The computer code is stored in ROM (1145) or RAM (1146). Transient data is stored in RAM (1146), and permanent data may be stored, for example, in an internal mass storage device (1147). A cache memory can be used to quickly store and retrieve any of the memory devices, and the cache memory can be closely associated with one or more of the CPU (1141), GPU (1142), mass storage device (1147), ROM (1145), RAM (1146), etc.

[0109] The computer readable medium bears computer code for performing various computer implemented operations, and may be media and computer code specially designed and constructed for the purposes of this disclosure, or may be of the type known and available to those skilled in the art of computer software.

[0110] By way of example and not limitation, a computer system having the architecture (1100), and in particular the core (1140), can provide functionality by having a processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media may be the media related to mass storage accessible by a user as introduced above, as well as storage of the core (1140) that is non-transitory in nature, such as mass storage (1147) internal to the core or ROM (1145). Software in various embodiments for implementing the present disclosure is stored in such devices and executed by the core (1140). Depending on the particular needs, the computer-readable media may include one or more storage devices or chips. The software may cause the core (1140), particularly the processor (including CPU, GPU, FPGA, etc.) therein to execute a particular process or a particular part of a particular process described herein, define data structures stored in RAM (1146), and modify such data structures based on the process defined by the software. Additionally or alternatively, the computer system may provide functionality embodied in logically hardwired or otherwise implemented circuitry (e.g., accelerator (1144)) that may operate in place of or in conjunction with the software to execute a particular process or a particular part of a particular process described herein. Where appropriate, references to software may include logic, and conversely, references to logic may include software. Where appropriate, references to computer-readable media may include circuitry (e.g., integrated circuits (ICs)) on which software is stored for execution, circuitry embodying logic for execution, or both. The present disclosure includes any appropriate combination of hardware and software.

[0111] While certain illustrative embodiments have been described in this disclosure, there are alterations, substitutions, and various substitute equivalents which fall within the scope of this disclosure. Thus, many systems and methods which embody the principles of this disclosure and are within the spirit and scope of this disclosure will come to mind to those skilled in the art, although not expressly described herein.

[0112] (1) A method for video decoding in a decoder, comprising: obtaining a current picture from an encoded video bitstream; identifying, for a current block included in the current picture, a reference block included in a reference picture different from the current picture; the current block is divided into a plurality of sub-blocks (CBSBs), the reference block has a plurality of sub-blocks (RBSBs), the plurality of sub-blocks (RBSBs) respectively corresponding to different CBSBs of the plurality of CBSBs; determining whether a reference picture of the RBSB is the current picture; In response to determining that the reference picture is the current picture, determine an encoding mode of the RBSB as an intra mode; and in response to determining that the reference picture of the RBSB is not the current picture, (i) for one of the plurality of CBSBs, determine whether the encoding mode of the corresponding RBSB is an intra mode or an inter mode, and (ii) determine a motion vector prediction value for the one CBSB of the plurality of CBSBs based on whether the encoding mode of the corresponding RBSB is the intra mode or the inter mode.

[0113] (2) The method according to (1), further comprising setting the motion vector predictor value determined for one of the CBSBs to a default motion vector in response to determining that the encoding mode of the corresponding RBSB is the intraframe mode.

[0114] (3) The method according to (2), wherein the default motion vector is (i) a zero motion vector or (ii) a deviation between the CBSB and the RBSB.

[0115] (4) A method according to any one of features (1) to (3), wherein in response to determining that the coding mode of the corresponding RBSB is the inter-frame mode, a motion vector predictor value determined for one CBSB is based on a motion vector predictor value associated with the corresponding RBSB.

[0116] (5) The method according to any one of the preceding claims, wherein the determined motion vector predictor is a scaled version of the motion vector predictor associated with the corresponding RBSB.

[0117] (6) The method according to any one of features (1) to (5), wherein the reference block is identified according to a motion vector predictor associated with a block adjacent to the current block.

[0118] (7) The method according to any one of features (1) to (6), wherein the reference picture is a first reference picture from a reference picture sequence related to the current picture.

[0119] (8) A video decoder for video decoding, comprising a processing circuit, the processing circuit for obtaining a current picture from an encoded video bitstream, identifying a reference block included in a reference picture different from the current picture for a current block included in the current picture, the current block being divided into a plurality of sub-blocks (CBSBs), the reference block having a plurality of sub-blocks (RBSBs), the plurality of sub-blocks (RBSBs) respectively corresponding to different CBSBs of the plurality of CBSBs, determining whether a reference picture of the RBSB is the current picture, and determining whether a reference picture of the RBSB is the current picture. In response to determining that a picture is the current picture, the coding mode of the RBSB is determined as an intra mode, and in response to determining that a reference picture of the RBSB is not the current picture, (i) for one of the plurality of CBSBs, determine whether the coding mode of a corresponding RBSB is an intra mode or an inter mode, and (ii) determine a motion vector prediction value for the one of the CBSBs based on whether the coding mode of the corresponding RBSB is the intra mode or the inter mode.

[0120] (9) The video decoder of feature (8), wherein the processing circuit is configured to set a motion vector predictor value for the one CBSB to a default motion vector in response to determining that the encoding mode of the corresponding RBSB is the intraframe mode.

[0121] (10) The video decoder of (9), wherein the default motion vector is: (i) a zero motion vector; or (ii) a deviation between the CBSB and the RBSB.

[0122] (11) A video decoder according to any one of features (8) to (10), wherein the processing circuit is configured to determine a motion vector prediction value for one of the CBSBs based on a motion vector prediction value associated with the corresponding RBSB in response to determining that the encoding mode of the corresponding RBSB is the inter-frame mode.

[0123] (12) The video decoder of (11), wherein the determined motion vector predictor is a scaled version of the motion vector predictor associated with the corresponding RBSB.

[0124] (13) The video decoder according to any one of features (8) to (12), wherein the processing circuit identifies the reference block based on a motion vector prediction value associated with a block adjacent to the current block.

[0125] (14) The video decoder according to any one of features (8) to (13), wherein the reference picture is a first reference picture from a reference picture sequence related to the current picture.

[0126] (15) A computer program having instructions, which when executed by a processor, cause the processor to perform a method, the method including: obtaining a current picture from an encoded video bitstream; identifying, for a current block included in the current picture, a reference block included in a reference picture different from the current picture; the current block being divided into a plurality of sub-blocks (CBSBs), the reference block having a plurality of sub-blocks (RBSBs), each of the plurality of sub-blocks (RBSBs) corresponding to a different CBSB of the plurality of CBSBs; and determining whether a reference picture of the RBSB is the current picture. and in response to determining that a reference picture of the RBSB is the current picture, determining an encoding mode of the RBSB as an intra mode; and in response to determining that the reference picture of the RBSB is not the current picture, (i) for one of the plurality of CBSBs, determining whether an encoding mode of a corresponding RBSB is an intra mode or an inter mode, and (ii) determining a motion vector prediction value for the one CBSB of the plurality of CBSBs based on whether an encoding mode of the corresponding RBSB is the intra mode or the inter mode.

[0127] (16) The computer program according to (15), further comprising setting the motion vector predictor value determined for one CBSB to a default motion vector in response to determining that the encoding mode of the corresponding RBSB is the intraframe mode.

[0128] (17) The computer program product according to (16), wherein the default motion vector is (i) a zero motion vector or (ii) a deviation between the CBSB and the RBSB.

[0129] (18) The computer program according to any one of features (15) to (17), wherein in response to determining that the coding mode of the corresponding RBSB is the inter-frame mode, a motion vector prediction value determined for one CBSB is based on a motion vector prediction value associated with the corresponding RBSB.

[0130] (19) The non-transitory computer-readable medium of (18), wherein the determined motion vector predictor is a scaled version of the motion vector predictor associated with the corresponding RBSB.

[0131] (20) A non-transitory computer-readable medium according to any one of features (15) to (19), wherein the reference block is identified according to a motion vector predictor value associated with a block adjacent to the current block.

[0132] Appendix A: Acronyms MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Replenishment Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: conversion unit PU: Prediction unit CTU: Coding Tree Unit CTB: coding tree block PB: Predicted block HRD: Hypothetical Reference Decoder SNR: Signal to Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: cathode ray tube LCD: Liquid crystal display OLED: Organic Light Emitting Diode CD:Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile Communications LTE: Long Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: Solid State Drive IC: Integrated Circuit CU: Coding Unit

Claims

1. 1. A method for video encoding performed by an encoder, comprising: obtaining a current picture from a source video sequence; identifying a reference block in a reference picture for a current block in the current picture, the reference picture being different from the current picture, the current block being divided into a first plurality of sub-blocks (CBSBs) coded in inter mode, the reference block having a second plurality of sub-blocks (RBSBs), and for each CBSB, a corresponding RBSB is determined by a motion vector with respect to the current block; For each CBSB, determining whether the corresponding RBSB is coded in inter mode; responsive to the corresponding RBSB not being coded in the inter mode, setting a first motion vector predictor of the CBSB to a default motion vector predictor, the default motion vector predictor being a zero motion vector; in response to the corresponding RBSB being coded in the inter mode, setting the first motion vector predictor of the CBSB using a second motion vector predictor of the corresponding RBSB as the first motion vector predictor of the CBSB; performing sub-block based temporal motion vector prediction for each CBSB based on the first motion vector predictor value for each CBSB.

2. 2. The method of claim 1, wherein for each CBSB, in response to the corresponding RBSB being coded in the inter mode, the first motion vector predictor is a scaled version of the second motion vector predictor of the corresponding RBSB.

3. The method of claim 1 or 2, wherein the reference picture is a first reference picture from a reference picture sequence associated with the current picture.

4. 1. A video encoder for video encoding, comprising: A processing circuit is provided. The processing circuitry includes: Get the current picture from the source video sequence; Identifying a reference block included in a reference picture for a current block included in the current picture, the reference picture being different from the current picture, the current block being divided into a first plurality of sub-blocks (CBSBs) coded in inter mode, the reference block having a second plurality of sub-blocks (RBSBs), and for each CBSB, a corresponding RBSB is determined by a motion vector for the current block; For each CBSB, determining whether the corresponding RBSB is coded in inter mode; responsive to the corresponding RBSB not being coded in the inter mode, setting a first motion vector predictor of the CBSB to a default motion vector predictor, the default motion vector predictor being a zero motion vector; responsive to the corresponding RBSB being coded in inter mode, setting the first motion vector predictor of the CBSB using a second motion vector predictor of the corresponding RBSB as the first motion vector predictor of the CBSB; performing sub-block based temporal motion vector prediction for each CBSB based on the first motion vector predictor value for each CBSB; It is configured as follows: Video encoder.

5. 5. The video encoder of claim 4, wherein for each CBSB, in response to the corresponding RBSB being coded in the inter mode, the first motion vector predictor is a scaled version of the second motion vector predictor of the corresponding RBSB.

6. 6. A video encoder as claimed in claim 4 or 5, wherein the reference picture is a first reference picture from a reference picture sequence associated with the current picture.

7. A computer program comprising instructions which, when executed by a processor, cause the processor to perform a method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Sub-prediction unit (pu)-based temporal motion vector prediction in hevc and sub-pu design in 3d-hevc

    JP2016537839A

  • Sub-prediction unit-based advanced temporal motion vector prediction

    JP2018506908A

  • Offset vector identification of temporal motion vector predictor

    US20180084260A1