Method, computing system, and computer program for video coding
The bilateral matching method for adaptive MVD resolution addresses the challenge of efficiently resolving motion vector differences in video coding, resulting in improved compression efficiency and video quality by refining MVDs adaptively.
Patent Information
- Application Number
- JP2024526501
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-03-29
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video coding technologies face challenges in efficiently resolving motion vector differences (MVDs) across varying pixel resolutions, which affects video compression efficiency and quality.
The implementation of a bilateral matching method for adaptive MVD resolution, where the system determines whether a joint adaptive MVD resolution mode is signaled, and then refines the motion vector of a video block based on identified predicted blocks from reference frames, allowing for adaptive pixel resolution.
This approach enhances video coding efficiency by refining MVDs adaptively, leading to improved compression performance and video quality, while also optimizing bandwidth and storage requirements.
Smart Images

Figure 2025516419000001_ABST
Abstract
Description
Technical Field
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 339,869, filed on May 9, 2022, entitled "Bilateral Matching for Adaptive Motion Vector Resolution", which is hereby incorporated by reference in its entirety.
[0002] The disclosed embodiments generally relate to video coding and include, but are not limited to, systems and methods for bilateral matching for adaptive motion vector difference (MVD) resolution.
Background Art
[0003] Digital video is supported by various electronic devices such as, for example, digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Those electronic devices transmit and receive digital video data across communication networks or communicate in other ways to store digital video data in storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before the video data is communicated or stored.
[0004] Multiple video codec standards have been developed. For example, video coding standards include AOMedia Video 1 (AV1), Versatile Video Coding (VVC), Joint Exploration Model (JEM), High Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), and Moving Picture Experts Group (MPEG) coding. Video coding generally utilizes prediction methods (e.g., inter prediction, intra prediction, or the like) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation to video quality.
[0005] HEVC, also known as H.265, is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC published the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC), also known as H.266, is a video compression standard intended as a successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AV1 is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the certified version 1.0.0 with errata of the specification was released. SUMMARY OF THE INVENTION
[0006] This disclosure describes advanced video coding techniques, and more specifically, a bilateral matching method for adaptive MVD resolution.
[0007] According to some embodiments, a video coding method is executed by a computing system. The method includes determining whether a joint adaptive motion vector difference (MVD) resolution mode is signaled based on one or more syntax elements from a video stream, where the joint adaptive MVD resolution mode is an inter prediction mode in which MVDs from a first and a second reference frame are signaled at an adaptive MVD pixel resolution; receiving the signaled MVD of a video block within a current frame from the video stream; in response to determining that the joint adaptive MVD resolution mode is signaled, searching for a first predicted video block within a first reference frame and a second predicted video block within a second reference frame for the video block, where the first predicted video block is a reconstructed / predicted video block in front of or behind the video block, and the second predicted video block is a reconstructed / predicted video block in front of or behind the video block; identifying the first predicted video block and the second predicted video block based on a minimum difference measured by a cost criterion between the first predicted block and the second predicted block; refining the signaled MVD of the video block based on the identified first predicted video block and the identified second predicted video block; refining a motion vector (MV) of the video block based on the refined MVD of the video block; and reconstructing / processing the video block based at least on the refined MV.
[0008] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and / or a decoder component.
[0009] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0010] Thus, methods, devices, and systems for coding video are disclosed. Such methods, devices, and systems can complement or replace conventional methods, devices, and systems for video coding.
[0011] The features and advantages described herein are not necessarily all-inclusive. In particular, in view of the drawings, specification, and claims provided in this disclosure, some additional features and advantages will become apparent to those skilled in the art. Also, note that the words used herein are mainly selected for readability and teaching purposes and are not necessarily selected to delineate or demarcate the subject matter described herein.
Brief Description of the Drawings
[0012] To better understand the present disclosure, a more specific description will be given by referring to the features of various embodiments, and some of those embodiments are shown in the accompanying drawings. However, the accompanying drawings merely illustrate the relevant features of the present disclosure and should not necessarily be regarded as limiting, and other valid features that those skilled in the art will recognize and understand when reading this disclosure may be admitted for the purpose of explanation.
Figure 1
Figure 2A
Figure 2B
Figure 3
Figure 4
Figure 5
[0013] In a general manner, the various features shown in the drawings are not necessarily drawn to scale, and the same reference numerals may be used throughout the specification and drawings to indicate similar features. DETAILED DESCRIPTION OF THE INVENTION
[0014] FIG. 1 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system for use with video-enabled applications such as, for example, video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0015] The source device 102 includes a video source 104 (e.g., a camera component or a media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may be of a larger data volume compared to the encoded video bitstream 108 generated by the encoder component 106. Since the encoded video bitstream 108 is of a smaller data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., is configured to transmit uncompressed video data to the (one or more) networks 110).
[0016] One or more networks 110 represent any number of networks that carry information between source device 102, server system 112, and / or electronic device 120, and include, for example, wireline (wired) and / or wireless communication networks. One or more networks 110 can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0017] One or more networks 110 include server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or deliver video content such as an encoded video stream from source device 102). Server system 112 includes coder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, coder component 114 includes an encoder component and / or a decoder component. In various embodiments, coder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, coder component 114 is configured to decode encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings from encoded video bitstream 108.
[0018] In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to prune the encoded video bitstream 108 to tailor potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, a MANE is provided separately from the server system 112.
[0019] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include media storage). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0020] The source device and / or the plurality of electronic devices 120 may be referred to as "terminal devices" or "user devices". In some embodiments, the source device 102, and / or one or more of the electronic devices 120, are instances of a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0021] In an operating example of communication system 100, source device 102 transmits encoded video bitstream 108 to server system 112. For example, source device 102 may encode a stream of pictures captured by the source device. Server system 112 receives encoded video bitstream 108 and may decode and / or encode encoded video bitstream 108 using coder component 114. For example, server system 112 may apply more optimal encoding for network transmission and / or storage to the video data. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of electronic devices 120. Each electronic device 120 may decode the encoded video data 116 to restore video pictures and optionally display them.
[0022] In some embodiments, the above-described transmission is unidirectional data transmission. Unidirectional data transmission may be used in media serving applications and the like. In some embodiments, the above-described transmission is bidirectional data transmission. Bidirectional data transmission may be used in video conferencing applications and the like. In some embodiments, encoded video bitstream 108 and / or encoded video data 116 are encoded and / or decoded according to any of the video coding / compression standards described herein, such as HEVC, VVC, and / or AV1.
[0023] FIG. 2A is a block diagram showing an example of elements of an encoder component 106 according to some embodiments. The encoder component 106 receives a source video sequence from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB, or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0, or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device storing pre-captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. The pictures themselves can be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can immediately understand the relationship between pixels and samples. The following description focuses on samples.
[0024] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence (643) in real time or under other timing constraints required by the application. Enforcing an appropriate encoding rate is one function of the controller 204. In some embodiments, the controller 204 controls other functional units as described hereinafter and is inductively coupled to those other functional units. Parameters set by the controller 204 may include rate control related parameters (including picture skip, quantizer, and / or lambda value of rate distortion optimization techniques, picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.). Those skilled in the art can immediately identify other functions of the controller 204, because they may relate to the encoder component 106 being optimized for a particular system design.
[0025] In some embodiments, encoder component 106 is configured to operate in a coding loop. In a simplified example, the coding loop includes a source coder 202 (which is responsible for creating symbols such as a symbol stream based on, for example, an input picture to be coded and one or more reference pictures), and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in the same manner as a (remote) decoder (when the compression between the symbols and the coded video bitstream is reversible). The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream yields a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. Thus, the prediction part of the encoder interprets the same sample values as reference picture samples as those that the decoder would interpret when using the prediction during decoding. This principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is known to those skilled in the art.
[0026] The operation of decoder 210 can be assumed to be the same as that of a remote decoder, such as decoder component 122, which will be described in detail later in relation to, for example, FIG. 2B. However, referring briefly to FIG. 2B, since symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder 214 and the parser 254 can be reversible, the entropy decoding part of decoder component 122, which includes buffer memory 252 and parser 254, need not be fully implemented in local decoder 210.
[0027] What can be noticed at this point is that decoder technologies other than syntax analysis / entropy decoding existing in the decoder must necessarily exist in substantially the same functional form in the corresponding encoder as well. For this reason, the matters related to the disclosure focus on decoder operations. The explanation of encoder technologies can be omitted because it is the reverse of the decoder technologies that will be thoroughly explained. Only in certain specific areas is further detailed explanation necessary, which will be provided below.
[0028] As part of its operation, the source coder 202 can perform motion compensation prediction coding, which predicts and codes an input frame by referring to one or more previously coded frames designated as reference frames from a video sequence. Thus, the coding engine 212 codes the difference between a pixel block of the input frame and pixel blocks of the (one or more) reference frames that can be selected as the (one or more) prediction references for the input frame. The controller 204 can manage the coding operation of the source coder 202, including, for example, setting the parameters and subgroup parameters used to encode video data.
[0029] The decoder 210 decodes the coded video data of frames that can be designated as reference frames based on the symbols created by the source coder 202. The operation of the coding engine 212 can advantageously be an irreversible process. When the coded video data is decoded by a video decoder (not shown in FIG. 2A), the reconstructed video sequence can be a replica of the source video sequence with some error. The decoder 210 can replicate the decoding process that can be performed by a remote video decoder on the reference frame and cause the reconstructed reference frame to be stored in the reference picture memory 208. Thus, the encoder component 106 locally stores a copy of the reconstructed reference frame that has the same content as the reconstructed reference frame that would be obtained by a remote video decoder.
[0030] Predictor 206 may perform prediction search for the coding engine 212. That is, for a new frame to be coded, predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) that can serve as an appropriate prediction reference for the new picture or for specific metadata such as reference picture motion vectors and block shapes. Predictor 206 may operate on a per pixel block basis to find an appropriate prediction reference. In some cases, the input picture may have prediction references drawn from a plurality of reference pictures stored in the reference picture memory 208 as determined by the search results obtained by predictor 206.
[0031] The outputs of all of the aforementioned functional units may be subjected to entropy coding in the entropy coder 214. The entropy coder 214 converts the symbols generated by the various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0032] In some embodiments, the output of the entropy coder 214 is coupled to a transmitter. The transmitter may be configured to buffer the (one or more) encoded video sequences generated by the entropy coder 214 and prepare them for transmission over the communication channel 218. The communication channel 218 may be a hardware / software link to a storage device storing the encoded video data. The transmitter may be configured to merge the encoded video data from the source coder 202 with other data to be transmitted, such as, for example, encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source coder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as, for example, redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user ability information (VUI) parameter set fragments, and the like.
[0033] The controller 204 may manage the operation of the encoder component 106. In coding, the controller 204 may assign to each coded picture a specific coded picture type that may affect the coding technique applied to that picture. For example, a picture may be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture). An intra picture may be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow multiple different types of intra pictures, including for example an Independent Decoder Refresh (IDR) picture. Those skilled in the art are aware of those variants of I pictures, as well as their respective uses and characteristics, and thus will not repeat them here. A predicted picture may be coded and decoded using intra prediction or inter prediction, using at most one motion vector and a reference index to predict the sample values of each block. A bi-directionally predicted picture may be coded and decoded using intra prediction or inter prediction, using at most two motion vectors and a reference index to predict the sample values of each block. Similarly, a multi-predicted picture may be able to use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0034] The source picture is generally spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be predictively coded with reference to other (already coded) blocks determined by the coding assignment applied to each of those blocks in the picture. For example, blocks of an I picture can be non-predictively coded, or they can be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively, or via spatial prediction, or via temporal prediction referring to a reference picture coded one ahead. Blocks of a B picture can be coded non-predictively, or via spatial prediction, or via temporal prediction referring to one or two reference pictures coded ahead.
[0035] Video can be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra picture prediction (often abbreviated as intra prediction) uses the spatial correlation within a given picture, and inter picture prediction uses the (temporal or other) correlation between pictures. In one example, a particular picture being coded / decoded, referred to as the current picture, is partitioned into a plurality of blocks. When a block within the current picture is similar to a reference block within a reference picture that has been coded earlier in the video and is still buffered, the block within the current picture can be coded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension identifying the reference picture if multiple reference pictures are being used.
[0036] The encoder component 106 may perform encoding operations according to a predetermined video coding technique or standard, such as any of those described herein. In its operation, the encoder component 106 can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Accordingly, the encoded video data may conform to the syntax defined by the video coding technique or standard being used.
[0037] FIG. 2B is a block diagram showing an example of elements of a decoder component 122 according to some embodiments. The decoder component 122 of FIG. 2B is coupled to a channel 218 and a display 124. In some embodiments, the decoder component 122 includes a transmitter configured to be coupled to a loop filter 256 to send data to the display 124 (e.g., via a wired or wireless connection).
[0038] In some embodiments, decoder component 122 includes a receiver configured to be coupled to channel 218 to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data together with other data, such as, for example, encoded audio data and / or auxiliary data streams. Those other data may be transferred to their respective using entities (not shown). The receiver can separate the encoded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the (one or more) encoded video sequences. The additional data may be used by decoder component 122 for decoding the data and / or for more accurately reconstructing the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.
[0039] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. In some embodiments, decoder component 122 is implemented at least partially in software.
[0040] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to handle network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., which is configured to handle playback timing), a separate buffer memory is provided external to decoder component 122 (e.g., to handle network jitter). When receiving data from a storage / transfer device with sufficient bandwidth and controllability or from an isosynchronous network, buffer memory 252 may not be required or may be made smaller. For use on a best effort packet network such as the Internet, buffer memory 252 may be required and may be made relatively large, advantageously of an adaptable size, and may be implemented at least in part by an operating system or similar element (not shown) external to decoder component 122.
[0041] Parser 254 is configured to reconstruct symbol 270 from the encoded video sequence. The symbol may include, for example, information used to manage the operation of decoder component 122 and / or information for controlling a rendering device such as display 124. The control information for the (one or more) rendering devices may be in the form of, for example, a supplemental enhancement information (SEI) message or a video user ability information (VUI) parameter set fragment (not shown). Parser 254 analyzes (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technology or standard and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context dependence, etc. Parser 254 may extract a set of subgroup parameters regarding at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to a group from the encoded video sequence. The subgroups can include, for example, group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. Parser 254 may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the encoded video sequence information.
[0042] For the reconstruction of symbol 270, multiple different units may be involved depending on the type of the encoded video picture or its part and other factors (e.g., inter-picture and intra-picture, inter-block and intra-block, etc.). Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by parser 254. Such a flow of subgroup control information between parser 254 and the following multiple units is not shown for clarity.
[0043] Beyond the functional blocks described above, the decoder component 122 can conceptually be subdivided into a number of functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of explaining the disclosed subject matter, the conceptual subdivision into the following functional units will be maintained.
[0044] The scaler / inverse transform unit 258 receives, as (one or more) symbols 270 from the parser 254, the quantized transform coefficients and control information (e.g., which transform to use, block size, quantization coefficients, and / or quantization scaling matrix). The scaler / inverse transform unit 258 can output a block containing sample values that can be input to the aggregator 268.
[0045] In some cases, the output samples of the scaler / inverse transform unit 258 are related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 can generate a block of the same size and shape as the block being reconstructed, using the surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 can add, for each sample, the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258.
[0046] In other cases, the output samples of the scaler / inverse transform unit 258 are related to the inter-coded, potentially motion-compensated blocks. In such cases, the motion compensation prediction unit 260 can access the reference picture memory 266 to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols 270 related to the block, these samples can be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or a residual signal) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 fetches the prediction samples can be controlled by the motion vector. The motion vector can be available to the motion compensation prediction unit 260 in the form of symbols 270 that can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of the sample values fetched from the reference picture memory 266 when an exact motion vector of sub-samples is used, and a motion vector prediction mechanism, etc.
[0047] The output samples of the aggregator 268 can be subjected to various loop filtering techniques in the loop filter unit 256. Video compression techniques can include in-loop filter techniques, which are controlled by parameters made available to the loop filter unit 256 as symbols 270 from the parser 254 included in the encoded video bitstream, but can also respond to meta-information obtained during decoding of the preceding part (in decoding order) of the encoded picture or encoded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.
[0048] The output of the loop filter unit 256 can be made into a sample stream that can be output to a rendering device such as the display 124, and can also be stored in the reference picture memory 266 for future use in inter-picture prediction.
[0049] When a particular encoded picture is completely reconstructed, it can be used as a reference picture for future prediction. When the encoded picture is completely reconstructed and that encoded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of the reference picture memory 266 and a new current picture memory can be reallocated before starting the reconstruction of the next encoded picture.
[0050] Decoder component 122 may perform a decoding operation according to a predetermined video compression technique that may be documented in a standard, such as any of the standards described herein. The encoded video sequence may conform to the syntax defined by the video compression technique or standard used, in the sense of faithfully adhering to the syntax of the video compression technique or standard defined in the video compression technique document or standard, and specifically the profile document therein. Also, for compliance with some video compression techniques or standards, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limitations set by the level may, in some cases, be further constrained through the Hypothetical Reference Decoder (HRD) specification and the metadata for HRD buffer management signaled within the encoded video sequence.
[0051] FIG. 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., a CPU, a GPU, and / or a DPU). In some embodiments, the control circuit includes one or more field programmable gate arrays (FPGAs), hardware accelerators, and / or one or more integrated circuits (e.g., application specific integrated circuits).
[0052] (One or more) network interfaces 304 may be configured to interface with one or more communication networks (e.g., wireless, wired, and / or optical networks). The communication network can be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of communication networks include local area networks such as Ethernet (registered trademark), wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE and the like, TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicle and industrial including CANbus, and the like. Such communication can be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANbus for a specific CANbus device), or bidirectional (e.g., for other computer systems using a local or wide area digital network). Such communication can include communication to one or more cloud computing networks.
[0053] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The (one or more) input devices 310 may include one or more of a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, or the like. The (one or more) output devices 308 may include one or more of an audio output device (e.g., a speaker), a visual output device (e.g., a display or a monitor), or the like.
[0054] The memory 314 may include high-speed random access memory (e.g., DRAM, SRAM, DDR RAM, and / or other random access solid state memory devices, etc.) and / or non-volatile memory (e.g., one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid state storage devices, etc.). The memory 314 optionally includes one or more storage devices that are remotely located from the control circuit 302. The memory 314, or alternatively a non-volatile solid state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314, or the non-transitory computer-readable storage medium of the memory 314, stores the following programs, modules, instructions, and data structures, or stores a subset or superset thereof: · An operating system 316 including procedures for handling various basic system services and for performing hardware-dependent tasks; · A network communication module 318 used to connect the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); · A coding module 320 that performs various functions regarding encoding and / or decoding data such as video data. In some embodiments, the coding module 320 is an instance of the coder component 114. The coding module 320 includes, but is not limited to, one or more of the following: · A decoding module 322 that performs various functions regarding decoding encoded data, such as those described above with respect to the decoder component 122; and · An encoding module 340 that performs various functions regarding encoding data, such as those described above with respect to the encoder component 106; and · A picture memory 352 that stores pictures and picture data for use with, for example, the coding module 320. In some embodiments, the picture memory 352 includes one or more of the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.
[0055] In some embodiments, the decoding module 322 includes an analysis module 324 (configured to perform various functions described above with respect to, for example, the parser 254), a conversion module 326 (configured to perform various functions described above with respect to, for example, the scaler / inverse transform unit 258), a prediction module 328 (configured to perform various functions described above with respect to, for example, the motion compensation prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (configured to perform various functions described above with respect to, for example, the loop filter 256).
[0056] In some embodiments, the encoding module 340 includes a code module 342 (configured to perform the various functions previously described with respect to the source coder 202 and / or the coding engine 212), and a prediction module 344 (configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes a subset of the modules shown in FIG. 3. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0057] Each of the modules identified above stored in the memory 314 corresponds to a set of instructions for performing the functions described herein. The modules identified above (e.g., the set of instructions) need not be implemented as separate software programs, procedures, or modules, and thus, in various embodiments, various subsets of these modules may be combined or otherwise reorganized. For example, the coding module 320 optionally uses the same set of modules to perform both sets of functions without separately including a decoding module and an encoding module. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above, such as, for example, an audio processing module.
[0058] In some embodiments, the server system 112 includes a web or Hypertext Transfer Protocol (HTTP) server, a File Transfer Protocol (FTP) server, and web pages and applications implemented using Common Gateway Interface (CGI) scripts, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), Hypertext Markup Language (HTML), Extensible Markup Language (XML), Java (registered trademark), JavaScript (registered trademark), Asynchronous JavaScript + XML (AJAX), XHP, Javelin, Wireless Universal Resource File (WURFL), and the like.
[0059] FIG. 3 shows a server system 112 according to some embodiments, but FIG. 3 is intended as a functional description of various features that may exist in one or more server systems rather than a structural schematic of the embodiments described herein. In practice, as will be recognized by those skilled in the art, separately shown items may be combined and some items may be separated. For example, some of the items separately shown in FIG. 3 may be implemented on a single server, or a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how features are allocated among them will vary from implementation to implementation and, optionally, will depend in part on the amount of data traffic the server system handles during peak and average usage periods.
[0060] In some implementations, a predicted block (PB or coding block (CB), also referred to as PB when not further partitioned into predicted blocks) obtained from any of the partitioning schemes can be an individual block for coding via either intra prediction or inter prediction. In the case of inter prediction for the current PB, a residual between the current block and the predicted block is generated, coded, and can be included in the coded bitstream.
[0061] In some implementations, inter prediction may be performed, for example, in a single-reference mode or a multiple-reference mode. In some implementations, first, a skip flag may be included in the bitstream for the current block (or at a higher level) to indicate whether the current block is to be inter-coded and not skipped. If the current block is inter-coded, another flag may be further included in the bitstream as a signal to indicate whether a single-reference mode or a multiple-reference mode is used for prediction of the current block. In the single-reference mode, one reference block may be used to generate a predicted block for the current block. In the multiple-reference mode, two or more reference blocks may be used to generate a predicted block, for example, by weighted averaging. The multiple-reference mode may sometimes be referred to as a more-than-one-reference mode, a two-reference mode, or a multiple-reference mode. One or more reference blocks may be identified using one or more reference frame indices and further using one or more corresponding motion vectors indicating (one or more) shifts between the (one or more) reference blocks and the current block at positions in, for example, horizontal and vertical pixels. For example, an inter-predicted block for the current block may be generated from one reference block identified by one motion vector in a reference frame as the predicted block in the single-reference mode, while in the multiple-reference mode, the predicted block may be generated by weighted averaging of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. The (one or more) motion vectors may be coded in various ways and included in the bitstream.
[0062] In some implementations, an encoding or decoding system may maintain a decoded picture buffer (DPB). Some images / pictures can be maintained in the DPB while waiting to be displayed (in the decoding system), and some images / pictures in the DPB can be used as reference frames to enable inter prediction (in the decoding or encoding system). In some implementations, the reference frames in the DPB can be tagged as either short-term references or long-term references for the currently encoded or decoded picture. For example, a short-term reference frame may include a frame used for inter prediction for blocks within the current frame, or blocks within a predetermined number (e.g., two) of subsequent video frames that are closest to the current frame in decoding order. A long-term reference frame may include a frame in the DPB that can be used to predict image blocks within a frame that is further away from the current frame than the predetermined number of frames in decoding order. Information about such tags for short-term and long-term reference frames may be referred to as a Reference Picture Set (RPS) and can be added to the header of each frame in the encoded bitstream. Each frame in the encoded video stream can be identified by a Picture Order Counter (POC), which is numbered according to the playback sequence in an absolute manner or, for example, in relation to a picture group starting from an I-frame.
[0063] In some implementation examples, one or more reference picture lists including the identification of short-term reference frames and long-term reference frames for inter prediction may be formed based on the information in the RPS. For example, in uni-directional inter prediction, a single picture reference list denoted as L0 reference (or reference list 0) may be formed, while in bi-directional inter prediction, for each of the two prediction directions, two picture reference lists denoted as L0 (or reference list 0) and L1 (or reference list 1) may be formed. The reference frames included in the L0 list and the L1 list may be ordered in various predetermined ways. The lengths of the L0 list and the L1 list may be signaled in the video bitstream. Uni-directional inter prediction may be either in a single-reference mode or in a composite-reference mode when multiple references for the generation of a predicted block by weighted average in the composite prediction mode are on the same side of the block to be predicted. Bi-directional inter prediction can only be in a composite mode in that at least two reference blocks are involved in bi-directional inter prediction.
[0064] In some implementations, a merge mode (MM) for inter prediction may be implemented. Generally, in the merge mode, instead of the motion vectors in a single reference prediction for the current PB or one or more of the motion vectors in a composite reference prediction being independently calculated and signaled, they may be derived from other (one or more) motion vectors. For example, in an encoding system, the (one or more) current motion vectors for the current PB may be represented by the difference between the current motion vector and one or more other already encoded motion vectors (referred to as reference motion vectors). Instead of encoding the (one or more) current motion vectors in their entirety, such (one or more) differences of the (one or more) motion vectors can be encoded and included in the bitstream, and can be linked to the (one or more) reference motion vectors. Correspondingly, in a decoding system, the motion vector corresponding to the current PB can be derived based on the decoded (one or more) motion vector differences and the decoded (one or more) reference motion vectors linked thereto. As a particular form of a general merge mode (MM) inter prediction, such an inter prediction based on motion vector differences may be referred to as Merge Mode with Motion Vector Difference (MMVD). Thus, generally MM, or particularly MMVD, can be implemented to utilize the correlation between motion vectors associated with different PBs to improve coding efficiency. For example, adjacent PBs may have similar motion vectors, and thus, the MVD can be small and can be efficiently coded. In another example, the motion vectors can be temporally correlated (between frames) for similarly located / positioned blocks in space.
[0065] In some implementation examples, to indicate whether the PB is currently in merge mode, an MM flag can be included in the bitstream during the encoding process. Additionally, or alternatively, to indicate whether the PB is currently in MMVD mode, an MMVD flag can be included in the bitstream and signaled during the encoding process. The MM and / or MMVD flags or indicators can be provided at levels such as the PB level, coding block (CB) level, coding unit (CU) level, coding tree block (CTB) level, coding tree unit (CTU) level, slice level, and picture level. In a specific example, both the MM flag and the MMVD flag can be included for the current CU, and the MMVD flag can be signaled immediately after the skip flag and the MM flag to define whether the MMVD mode is used for the current CU.
[0066] In some implementation examples of MMVD, a list of reference motion vector (RMV) candidates or MV predictor candidates for motion vector prediction can be formed for the block being predicted. The list of RMV candidates can include a predetermined number (e.g., two) of MV predictor candidate blocks having motion vectors that can be used to predict the current motion vector. The RMV candidate blocks can include blocks selected from adjacent blocks and / or temporal blocks (e.g., blocks at the same position in a previous or subsequent frame of the current frame) within the same frame. These options represent blocks at spatial or temporal positions relative to the current block that are likely to have similar or the same motion vector as the current block. The size of the list of MV predictor candidates can be predetermined. For example, the list can include two or more candidates. To be included in the list of RMV candidates, a candidate block may be required to have, for example, the same (one or more) reference frames as the current block, must exist (e.g., boundary checking may be required when the current block is near the edge of the frame), and must not have already been encoded during the encoding process and / or decoded during the decoding process. In some implementations, the list of merge candidates can first be populated with spatial adjacent blocks (scanned in a particular predetermined order) if available and meeting the above conditions, and then temporal blocks can be added if there is still room in the list. Adjacent RMV candidate blocks can be selected, for example, from the left and upper blocks of the current block. The list of RMV predictor candidates can be dynamically formed at various levels (sequence, picture, frame, slice, superblock, etc.) as a Dynamic Reference List (DRL). The DRL can be signaled within the bitstream.
[0067] In some implementations, the actual MV predictor candidates that are currently being used as reference motion vectors for predicting the motion vector of a block can be signaled. If the RMV candidate list contains two candidates, a 1-bit flag called the merge candidate flag can be used to indicate the selection of the reference merge candidate. For a current block being predicted in a composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate predicts the current coding block more closely and signal that selection as an index to the DRL.
[0068] In some implementation examples of MMVD, after an RMV candidate is selected and used as the base motion vector predictor (MVP) for the motion vector to be predicted, the motion vector difference (MVD or delta MV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the coding system. Such an MVD can include information representing the magnitude and direction of the MV difference, and both of them can be signaled in the bitstream. The magnitude of the motion difference and the direction of the motion difference can be signaled in various ways.
[0069] In some implementation examples of MMVD, a distance index can be used to define the magnitude information of the motion vector difference and indicate one of a set of predetermined offsets representing a predetermined motion vector difference from the starting point (reference motion vector). Then, the MV offset according to the signaled index can be added to either the horizontal component or the vertical component of the starting (reference) motion vector. Whether the horizontal component or the vertical component of the reference motion vector should be offset can be determined by the direction information of the MVD. An example of the predetermined relationship between the distance index and the predetermined offset is defined in Table 1. [Table 1]
[0070] In some implementation examples of MMVD, the direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be restricted to either the horizontal or vertical direction. An example of a 2-bit direction index is shown in Table 2. In the example of Table 2, the interpretation of the MVD may vary according to the information of the start / reference MV. For example, when the start / reference MV corresponds to a single prediction block, or when both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are both greater than or both less than the POC of the current picture), the signs in Table 2 may define the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bi-prediction block where the two reference pictures are on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 2 can define the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign for the offset of the MV corresponding to the reference picture in picture reference list 1 can have the opposite value (opposite sign to the offset). Otherwise, when the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, the signs in Table 2 can define the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign for the offset of the reference MV associated with picture reference list 0 has the opposite value.
Table 2
[0071] In some implementation examples, the MVD can be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Instead, if the difference in POC in reference list 0 is greater than that in reference list 1, the MVD for reference list 1 is scaled. If the POC difference in reference list 1 is greater than that in list 0, the MVD for list 0 can be scaled in a similar manner. When the starting MV is a uni-prediction, the MVD is added to the available or reference MV.
[0072] In some implementation examples of MVD coding and signaling in bidirectional composite prediction, in addition to or instead of coding and signaling two MVDs separately, symmetric MVD coding can be implemented such that only one of the MVDs requires signaling and the other MVD can be derived from the signaled MVD. In such an implementation, motion information including reference picture indices for both list 0 and list 1 is signaled. However, for example, only the MVD associated with reference list 0 is signaled and the MVD associated with reference list 1 is derived without being signaled. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" can be included in the bitstream to indicate whether reference list 1 is not signaled within the bitstream. If this flag is 1, it indicates that reference list 1 is equal to 0 (thus not signaled), and a bidirectional prediction flag called "BiDirPredFlag" can be set to 0, meaning there is no bidirectional prediction. Otherwise, when mvd_l1_zero_flag is 0 and the closest reference picture in list 0 and the closest reference picture in list 1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, BiDirPredFlag can be set to 1 and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 can indicate that a symmetric mode flag is further signaled within the bitstream. When BiDirPredFlag is 1, the decoder can extract the symmetric mode flag from the bitstream. The symmetric mode flag can be signaled, for example, at the CU level (if necessary), and can indicate whether the symmetric MVD coding mode is used for the corresponding CU.When the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, and only the reference picture indices of both list 0 and list 1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") are signaled together with the MVD related to list 0 (referred to as "MVD0"), and the other motion vector difference "MVD1" should not be signaled but should be derived. For example, MVD1 can be derived as -MVD0. Therefore, in the example of the symmetric MVD mode, only one MVD is signaled. In some other implementation examples for MV prediction, a harmonized way can be used to implement a general merge mode, MMVD, and some other types of MV prediction for both single-reference mode MV prediction and composite-reference mode MV prediction. Various syntax elements can be used to signal the way the MV for the current block is predicted.
[0073] For example, in the single-reference mode, the following MV prediction modes can be signaled.
[0074] NEARMV - Use one of the motion vector predictors (MVP) in the list indicated by the DRL (dynamic reference list) index directly without MVD.
[0075] NEWMV - Use one of the motion vector predictors (MVP) in the list signaled by the DRL index as a reference and apply a delta to the MVP (e.g., using MVD).
[0076] GLOBALMV - Use a motion vector based on frame-level global motion parameters.
[0077] Similarly, in the composite-reference inter-prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes can be signaled.
[0078] NEAR_NEARMV - For each of the two MVs to be predicted, without MVD, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index.
[0079] NEAR_NEWMV - To predict the first of the two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV without MVD, and to predict the second of the two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV together with the additionally signaled delta MV (MVD).
[0080] NEW_NEARMV - To predict the second of the two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV without MVD, and to predict the first of the two motion vectors, use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV together with the additionally signaled delta MV (MVD).
[0081] NEW_NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as the reference MV and use it together with the additionally signaled delta MV to predict each of the two MVs.
[0082] GLOBAL_GLOBALMV - Use the MV from each reference based on the frame-level global motion parameters.
[0083] Thus, the term "NEAR" refers to MV prediction that uses a reference MV without MVD as a general merge mode, while the term "NEW" refers to MV prediction that uses a reference MV and offsets it with the signaled MVD, as in the MMVD mode. In composite inter prediction, although they are correlated and such correlation can be utilized to reduce the amount of information required to signal two motion vector deltas, both the above reference-based motion vectors and motion vector deltas can be generally different or independent between two references. In such a situation, joint signaling of two MVDs can be implemented and shown in the bitstream.
[0084] The above dynamic reference list (DRL) can be used to dynamically maintain and hold a set of indexed motion vectors considered as candidate motion vector predictors.
[0085] In some implementations, a predetermined resolution can be enabled for the MVD. For example, 1 / 8 pixel motion vector accuracy (or precision) can be enabled. The above MVDs in various MV prediction modes can be configured and signaled in various ways. In some implementations, various syntax elements can be used to signal the above (one or more) motion vector differences in reference frame list 0 or list 1.
[0086] For example, a syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. For the MVD, this is signaled together for all non-zero components. For example: · A mv_joint with a value of 0 can indicate that there is no non-zero MVD along either the horizontal or vertical direction; · A mv_joint with a value of 1 can indicate that there is a non-zero MVD only along the horizontal direction; · A mv_joint with a value of 2 can indicate that there is a non-zero MVD only along the vertical direction; An mv_joint having a value of ·3 may indicate that there are non-zero MVDs along both the horizontal and vertical directions.
[0087] When the "mv_joint" syntax element for MVD signals that there are no non-zero MVD components, no further MVD information may be signaled. However, when the "mv_joint" syntax signals the presence of one or two non-zero components, additional syntax elements may be further signaled for each of the non-zero MVD components, as described below.
[0088] For example, a syntax element called "mv_sign" may be used to further specify whether the corresponding motion vector difference component is positive or negative.
[0089] In another example, a syntax element called "mv_class" may be used to specify the class of the motion vector difference for the corresponding non-zero MVD component within a predetermined set of classes. The predetermined classes for the motion vector difference may be used, for example, to divide the continuous magnitude space of the motion vector difference into a plurality of non-overlapping ranges, each range corresponding to an MVD class. Thus, the signaled MVD class indicates the magnitude range of the corresponding MVD component. In the implementation example shown in Table 3 below, higher classes correspond to motion vector differences with larger magnitude ranges. In Table 3, the symbol (n,m] is used to represent the range of motion vector differences that are greater than n pixels and less than or equal to m pixels.
Table 3
[0090] In some other examples, a syntax element called "mv_bit" may be further used to define the integer part of the offset between a non-zero motion vector difference component and the starting size of the correspondingly signaled MV class size range. The number of bits required for "mv_bit" to signal the entire range of each MVD class may vary as a function of the MV class. In this example, MV_CLASS_0 and MV_CLASS_1 in the implementation of Table 3 can be assumed to require only a single bit to indicate an integer pixel offset of 1 or 2 from a starting MVD of 0, and the higher MV_CLASSES in the implementation example of Table 3 may each require one more bit for "mv_bit" than the previous MV_CLASS progressively.
[0091] In some other examples, a syntax element called "mv_fr" may be further used to define the first two fractional bits of the motion vector difference for the corresponding non-zero MVD component, while a syntax element called "mv_hp" may be used to define the third fractional bit (high-precision bit) of the motion vector difference for the corresponding non-zero MVD component. The 2-bit "mv_fr" basically provides a 1 / 4 pixel MVD resolution, while the "mv_hp" bit may further provide a 1 / 8 pixel resolution. In some other implementations, more than two "mv_hp" bits may be used to provide an MVD pixel resolution finer than 1 / 8 pixel. In some implementation examples, one or more additional flags may be signaled at one or more of various levels to indicate whether an MVD resolution of 1 / 8 pixel or higher is supported. If the MVD resolution is not applicable to a particular coding unit, the above-mentioned syntax elements for the corresponding unsupported MVD resolution may not be signaled.
[0092] In some of the above implementation examples, the fractional resolution may be independent of the various classes of MVDs. In other words, regardless of the magnitude of the motion vector difference, similar options for the motion vector resolution may be provided using a predetermined number of "mv_fr" bits and "mv_hp" bits for signaling the fractional MVDs of the non-zero MVD components.
[0093] However, in some other implementation examples, the resolution for motion vector differences in various MVD size classes can be made different. Specifically, the high-resolution MVD for larger MVD sizes in higher MVD classes may not provide a statistically significant improvement in compression efficiency. Therefore, for larger MVD size ranges corresponding to higher MVD size classes, the MVD can be coded at a reduced resolution (integer pixel resolution or fractional pixel resolution). Similarly, generally, for larger MVD values, the MVD can be coded at a reduced resolution (integer pixel resolution or fractional pixel resolution). Such MVD class-dependent or MVD size-dependent MVD resolution can generally be referred to as adaptive MVD resolution, amplitude-dependent adaptive MVD resolution, or size-dependent MVD resolution. The term "resolution" may further be referred to as "pixel resolution". Adaptive MVD resolution can be implemented in various ways as described by the following implementation examples to achieve overall better compression efficiency. In particular, treating the MVD resolution for larger sizes or higher classes of MVD at a level similar to that for smaller sizes or lower classes of MVD in a non-adaptive manner may not significantly enhance the inter-prediction residual coding efficiency for blocks having larger sizes or higher classes of MVD. According to the statistical observation that this may not be the case, reducing the number of signaling bits by aiming for a less accurate MVD can be greater than the additional bits required to code the inter-prediction residual as a result of such a less accurate MVD. In other words, it can be said that using a higher MVD resolution for larger sizes or higher classes of MVD does not produce many coding gains compared to using a lower MVD resolution.
[0094] In some general implementation examples, the pixel resolution or accuracy for MVD may decrease or not increase as the number of MVD classes increases. Decreasing the pixel resolution for MVD corresponds to a coarser MVD (or a larger step from one MVD level to the next MVD level). In some implementations, the correspondence between MVD pixel resolution and MVD class can be defined, predetermined, or preconfigured, and thus may not need to be signaled within the encoded bitstream.
[0095] In some implementation examples, each of the MV classes in Table 3 may be associated with a different MVD pixel resolution.
[0096] In some implementation examples, each MVD class may be associated with a single allowable resolution. In some other implementations, one or more MVD classes may be associated with two or more alternative MVD pixel resolutions. Thus, the signal within the bitstream for the current MVD component having such an MVD class may be followed by additional signaling to indicate which alternative pixel resolution is selected for the current MVD component.
[0097] In some implementation examples, the adaptively allowed MVD pixel resolutions may include, but are not limited to, 1 / 64 pel (pixel), 1 / 32 pel, 1 / 16 pel, 1 / 8 pel, 1 / 4 pel, 1 / 2 pel, 1 pel, 2 pel, 4 pel, … (in descending order of resolution). Thus, each of the MVD classes in ascending order can be associated with one of these resolutions in a non-ascending manner. In some implementations, an MVD class may be associated with two or more of the above resolutions, and the higher resolution may be less than or equal to the lower resolution for the previous MVD class. For example, when MV_CLASS_3 in Table 3 can be associated with the selectable 1-pel and 2-pel resolutions, the highest resolution that MV_CLASS_4 in Table 3 can be associated with is 2 pels. In some other implementations, the highest allowable resolution for a certain MV class may be higher than the lowest allowable resolution of the previous (lower) MV class. However, the average of the allowable resolutions for the MV classes in ascending order can only be non-increasing.
[0098] In some implementations, when a fractional pixel resolution higher than 1 / 8 pel is allowed, correspondingly, the “mv_fr” and “mv_hp” signaling can be extended to a total of more than 3 fractional bits.
[0099] In some implementation examples, fractional pixel resolution may be allowed only for MVD classes below a threshold MVD class. For example, fractional pixel resolution may be allowed only for MVD-CLASS_0 and not allowed for all other MV classes in Table 3. Similarly, fractional pixel resolution may be allowed only for MVD classes below any one of the other MV classes in Table 3. For other MVD classes above the threshold MVD class, only integer pixel resolution for the MVD is allowed. In such a manner, fractional resolution signaling, such as one or more of the "mv-fr" and / or "mv-hp" bits, may not need to be signaled for MVDs signaled in MVD classes above the threshold MVD class. For MVD classes with a resolution lower than 1 pixel, the number of bits in the "mv-bit" signaling can be further reduced. For example, in MV_CLASS_5 of Table 3, the range of the MVD pixel offset is (32,64], and thus 5 bits are required to signal the entire range at 1 pel resolution. However, if MV_CLASS_5 is associated with a 2 pel MVD resolution (a resolution lower than 1 pel resolution), 4 bits instead of 5 bits are required for "mv-bit", and neither "mv_fr" nor "mv-hp" needs to be signaled following the signaling of "mv_class" as MV_CLASS_5.
[0100] In some implementation examples, fractional pixel resolution may be allowed only for MVDs having integer values below a threshold integer pixel value. For example, fractional pixel resolution may be allowed only for MVDs smaller than 5 pixels. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 in Table 3 and not allowed for all other MV classes. In another example, fractional pixel resolution may be allowed only for MVDs smaller than 7 pixels. Corresponding to this example, fractional resolution may be allowed for MV_CLASS_0 and MV_CLASS_1 (having a range below 5 pixels) in Table 3 and not allowed for MV_CLASS_3 and later (having a range above 5 pixels). For MVDs belonging to MV_CLASS_2 whose pixel range includes 5 pixels, fractional pixel resolution for the MVD may be allowed or not allowed depending on the “mv bit” value. When the “m-bit” value is signaled as 1 or 2 (such that, calculating from the start of the pixel range of MV_CLASS_2 having offset 1 or 2 indicated by “m-bit”, the integer part of the signaled MVD is 5 or 6), fractional pixel resolution may be allowed. Otherwise, when the “mv bit” value is signaled as 3 or 4 (such that the integer part of the signaled MVD is 7 or 8), fractional pixel resolution may not be allowed.
[0101] In some other implementations, for MV classes above a threshold MV class, only a single MVD value may be allowed. For example, such a threshold MV class may be MV_CLASS_2. Thus, after MV_CLASS_2, only having a single MVD value is allowed, and it does not have fractional pixel resolution. The single allowable MVD value for these MV classes may be predetermined. In some examples, the single allowed value may be the upper end value of each range for these MV classes in Table 3. For example, MV_CLASS_2 to MV_CLASS_10 can be considered above the threshold class of MV_CLASS_2, and the single allowable MVD value for these classes can be predetermined as 8, 16, 32, 64, 128, 256, 512, 1024, and 2048 respectively. In some other examples, the single allowed value may be the middle value of each range for these MV classes in Table 3. For example, MV_CLASS_2 to MV_CLASS_10 can be considered to exceed the class threshold, and the single allowable MVD value for these classes can be predetermined as 3, 6, 12, 24, 48, 96, 192, 384, 768, and 1536 respectively. Any other value within the range can also be defined as the single allowable resolution for each MVD class.
[0102] In the above implementation, when the "mv_class" signaled is above a predetermined MVD class threshold, only the "mv_class" signaling is sufficient to determine the MVD value. Then, the magnitude and direction of the MVD are determined using "mv_class" and "mv_sign".
[0103] Therefore, when the MVD is signaled for only one reference frame (but not both, from either reference frame list 0 or list 1) or together for two reference frames, the accuracy (or resolution) of the MVD may depend on the class of the motion vector difference in the associated Table 3 and / or the magnitude of the MVD.
[0104] In some other implementations, the pixel resolution or accuracy of the MVD may decrease or not increase as the size of the MVD increases. For example, the pixel resolution may depend on the integer part of the size of the MVD. In some implementations, fractional pixel resolution may only be allowed for MVD sizes below the amplitude threshold. In the decoder, first, the integer part of the size of the MVD can be extracted from the bitstream. Then, the pixel resolution is determined, and a decision can be made as to whether there is a fractional MVD in the bitstream that needs to be parsed (e.g., if fractional pixel resolution is not allowed for a particular extracted MVD integer absolute value, the fractional MVD bits that would be needed for extraction may not be included in the bitstream). The above implementation example related to MVD class - dependent adaptive MVD pixel resolution is applied to MVD size - dependent adaptive MVD pixel resolution. In a specific example, MVD classes above or including the size threshold may only be allowed to have one predetermined value.
[0105] The various implementation examples above are applied to the single - reference mode. These implementations are also applicable to examples of the NEW_NEARMV, NEAR_NEWMV, and / or NEW_NEWMV modes in composite prediction under MMVD. These implementations generally apply to the adaptive resolution for any MVD.
[0106] In some implementation examples, the adaptive MVD resolution is further described below. In the NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD depends on the associated class and the size of the MVD.
[0107] In some examples, fractional MVD is only allowed if the size of the MVD is 1 pixel or less.
[0108] In some examples, when the value of the related MV class is greater than or equal to MV_CLASS_1, only one MVD value is allowed, and the MVD value in each MV class is derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5).
[0109] Table 4 shows the allowed MVD values for each MV class.
Table 4
[0110] In some examples, when the current block is coded as the NEW_NEARMV mode or the NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class. Otherwise, a different context is used to signal mv_joint or mv_class.
[0111] In some implementation examples, joint MVD coding (JMVD) is further described below. To indicate whether the MVDs for two reference lists are signaled together, a new inter-coding mode called JOINT_NEWMV can be applied. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs for reference list 0 and reference list 1 can be signaled together. Therefore, only one MVD called joint_mvd can be signaled and sent to the decoder, and the delta MVs for reference list 0 and reference list 1 can be derived from joint_mvd.
[0112] In some examples, the JOINT_NEWMV mode can be signaled together with the NEAR_NEARMV mode, the NEAR_NEWMV mode, the NEW_NEARMV mode, the NEW_NEWMV mode, and the GLOBAL_GLOBALMV mode. No additional context is added.
[0113] In some examples, when the JOINT_NEWMV mode is signaled and the POC distances between two reference frames and the current frame are different, the MVD can be scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, let the distance between reference frame list 0 and the current frame be represented as td0, and the distance between reference frame list 1 and the current frame be represented as td1. When td0 is greater than or equal to td1, joint_mvd can be directly used for reference list 0, and the mvd for reference list 1 is derived from joint_mvd according to the following formula (1): [Number] is derived based on this.
[0114] Otherwise, when td1 is greater than or equal to td0, joint_mvd can be directly used for reference list 1, and the mvd for reference list 0 is derived from joint_mvd according to the following formula (2): [Number] can be derived based on this.
[0115] In some implementation examples, the improvement of the adaptive MVD resolution is described below.
[0116] In some examples, for the single-reference case, a new inter-coding mode called AMVDMV is added. When the AMVDMV mode is selected, it indicates that AMVD is applied to signal the MVD.
[0117] In some examples, to indicate whether AMVD is applied to the joint MVD coding mode under the JOINT_NEWMV mode, one flag called amvd_flag is added. When the adaptive MVD resolution is applied to the joint MVD coding mode, it is called joint AMVD coding, where the MVDs for two reference frames are signaled together and the accuracy of the MVD is implicitly determined by the magnitude of the MVD. Otherwise, the MVDs for two (or more) reference frames are signaled together and the conventional MVD coding is applied.
[0118] In some implementation examples, the adaptive motion vector resolution (AMVR) is further described below. AMVR that supports a total of seven MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) per (pixel) was first implemented. For each prediction block, the AOMedia Video Model (AVM) encoder can search all the supported accuracy values and signal the best accuracy to the decoder.
[0119] In some examples, to shorten the encoder execution time, two accuracy sets can be supported. Each accuracy set can include four predetermined accuracies. At the frame level, the accuracy set can be adaptively selected based on the value of the maximum accuracy of that frame. The maximum accuracy can be signaled in the frame header. Table 5 below summarizes the accuracy values supported based on the frame-level maximum accuracy.
Table 5
[0120] In some examples, in the AVM software (similar to AV1), there is a frame-level flag indicating whether the MV of a frame includes sub-pel accuracy. The AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, when the accuracy of a block is lower than the maximum accuracy, the motion model and interpolation filter are not signaled. When the accuracy of a block is lower than the maximum accuracy, the motion mode can be estimated as translational motion, and the interpolation filter can be estimated as the REGULAR interpolation filter. Similarly, when the accuracy of a block is either 4 pels or 8 pels, the inter-intra mode is not signaled and is assumed to be 0.
[0121] In some approaches, when the adaptive MVD resolution method is applied, like adaptive MVD coding, the accuracy of the MVD depends on the magnitude of the MVD. The accuracy of the MVD decreases as the magnitude of the MVD increases. As a result, when the adaptive MVD resolution is applied, the prediction for large MVDs may not be very accurate.
[0122] In some approaches, when the adaptive motion vector resolution is explicitly signaled, like AMVR, the accuracy of the MVD depends on the flag being signaled. If the flag being signaled indicates that the accuracy of the MVD is coarse, the MVD can be less accurate.
[0123] In some examples, the methods disclosed herein may be used separately or combined in any order. Also, each of these methods (or embodiments), encoders, and decoders can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium. The term block can be interpreted as a prediction block, a coding block, or a coding unit i.e., CU.
[0124] In this disclosure, the direction of a reference frame can be determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.
[0125] In this disclosure, an explanation of the maximum or highest accuracy for MVD signaling refers to the finest granularity of MVD accuracy. For example, 1 / 16 pel MVD signaling represents a higher accuracy level than the accuracy level of 1 / 8 pel MVD signaling.
[0126] In this disclosure, an explanation of the finest allowable MVD resolution refers to the resolution at which the MVD is signaled. For example, when adaptive MVD resolution is applied, the MVD can be signaled at 1 / 4 pel. However, when bilateral matching is also applied, the actual MVD used for motion compensation can be refined to 1 / 8 pel or higher accuracy without further signaling.
[0127] In some implementations, a Motion Vector Predictor (MVP) and a Motion Vector Difference (MVD) are two important parameters used to represent the motion vector (MV) of a current block. In the inter prediction mode, the MVP and the MVD are used to represent the motion vector of the current block with respect to a reference block in a previous / next frame.
[0128] For example, the MVP is typically calculated by using the motion vectors of adjacent blocks in the same frame or by using the motion vectors of corresponding blocks in a reference frame. The purpose of the MVP is to predict the motion of the current block based on the motion of adjacent or corresponding blocks in the reference frame.
[0129] For example, the MVD is the difference between the motion vector of the current block and the MVP. The MVD represents the deviation of the actual motion vector of the current block from the predicted motion vector based on adjacent blocks or corresponding blocks in the reference frame. The MVD is typically encoded together with the motion vector predictor and sent to the decoder to enable the decoder to reconstruct the motion vector of the current block.
[0130] FIG. 4 is a diagram showing an example of a bilateral matching method for refining the MVD according to some embodiments.
[0131] In some examples, the block matching method utilizes the correlation between the pixels in the block and the pixels in the predicted block. For example, for the pixels of a given block in a frame, the best match with the pixels of the corresponding block in the reference frame is found. The pixel values of the encoded / decoded block are compared with the pixel values of each block in the reference frame, and the block with the closest match is selected. The pixels in the current block are predicted based on the closest match block of the pixels in the reference frame.
[0132] In some aspects / embodiments, when an adaptive MVD resolution (or AMVR) is applied to joint MVD coding, also known as joint AMVD coding, bilateral matching can be used to further refine the MV for the current block. The starting point for MV refinement using bilateral matching is the MV of the current block 402, which is the sum of the MVP for the current block 402 and the MVD signaled (or the MVD derived from the joint MVD). MV refinement by bilateral matching is performed on both the encoder side and the decoder side, and thus the difference between the refined MV and the starting point of MV refinement is not signaled in the bitstream. The predicted block P0 404 is the block behind the current block 402, and the predicted block P1 406 is the block in front of the current block 402.
[0133] FIG. 5 is an exemplary flowchart showing a method 500 for coding a video according to some embodiments. The method 500 may be executed in a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, the method 500 may be executed by executing instructions stored in a memory (e.g., memory 314) of the computing system. The method 500 may be executed by an encoder (e.g., encoder 106) and / or a decoder (e.g., decoder 122).
[0134] Referring to FIG. 5, in one aspect, a video decoder (e.g., decoder 122 of FIG. 2B) and / or a video encoder (e.g., encoder 106 of FIG. 2B) determines whether a joint adaptive motion vector difference (MVD) resolution mode is signaled based on one or more syntax elements from a video stream, where the joint adaptive MVD resolution mode is an inter prediction mode in which MVDs from a first and a second reference frame are signaled at an adaptive MVD pixel resolution (510).
[0135] The video decoder and / or the video encoder receives the signaled MVD of a video block within a current frame from the video stream (520).
[0136] In response to the determination that the joint adaptive MVD resolution mode is signaled, the video decoder and / or the video encoder searches for a first predicted video block in a first reference frame and a second predicted video block in a second reference frame for the video block, where the first predicted video block is a reconstructed / predicted video block before or after the video block, and the second predicted video block is a reconstructed / predicted video block before or after the video block (530).
[0137] A video decoder and / or a video encoder locates (540) a first predicted video block and a second predicted video block based on a minimum difference measured by a cost criterion between the first predicted block and the second predicted block.
[0138] The video decoder and / or the video encoder refines (550) the signaled MVD of the video block based on the located first predicted video block and the located second predicted video block.
[0139] The video decoder and / or the video encoder refines (560) the motion vector (MV) of the video block based on the refined MVD of the video block.
[0140] The video decoder and / or the video encoder reconstructs / processes (570) the video block based at least on the refined MV.
[0141] In one embodiment and / or any combination of the embodiments disclosed herein, for each MVD within an allowable / given search area surrounding the MV of the current block, predicted blocks P0 404 and P1 406 are generated using an MV equal to the sum of the MV (MVP + signaled MVD) and the refined MVD. Then, the difference between P0 404 and P1 406 is calculated, measured by a cost criterion, and the refined MVD having the minimum cost is used as the refined MVD for the current block.
[0142] In some examples, a refined MVD for one reference frame (e.g., reference frame list 0) can be derived from a refined MVD for another reference frame (e.g., reference frame list 1) based on the distances between the two reference frames and the current frame. For example, the refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is derived from the first refined MVD of the first reference frame.
[0143] In some examples, refined_mvd_1 = (td1 / td0) * refined_mvd_0. In this formula, the distance between reference frame list 0 and the current frame is denoted as td0, and the distance between reference frame list 1 and the current frame is denoted as td1. refined_mvd_0 and refined_mvd_1 are the refined MVDs for reference frame list 0 and reference frame list 1, respectively. For example, the refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is derived from the first refined MVD of the first reference frame according to refined_mvd_1 = (td1 / td0) * refined_mvd_0, where td0 is the distance between the first reference frame and the current frame, td1 is the distance between the second reference frame and the current frame, and refined_mvd_0 and refined_mvd_1 are the first refined MVD of the first reference frame and the second refined MVD of the second reference frame, respectively.
[0144] In some examples, the refined MVD for one reference frame (e.g., reference frame list 0) may be mirrored from the other reference frame (e.g., reference frame list 1), i.e., refined_mvd_1 = -refined_mvd_0. Additional constraints may apply to this example. It is that the relative distances between the current frame and the two reference frames are equal, i.e., td0 = td1. For example, the refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is mirrored from the first refined MVD of the first reference frame.
[0145] In one embodiment and / or any combination of the embodiments disclosed herein, only one MVD associated with reference frame list 0 or reference frame list 1 may be refined using bilateral matching, and the other MVD may be derived only from the signaled MVD without further refinement. For example, the refined MVD of the video block is the first refined MVD of the first reference frame, and the second MVD of the second reference frame is the signaled MVD.
[0146] In some examples, when the MVD is signaled for reference frame list 0 (or reference frame list 1) and the MVD for reference frame list 1 (or reference frame list 0) is derived from the signaled MVD, the refinement using bilateral matching is applied to the MVD applied to list 1 (or list 0), but not to the MVD for list 0 (or list 1).
[0147] In one embodiment and / or any combination of the embodiments disclosed herein, the cost criteria for bilateral matching include, but are not limited to, SAD (Sum of Absolute Differences), SSE (Sum of Squared Errors), and / or SATD (Sum of Absolute Transformed Differences).
[0148] In one embodiment and / or any combination of the embodiments disclosed herein, the distortion cost for bilateral matching at one or more specific positions can be modified by a coefficient for making this (these) position(s) more or less preferable during comparison. When the coefficient is greater than 1, the position becomes less preferable. When the coefficient is less than 1, the position becomes more preferable. For example, the cost criterion includes the distortion cost of one or more positions modified by a coefficient for making one or more positions more or less preferable during minimum difference measurement.
[0149] In some examples, the distortion cost of the starting position is scaled by a coefficient less than 1 to make this position more preferable during selection. One further advantage of this approach is that the computational complexity is reduced.
[0150] In one embodiment and / or any combination of the embodiments disclosed herein, the search area size for bilateral matching may depend on the accuracy of the MVD for the current block or the associated MVD class. For example, for the video block described above, searching for a first predicted video block in a first reference frame and a second predicted video block in a second reference frame (530) includes determining the search area size based on the accuracy of the MVD and searching based on the search area size.
[0151] In one embodiment and / or any combination of the embodiments disclosed herein, when AMVD is implicitly applied to joint MVD coding, as the magnitude of the MVD increases, for bilateral matching, the search area size either increases monotonically or remains unchanged.
[0152] In some examples, the search area size is the same for one MVD accuracy but different among multiple different MVD accuracies.
[0153] In some examples, when AMVD is implicitly applied to joint MVD coding, if the MV class of the MVD is above one threshold, such as MV_CLASS_1, the search area size is the same for all MVDs within one MV class.
[0154] In one embodiment and / or any combination of the embodiments disclosed herein, the accuracy / granularity for MV refinement within a given search area for bilateral matching may depend on the accuracy of the MVD and / or the size of the MVD and / or the associated MV class. The accuracy may include, but is not limited to, accuracies of 1 / 64 pel, 1 / 32 pel, 1 / 16 pel, 1 / 8 pel, 1 / 4 pel, 1 / 2 pel, integer pel, 1 pel, 2 pels, 3 pels, 4 pels, ···. For example, refining (550) the signaled MVD of the video block has determining the refinement granularity of the MVD based on the accuracy, size, and / or associated MV class of the MVD.
[0155] In some examples, when AMVD is implicitly applied to joint MVD coding, fractional accuracy MV refinement by bilateral matching is only allowed when the size of the MVD is below one threshold or the associated MV class is below another threshold. In one example, fractional accuracy MV refinement by bilateral matching is only allowed when the size of the MVD is below 1 pel sample. In one example, fractional accuracy MV refinement by bilateral matching is only allowed when the associated MV class is below MV_CLASS_0. For example, determining the refinement granularity of the MVD has performing fractional accuracy MVD refinement only when the size of the MVD is below the threshold.
[0156] In some examples, when AMVD is implicitly applied to joint MVD coding, the accuracy / granularity for MV refinement using bilateral matching may monotonically become coarser as the size of the MVD (or MVD class) increases.
[0157] In some examples, when the AMVR is explicitly signaled for joint MVD coding, the accuracy / granularity for MV refinement using bilateral matching may coarsen monotonically as the accuracy of the MVD decreases. In one example, when the accuracy of the MVD is coarser than 1 pel, such as 2 pel or 4 pel for example, only full pel MVD refinement is supported.
[0158] In some examples, when an adaptive MVD resolution is applied, the finest allowable MVD resolution depends on whether bilateral matching is applied. In one example, when bilateral matching is applied, the finest allowable MVD resolution is lower than the finest allowable MVD resolution without the application of bilateral matching. In one example, when an adaptive MVD resolution is applied, if the finest allowable MVD resolution without the application of bilateral matching is 1 / 8 pel, the finest allowable MVD resolution when bilateral matching is applied is 1 / 4 pel or 1 / 2 pel.
[0159] In one embodiment and / or any combination of the embodiments disclosed herein, the MV refinement for bilateral matching is restricted to a specific predetermined direction, such as the horizontal direction, the vertical direction, or the diagonal direction for example.
[0160] In some examples, the predetermined search direction can be signaled with a high level syntax such as at the sequence level, the frame level, or the slice level for example.
[0161] In one embodiment and / or any combination of the embodiments disclosed herein, the search direction for MV refinement using bilateral matching may depend on the direction of the MVD. For example, for the video block described above, searching for the first predicted video block in the first reference frame and the second predicted video block in the second reference frame (530) includes determining the search direction based on the direction of the MVD and searching based on the search direction.
[0162] In some examples, when the direction of the MVD is along the horizontal or vertical direction, the search direction for MV refinement using bilateral matching is also restricted to the horizontal or vertical direction.
[0163] In some examples, the search direction for MV refinement using bilateral matching can be the same as or perpendicular to the direction of the MVD.
[0164] In one embodiment and / or any combination of the embodiments disclosed herein, one high-level syntax can be signaled to indicate whether bilateral matching is applied to the adaptive MVD resolution (or AMVR). For example, before searching, the decoder / encoder determines whether the bilateral matching mode is signaled based on a second syntax element from the video stream, and searches in response to the determination that the bilateral matching mode is signaled.
[0165] In some examples, this high-level syntax can be signaled at the sequence level, frame level, or slice level. For example, the second syntax element is signaled at one or more of the sequence level, frame level, and / or slice level.
[0166] FIG. 5 shows several logical stages in a particular order, but the stages that are not order-dependent may be rearranged, and other stages may be combined or decomposed. Any rearrangement or other grouping not specifically mentioned will be apparent to those skilled in the art, and the ordering and grouping presented here are not exhaustive. Also, it should be recognized that these stages can be implemented in hardware, firmware, software, or any combination thereof.
[0167] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more sets of instructions configured to be executed by the control circuit, the one or more sets of instructions including instructions for performing any of the methods described herein.
[0168] In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by a control circuit of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein.
[0169] It should be understood that the terms "first", "second", etc. may be used herein to describe various elements, but these elements should not be limited by these terms. These terms are only used to distinguish one element from another.
[0170] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. When used in the description of embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Further understood is that the terms "comprising" and / or "including" when used in this specification specify the presence of the features, integers, steps, operations, elements, and / or components recited, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0171] As used herein, the term "when" can be construed to mean "when" the stated precondition is true, or "in response to" it, or "in response to deciding so", or "in accordance with that decision", or "in response to detecting so", depending on the context. Similarly, the phrases "[when it is determined that the stated precondition is true]", or "[when the stated precondition is true]", or "[when the stated precondition is true]" can, depending on the context, mean "in response to determining" that the stated precondition is true, or "in response to deciding", or "in accordance with the decision" or "in response to detecting", or "in response to detecting so".
[0172] The foregoing description has been presented for purposes of illustration and description with reference to specific embodiments. However, the foregoing exemplary description is not intended to be exhaustive or to limit the claims to the forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments were chosen and described in order to best explain the principles of operation and practical applications, thereby enabling others skilled in the art to understand.
Claims
1. A method for decoding a video stream executed by a computing system having a memory and a control circuit, comprising: Determining whether a joint adaptive motion vector difference (MVD) resolution mode is signaled based on one or more syntax elements from the video stream, wherein the joint adaptive MVD resolution mode is an inter prediction mode in which MVDs from a first and a second reference frame are signaled at an adaptive MVD pixel resolution; Receiving the signaled MVD of a video block within a current frame from the video stream; In response to a determination that the joint adaptive MVD resolution mode is signaled, searching for a first predicted video block within the first reference frame and a second predicted video block within the second reference frame for the video block, wherein the first predicted video block is a reconstructed video block preceding or following the video block, and the second predicted video block is a reconstructed video block preceding or following the video block; Identifying the first predicted video block and the second predicted video block based on a minimum difference measured by a cost criterion between the first predicted video block and the second predicted video block; Refining the signaled MVD of the video block based on the identified first predicted video block and the identified second predicted video block; Refining a motion vector (MV) of the video block based on the refined MVD of the video block; Reconstructing the video block based at least on the refined MV; A method comprising the steps of.
2. The method according to claim 1, wherein the refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is derived from the first refined MVD of the first reference frame.
3. The refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is derived from the first refined MVD of the first reference frame according to refined_mvd_1 = (td1 / td0) * refined_mvd_0, where td0 is the distance between the first reference frame and the current frame, td1 is the distance between the second reference frame and the current frame, and refined_mvd_0 and refined_mvd_1 are the first refined MVD of the first reference frame and the second refined MVD of the second reference frame, respectively. The method according to claim 1.
4. The refined MVD of the video block is the first refined MVD of the first reference frame, and the second refined MVD of the second reference frame is mirrored from the first refined MVD of the first reference frame, the method according to claim 1.
5. The refined MVD of the video block is the first refined MVD of the first reference frame, and the second MVD of the second reference frame is the signaled MVD, the method according to claim 1.
6. The cost criterion includes those obtained by modifying the distortion cost at one or more positions by a coefficient that makes the one or more positions more or less preferable during the measurement of the minimum difference, the method according to claim 1.
7. For the video block, the step of searching for the first predicted video block in the first reference frame and the second predicted video block in the second reference frame includes determining a search area size based on the accuracy of the MVD and searching based on the search area size, the method according to claim 1.
8. The step of refining the signaled MVD of the video block includes determining the refinement granularity of the MVD based on the accuracy, size, and / or related MV class of the MVD, the method according to claim 1.
9. The method according to claim 8, wherein determining the refined granularity of the MVD comprises performing fractional precision MVD refinement only when the size of the MVD is below a threshold value.
10. The method according to claim 1, wherein for the video block, the step of searching for the first predicted video block in the first reference frame and the second predicted video block in the second reference frame comprises determining a search direction based on the direction of the MVD and searching based on the search direction.
11. The method according to claim 1, further comprising determining, based on a second syntax element from the video stream, whether a bilateral matching mode is signaled before searching, and searching in response to the determination that the bilateral matching mode is signaled.
12. The method according to claim 11, wherein the second syntax element is signaled at one or more of sequence level, frame level, and / or slice level.
13. The method according to claim 11, wherein when the joint adaptive MVD resolution mode is signaled, the finest allowable MVD resolution depends on whether the bilateral matching mode is signaled.
14. A computing system having a memory storing computer instructions and a control circuit communicating with the memory, wherein the control circuit is configured to cause the computing system to execute the method according to any one of claims 1 to 13 when executing the computer instructions.
15. A computer program for causing a computer to execute the method according to any one of claims 1 to 13.