Flexible Scaling Factors for Joint MVD Coding

JP2025517262A5Pending Publication Date: 2025-11-17TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518496
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-07
Filing Date
2022-11-08
Publication Date
2025-11-17

AI Technical Summary

Technical Problem

Existing video coding technologies, such as AV1 and HEVC, assume linear motion in joint motion vector difference (JMVD) coding modes, which is not always accurate as motion between reference frames can be non-linear.

Method used

A method for video coding that involves determining whether JMVD is used to predict a coding block, obtaining a scaling factor, and applying it to the components of the joint MVD along predefined directions to derive a motion vector difference (MVD) for reconstructing the coding block.

Benefits of technology

This approach improves the accuracy of motion prediction by accounting for non-linear motion between reference frames, leading to more efficient video coding and better compression performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for video coding includes the steps of obtaining a coding block of a video bitstream, determining whether joint coding of motion vector difference (JMVD) is used to predict the coding block, obtaining a scaling factor based on a determination that the JMVD is selected and used to predict the coding block, deriving a motion vector difference (MVD) for one or more reference frame lists based on application of the scaling factor to one or more components of the JMVD along one or more predefined directions, and reconstructing the coding block based on at least the derived MVD.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 342,441, filed May 16, 2022, and U.S. Patent Application No. 17 / 982,139, filed November 7, 2022, the contents of which are expressly incorporated by reference in their entireties into this application.

[0002] This disclosure relates to a set of advanced image and video coding techniques, and more specifically, to an improved scheme for joint coding of motion vector differences (JMVD). [Background technology]

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. As a successor to VP9, ​​it was developed by the Alliance for Open Media (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project were contributed by previous research efforts by members of the Alliance. Individual contributors started experimental technology platforms several years ago, namely Xiph's / Mozilla's Daala, which already released its code in 2010, Google's experimental VP9 evolution project VP10, announced on September 12, 2014, and Cisco's Thor, on August 11, 2015. Building on the VP9 code base, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version 0.1.0 of the AV1 reference codec was released on April 7, 2016. The Alliance announced the release of the AV1 Bitstream Specification on March 28, 2018, along with reference software-based encoders and decoders. The working version 1.0.0 of the specification was published on June 25, 2018. The working version 1.0.0 of the specification, including Errata 1, was released on January 8, 2019. The AV1 Bitstream Specification includes reference video codecs.

[0004] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3) and 2016 (version 4). Since then, they have been studying the potential need for standardization of future video coding technologies that may significantly exceed HEVC in compression capabilities. In October 2017, they announced a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for 360 video categories had been submitted, respectively. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET (Joint Video Exploration Team-Joint Video Expert Team) meeting. After careful evaluation, JVET officially started the standardization of the next generation video coding beyond HEVC, namely the so-called Versatile Video Coding (VVC).

[0005] Also, in the case of JMVD, there are technical problems in assuming linear motion in the JMVD coding mode because the motion between two reference frames is not always linear, e.g., the motion may be slower or faster from a backward reference frame to a forward reference frame. Therefore, a technical solution to such problems is desired. Summary of the Invention [Means for solving the problem]

[0006] According to aspects of some embodiments, a method for video coding performed by at least one processor is provided, the method including obtaining a coding block of video data, determining whether joint coding of motion vector difference (JMVD) is used to predict the coding block, obtaining a scaling factor based on a determination that JMVD is used to predict the coding block, and deriving a motion vector difference (MVD) of at least one reference frame list based on application of the scaling factor to one or more components of the joint MVD along one or more predefined directions, and reconstructing the coding block based on at least the derived MVD.

[0007] According to other aspects of some embodiments, there is also provided an apparatus and computer readable medium consistent with the present methods.

[0008] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Diagram 2] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Diagram 3] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 4] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Diagram 5] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 6] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 7] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 8] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 9A] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 9B] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 10A] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 10B] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 10C] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 11] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 12] FIG. 1 is a simplified diagram of a diagram according to some embodiments. [Figure 13] FIG. 1 is a simplified flow diagram according to some embodiments. [Figure 14] FIG. 1 is a schematic diagram of a diagram according to some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] The proposed functions described below may be used separately or combined in any order. Furthermore, the embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0011] According to some embodiments, there is at least one memory configured to store computer program code and at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code including: an acquisition code configured to cause the at least one processor to acquire a coding block of the video data; a decision code configured to cause the at least one processor to determine whether joint coding of motion vector difference (JMVD) is used to predict the coding block; a further acquisition code configured to cause the at least one processor to acquire a scaling factor in response to determining that JMVD is used to predict the coding block; a derivation code configured to cause the at least one processor to derive a motion vector difference (MVD) of the at least one reference frame list based on application of the scaling factor to one or more components of the JMVD along one or more predefined directions; and a reconstruction code configured to cause the at least one hardware processor to reconstruct the coding block based at least on the derived MVD.

[0012] According to some example embodiments, the MVD is derived further based on any of the distances between the reference frame and the current frame.

[0013] According to some example embodiments, the step of deriving the MVD includes a step of determining whether the flag indicates that at least one of the scaling factors is not equal to a first predefined default value, where at least one of the scaling factors is used to derive the MVD from the JMVD for one of the reference frames, and another one of the scaling factors used to derive the MVD from the JMVD for a second one of the reference frames is set to a second predefined default value.

[0014] According to some example embodiments, the step of obtaining the scaling factor is based on obtaining at least one flag signaled in a bitstream of the coding block, the at least one flag indicating a scaling factor for at least one of the components along one or more predefined directions.

[0015] According to some example embodiments, deriving the MVD includes determining whether a first flag indicates that at least one of the scaling factors is not equal to a first predefined default value, and in response to determining that the first flag indicates that at least one of the scaling factors is not equal to the first predefined default value, determining whether a scaling factor is applied to at least one direction of the MVD and at least one of the predefined directions based on a value of a second flag.

[0016] According to some example embodiments, deriving the MVD includes applying one or more scaling factors equally to both of the predefined directions based on determining that the second flag indicates both of the predefined directions.

[0017] According to some example embodiments, obtaining the scaling coefficients includes obtaining indices of the scaling coefficients in a lookup table, where the lookup table indicates that at least one pair at a first index of the scaling coefficient indices has the same scaling coefficient value in both of the predefined directions and the lookup table indicates that at least a second pair at a second index of the scaling coefficient indices has different scaling coefficient values ​​between one of the predefined directions.

[0018] According to some example embodiments, at least one of the same scaling factor value and the different scaling factor value is a fractional scaling factor value, and at least one other of the same scaling factor value and the different scaling factor value is m / M, where M is a power of 2 and m and n are integers.

[0019] According to some example embodiments, the scaling factor is derived based on at least one coding information of a quantization step size, a quantization parameter, a block size, a difference between the motion vector prediction block of the current block, an MVD class, a reference picture, and an MVD scaling factor of a neighboring block in the vicinity of the coding block.

[0020] According to some example embodiments, at least one of the frame header, slice header, and sequence header indicates whether to signal a scaling factor for the bitstream of the coding block.

[0021] 1 illustrates a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The communication system 100 may include at least two terminals 102, 103 interconnected via a network 105. For unidirectional transmission of data, a first terminal 103 may code video data at a local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the coded video data of the other terminal from the network 105, decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media provision applications, etc.

[0022] 1 shows a second pair of terminals 101 and 104 provided to support bidirectional transmission of coded video, such as may occur during a video conference. For the bidirectional transmission of data, each terminal 101 and 104 may code captured video data at a local location for transmission to the other terminal over network 105. Each terminal 101 and 104 may also receive coded video data transmitted by the other terminal, decode the coded data, and display the reconstructed video data on a local display device.

[0023] In FIG. 1, terminals 101, 102, 103, and 104 may be illustrated as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 105 represents any number of networks that convey coded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 105 may not be important to the operation of the present disclosure unless described below herein.

[0024] 2 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as one example of an application of the disclosed subject matter. The subject matter of this disclosure is equally applicable to other video-enabled applications, such as, for example, video conferencing, digital television, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0025] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213. The sample stream 213 may be emphasized as a high data volume compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may be emphasized as a lower data volume compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 to decode the incoming copy 208 of the encoded video bitstream and create an outgoing video sample stream 210 that may be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to a particular video coding / compression standard, examples of which are mentioned above and further described herein.

[0026] FIG. 3 may be a functional block diagram of a video decoder 300 according to one embodiment of the present invention.

[0027] The receiver 302 may receive one or more codec video sequences to be decoded by the decoder 300, in the same or other embodiments, one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel 301, which may be a hardware / software link to a storage device that stores the encoded video data. The receiver 302 may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, that may be forwarded to a respective usage entity (not shown). The receiver 302 may separate the coded video sequences from the other data. To combat network jitter, a buffer memory 303 may be coupled between the receiver 302 and the entropy decoder / parser 304 (hereinafter "parser"). If the receiver 302 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer 303 may not be needed or may be small. For use in a best effort packet network such as the Internet, buffer 303 may be required and may be relatively large, and may advantageously be of adaptive size.

[0028] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from the entropy coded video sequence. These categories of symbols include information used to manage the operation of the decoder 300 and potentially information for controlling a rendering device, such as a display 312, that is not an integral part of the decoder but may be coupled to it. The rendering device control information may be in the form of supplemental enhancement information (SEI message) or video usability information parameter set fragments (not shown). The parser 304 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without contextual dependency, etc. The parser 304 may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to that group. The subgroups may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The entropy decoder / parser may also extract from the coded video sequence information such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0029] The parser 304 may perform an entropy decoding / parsing operation on the video sequence received from the buffer 303 to produce symbols 313. The parser 304 may receive the encoded data and selectively decode particular symbols 313. Additionally, the parser 304 may determine whether a particular symbol 313 should be provided to the motion compensated prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0030] The reconstruction of symbols 313 may involve several different units depending on the type of coded video picture or part thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be governed by subgroup control information parsed from the coded video sequence by parser 304. The flow of such subgroup control information between parser 304 and the following units is not shown for clarity.

[0031] In addition to the functional blocks already mentioned, the decoder 300 can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate:

[0032] The first unit is a scalar / inverse transform unit 305. The scalar / inverse transform unit 305 receives quantized transform coefficients and control information including the transform to use, block size, quantization factor, quantization scaling matrix, etc. as symbols 313 from the parser 304. It can output blocks containing sample values ​​that can be input to the aggregator 310.

[0033] In some cases, the output samples of the scalar / inverse transform unit 305 may relate to intra-coded blocks, i.e. blocks that do not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309 to generate a block of the same size and shape as the block being reconstructed. The aggregator 310 adds, possibly on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the scalar / inverse transform unit 305.

[0034] In other cases, the output samples of the scalar / inverse transform unit 305 may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion compensated prediction unit 306 may access the reference picture memory 308 to fetch samples used for prediction. After motion compensating the fetched samples according to the symbols 313 relating to the block, these samples may be added by the aggregator 310 to the output of the scalar / inverse transform unit to generate output sample information (in this case referred to as residual samples or residual signals). The addresses in the reference picture memory format from which the motion compensation unit fetches the prediction samples may be controlled by a motion vector and may be made available to the motion compensation unit in the form of symbols 313, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​fetched from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0035] The output samples of the aggregator 310 may be subjected to various loop filtering techniques in a loop filter unit 311. The video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit 311 as symbols 313 from the parser 304, but may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) part of a coded video sequence, or to previously reconstructed and loop filtered sample values.

[0036] The output of the loop filter unit 311 can be a sample stream that can be output to the rendering device 312 as well as stored in the reference picture memory 557 for use in future inter-picture prediction.

[0037] Once fully reconstructed, a particular coded picture can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 304), the current reference picture 309 can become part of the reference picture buffer 308, and the new current picture memory can be reallocated before starting reconstruction of the following coded picture.

[0038] The video decoder 300 may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as ITU-T Rec. H.265. The coded video sequence may comply with the syntax specified by the video compression technique or standard being used in the sense of adhering to the syntax of the video compression technique or standard, as specified in the video compression technique document or standard, specifically in a profile document therein. Also, what is required for compliance may be that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits a maximum picture size, a maximum frame rate, a maximum reconstruction sample rate (e.g., measured in megasamples per second), a maximum reference picture size, etc. The limits set by the level may be further limited in some cases by the specification of a hypothetical reference decoder (HRD) and metadata for HRD buffer management signaled within the coded video sequence.

[0039] In one embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal layer, a spatial layer, or a signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0040] FIG. 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.

[0041] The encoder 400 may receive video samples from a video source 401 (not part of the encoder), which may capture video images to be coded by the encoder 400 .

[0042] The video source 401 may provide the source video sequence to be coded by the encoder (303) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...) and suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media delivery system, the video source 401 may be a storage device that stores previously prepared video. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a number of separate pictures that give motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0043] According to one embodiment, the encoder 400 may code and compress pictures of a source video sequence into a coded video sequence 410 in real-time or under any other time constraint as required by an application. Providing an appropriate coding rate is one function of the controller 402. The controller controls and is operatively coupled to other functional units as described below. Coupling is not shown for clarity. Parameters set by the controller may include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, etc.), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art may easily identify other functions of the controller 402 as they may relate to the video encoder 400 being optimized for a particular system design.

[0044] Some video encoders operate in what those skilled in the art will readily recognize as a "coding loop." As an oversimplified explanation, the coding loop can consist of an encoding part of the encoder 402 (hereafter "source coder") (responsible for creating symbols based on the input picture to be coded and reference pictures), and a (local) decoder 406 built into the encoder 400 that reconstructs the symbols to create sample data that the (remote) decoder will also create (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream is input to a reference picture memory 405. Since decoding of the symbol stream results in bit-exact results regardless of the location of the decoder (local or remote), the reference picture buffer contents are also bit-exact between the local and remote encoders. In other words, the predictive part of the encoder "sees" exactly the same sample values ​​as the decoder "sees" when using prediction during decoding as reference picture samples. This basic principle of reference picture synchrony (and the resulting drift when synchrony cannot be maintained, for example due to channel errors) is well known to those skilled in the art.

[0045] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300, which has already been described in detail above in relation to Figure 3. However, with brief reference also to Figure 4, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder 408 and parser 304 may be lossless, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303 and parser 304, may not be fully implemented in the local decoder 406.

[0046] At this point, it can be said that any decoder technique, other than analysis / entropy decoding, present in the decoder must necessarily also be present in the corresponding encoder in substantially the same functional form. The description of the encoder technique can be omitted, since it is the inverse of the decoder technique, which has been described generically. Only in certain areas is a more detailed description required, which is provided below.

[0047] As part of its operation, the source coder 403 may perform motion compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence, designated as “reference frames.” In this method, the coding engine 407 codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.

[0048] The local video decoder 406 may decode the coded video data of frames that may be designated as reference frames based on the symbols created by the source coder 403. The operation of the coding engine 407 may advantageously be a lossy process. When the coded video data may be decoded in a video decoder (not shown in FIG. 4), the reconstructed video sequence may usually be a copy of the source video sequence with some errors. The local video decoder 406 may replicate the decoding process that may be performed by the video decoder on the reference frames and store the reconstructed reference frames in the reference picture cache 405. In this way, the encoder 400 may locally store copies of reconstructed reference frames that have common content as the reconstructed reference frames that will be obtained by the far-end video decoder (without transmission errors).

[0049] The predictor 404 may perform a prediction search for the coding engine 407. That is, for a new frame to be coded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor 404 may operate on a pixel block by pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory 405.

[0050] The controller 402 may manage the coding operations of the video coder 403, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0051] The output of all the aforementioned functional units may undergo entropy coding in an entropy coder 408. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.

[0052] The transmitter 409 may buffer the coded video sequence created by the entropy coder 408 and prepare it for transmission over a communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 may merge the coded video data from the video coder 403 with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0053] The controller 402 may manage the operation of the encoder 400. During coding, the controller 405 may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following frame types:

[0054] An intra picture (I-picture) may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow various types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0055] A predicted picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block.

[0056] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0057] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and may be coded block by block. A block may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, a block of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). A pixel block of a P-picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. A pixel block of a B-picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0058] Video coder 400 may perform coding operations in accordance with a given video coding technique or standard, such as ITU-T Rec. H.265. In its operations, video coder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0059] In one embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source coder 403 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0060] FIG. 5 illustrates the intra prediction modes used in HEVC and JEM. In order to capture any edge direction presented in natural video, the number of directional intra modes is expanded from 33 used in HEVC to 65. The additional directional modes in JEM above HEVC are illustrated as dotted arrows in FIG. 1(b), while the planar and DC modes remain the same. These denser directional intra prediction modes apply to all block sizes and to both luma and chroma intra prediction. As shown in FIG. 5, directional intra prediction modes identified by dotted arrows associated with odd intra prediction mode indexes are referred to as odd intra prediction modes. Directional intra prediction modes identified by solid arrows associated with even intra prediction mode indexes are referred to as even intra prediction modes. In this specification, the directional intra prediction modes indicated by solid or dotted arrows in FIG. 5 are also referred to as angular modes.

[0061] In JEM, a total of 67 intra prediction modes are used for luma intra prediction. To code intra modes, a Most Probable Mode (MPM) list of size 6 is constructed based on the intra modes of neighboring blocks. If the intra mode is not from the MPM list, a flag is signaled to indicate whether the intra mode belongs to the selected modes. In JEM-3.0, there are 16 selected modes, which are uniformly selected every 4th angular mode. In JVET-D0114 and JVET-G0060, 16 secondary MPMs are derived to replace the uniformly selected modes.

[0062] 6 illustrates N reference layers utilized for the intra-directional mode. There is a block unit 611, a segment A 601, a segment B 602, a segment C 603, a segment D 604, a segment E 605, a segment F 606, a first reference layer 610, a second reference layer 609, a third reference layer 608, and a fourth reference layer 607.

[0063] In both HEVC and JEM, as well as some other standards such as H.264 / AVC, the reference samples used to predict the current block are limited to the nearest reference line (row or column). In the method of multiple reference line intra prediction, the number of candidate reference lines (rows or columns) is increased from 1 (i.e., the nearest) to N for intra-directional mode, where N is an integer equal to or greater than 1. Figure 2 cites a 4x4 prediction unit (PU) as an example to illustrate the concept of the multiple line intra-directional prediction method. The intra-directional mode can arbitrarily select one of N reference layers to generate a predictor. In other words, the predictor p(x,y) is generated from one of the reference samples S1, S2, ..., SN. A flag is signaled to indicate which reference layer is selected for the intra-directional mode. If N is set to 1, the intra-directional prediction method is the same as the conventional method of JEM2.0. In Fig. 6, the reference lines 610, 609, 608 and 607 are composed of six segments 601, 602, 603, 604, 605 and 606 with an upper-left reference sample. In this specification, the reference hierarchy is also called a reference line. The coordinates of the upper-left pixel in the current block unit are (0,0), and the upper-left pixel of the first reference line is (-1,-1).

[0064] In JEM, for the luma component, the neighboring samples used for intra prediction sample generation are filtered before the generation process. The filtering is controlled by a given intra prediction mode and transform block size. If the intra prediction mode is DC or the transform block size is equal to 4×4, the neighboring samples are not filtered. If the distance between a given intra prediction mode and the vertical mode (or horizontal mode) is greater than a predefined threshold, the filtering process is enabled. A [1,2,1] filter and a bilinear filter are used for filtering the neighboring samples.

[0065] The position-dependent intra prediction synthesis (PDPC) method is an intra prediction method that induces a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. Each prediction sample pred[x][y] located at (x, y) is calculated as follows: pred[x][y]=(wL*R -1,y +wT*R x,-1 +wTL*R -1,-1 +(64 wL-wT-wTL)*pred[x][y]+32)>>6 (Formula 2-1) In the formula, R x,-1 , R -1,y represent the unfiltered reference samples located above and to the left of the current sample (x, y), respectively, and R -1,-1 represents the unfiltered reference sample located in the top-left corner of the current block. The weights are calculated as follows: wT=32>>((y<<1)>>shift) (Eq. 2-2) wL=32>>((x<<1)>>shift) (Formula 2-3) wTL=-(wL>>4)-(wT>>4) (Equation 2-4) shift = (log2(width) + log2(height) + 2) >> 2 (equation 2-5).

[0066] FIG. 7 illustrates a diagram 700 of DC mode PDPC weights (wL, wT, wTL) for (0,0) and (1,0) positions in one 4×4 block. When PDPC is applied to DC, planar, horizontal, and vertical intra modes, no additional boundary filters such as HEVC DC mode boundary filters or horizontal / vertical mode edge filters are required. FIG. 7 illustrates the definition of reference samples Rx,-1, R-1,y, and R-1,-1 for PDPC applied to the top right diagonal mode. The predicted sample pred(x',y') is located at (x',y') in the predicted block. The coordinate x of the reference sample Rx,-1 is given by x=x'+y'+1, and similarly the coordinate y of the reference sample R-1,y is given by y=x'+y'+1.

[0067] 8 illustrates a Local Illumination Compensation (LIC) diagram 800, which is based on a linear model of illumination changes with a scaling factor a and an offset b, and which is adaptively enabled or disabled for each coding unit (CU) coded in inter mode.

[0068] When LIC is applied to a CU, a least square error method is adopted to derive parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as illustrated in Figure 8, the subsampled (2:1 subsampled) neighboring samples of the CU and the corresponding samples in the reference picture (identified by the motion information of the current CU or sub-CU) are used. The IC parameters are derived and applied separately for each prediction direction.

[0069] If the CU is coded in merge mode, the LIC flag is copied from the neighboring block in a manner similar to the motion information copy in merge mode, otherwise the LIC flag is signaled to the CU to indicate whether LIC is applied or not.

[0070] 9A illustrates intra-prediction modes 900 used in HEVC. In HEVC, there are a total of 35 intra-prediction modes, of which mode 10 is a horizontal mode, mode 26 is a vertical mode, and modes 2, 18, and 34 are diagonal modes. The intra-prediction modes are signaled by three most probable modes (MPMs) and the remaining 32 modes.

[0071] 9B illustrates that in a VVC embodiment, there are a total of 87 intra-prediction modes, with mode 18 being the horizontal mode, mode 50 being the vertical mode, and modes 2, 34 and 66 being diagonal modes. Modes -1 through -10 and modes 67 through 76 are referred to as wide-angle intra-prediction (WAIP) modes.

[0072] A prediction sample pred(x,y) located at position (x,y) is predicted using a linear combination of reference samples according to the intra prediction mode (DC, planar, angular) and the PDPC representation. pred(x,y)=(wL×R-1,y+wT×Rx,-1-wTL×R-1,-1+(64-wL-wT+wTL)×pred(x,y)+32)>>6 In the formula, Rx,-1, R-1,y represent the reference samples located above and to the left of the current sample (x, y), respectively, and R-1,-1 represents the reference sample located in the upper left corner of the current block.

[0073] For DC mode, the weights are calculated for a block with dimensions width and height as follows: wT=32>>((y<<1)>>nScale), wL=32>>((x<<1)>>nScale), wTL=(wL>>4)+(wT>>4), Here, nScale=(log2(width)-2+log2(height)-2+2)>>2, where wT represents the weighting factor of the reference sample located on the above reference line with the same horizontal coordinate, wL represents the weighting factor of the reference sample located on the left reference line with the same vertical coordinate, and wTL represents the weighting factor of the top-left reference sample of the current block, and nScale specifies how fast the weighting factor decreases along the axis (wL decreases from left to right, or wT decreases from top to bottom), i.e., specifies the weighting factor decrease rate, which is the same along the x-axis (from left to right) and y-axis (from top to bottom) in the current design. And 32 represents the initial weighting factor of the neighboring samples, and the initial weighting factor is also the top (left or top-left) weighting assigned to the top-left sample in the current CB, and the weighting factor of the neighboring samples in the PDPC process should be less than or equal to this initial weighting factor.

[0074] For planar mode, wTL=0, while for horizontal mode, wTL=wT, and for vertical mode, wTL=wL. The PDPC weights can be calculated with additions and shifts only. The value of pred(x,y) can be calculated in one step using Equation 1.

[0075] In this specification, the proposed methods may be used separately or combined in any order. Moreover, each of the methods (or embodiments), the encoder, and the decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium. In the following, the term block may be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU.

[0076] FIG. 10A shows an example 1000 of block partitioning using QTBT, and FIG. 10B shows the corresponding tree representation 1001. The solid lines indicate quadtree partitioning, and the dotted lines indicate binary tree partitioning. At each partition (i.e., non-leaf) node of the binary tree, one flag is signaled to indicate which partition type (i.e., horizontal or vertical) is used, with 0 indicating horizontal partitioning and 1 indicating vertical partitioning. In the case of quadtree partitioning, there is no need to specify the partition type, since quadtree partitioning always partitions a block both horizontally and vertically to generate four sub-blocks of equal size.

[0077] In HEVC, CTUs are partitioned into CUs using a quadtree structure called a coding tree to adapt to various local characteristics. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can be further partitioned into one, two, or four PUs depending on the PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be partitioned into transform units (TUs) according to another quadtree structure, such as the coding tree of the CU. One of the key features of the HEVC structure is that it has a multiple partition concept, including CUs, PUs, and TUs.

[0078] According to an embodiment, the QTBT structure removes the concept of multiple partition types, i.e., removes the concept of separation of CU, PU, ​​and TU, and supports higher flexibility of CU partition shape. In the QTBT block structure, a CU can have either a square or rectangular shape. In the flow diagram 1100 of FIG. 11, according to an exemplary embodiment, a coding tree unit (CTU) or CU obtained in S11 is first partitioned by a quadtree structure in S12. It is further determined whether the quadtree leaf node is partitioned by a binary tree structure in S14, and if so, in S15, as described in FIG. 10C, for example, there are two partition types in binary tree partitioning: symmetric horizontal partitioning and symmetric vertical partitioning. The binary tree leaf node is called a coding unit (CU), and its segmentation is used for prediction and transformation processing without further partitioning. This means that the CU, PU, ​​and TU have the same block size in the QTBT coding block structure. In VVC, a CU may be composed of coding blocks (CBs) of different color components, e.g., for P and B slices in 4:2:0 chroma format, a CU contains one luma CB and two chroma CBs, or it may be composed of a single component CB, e.g., for an I slice, a CU contains only one luma CB or only two chroma CBs.

[0079] According to an embodiment, the following parameters are defined for the QTBT splitting scheme: -CTU size: Root node size of the quadtree, same concept as HEVC, -MinQTSize: The minimum allowed quadtree leaf node size, -MaxBTSize: Maximum allowed binary tree root node size, -MaxBTDepth: The maximum allowed binary tree depth, and -MinBTSize: The minimum allowed binary tree leaf node size.

[0080] In one example of a QTBT partitioning structure, the CTU size is set as 128×128 luma samples, with two corresponding 64×64 blocks of chroma samples. Here, QT is a quadtree, MinQTSize is set as 16×16, MaxBTSize is set as 64×64, MinBTSize (both width and height) is set as 4×4, and MaxBTDepth is set to 4. The quadtree partitioning is first applied to the CTU, generating quadtree leaf nodes in S12 or S15. The quadtree leaf nodes can have sizes from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). If the leaf quadtree node is 128×128, it is not further partitioned by the binary tree since its size exceeds MaxBTSize (i.e., 64×64) as checked in S14. Otherwise, the leaf quadtree node may be further split by the binary tree in S15. Thus, the quadtree leaf node is also the root node of the binary tree, and the depth of the binary tree is 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further splits are considered in S14. When the width of the binary tree node is equal to MinBTSize (i.e., 4), no further horizontal splits are considered in S14. Similarly, when the height of the binary tree node is equal to MinBTSize, no further vertical splits are considered in S14. Signaling in S16 is provided for the progression of the leaf node of the binary tree to be further processed by prediction and transformation processing, as described below with respect to the syntax describing QT / TT / BT sizes, and without further splitting, as described herein with respect to such prediction and transformation processing. Such signaling may be provided in S13 after S12, as shown in FIG. 11, according to an exemplary embodiment. In JEM, the maximum CTU size is 256x256 luma samples.

[0081] In addition, according to an embodiment, the QTBT scheme supports the ability / flexibility for luma and chroma to have separate QTBT structures. Currently, for P and B slices, the luma CTB and chroma coding tree block (CTB) in one CTU share the same QTBT structure. However, for I slices, the luma CTB is divided into CUs by a QTBT structure, and the chroma CTB is divided into chroma CUs by another QTBT structure. This means that a CU in an I slice is composed of a coding block of a luma component or a coding block of two chroma components, and a CU in a P or B slice is composed of coding blocks of all three color components.

[0082] In HEVC, inter prediction for small blocks is restricted to reduce memory access for motion compensation, so bi-prediction is not supported for 4 × 8 and 8 × 4 blocks, and inter prediction is not supported for 4 × 4 blocks. These restrictions have been removed in QTBT implemented in JEM-7.0.

[0083] FIG. 10C shows a simplified block diagram 1100 VVC for the included multi-type tree (MTT) structure 1002, which is a combination of the illustrated quadtree (QT) with nested binary trees (BT) and ternary / ternary trees (TT), QT / BT / TT. A CTU or CU is first recursively divided into square shaped blocks by QT. Then, each QT leaf may be further divided by BT or TT, and the BT and TT divisions can be applied recursively and interleaved, but no further QT divisions can be applied. In all related proposals, TT divides a rectangular block into three blocks vertically or horizontally using a 1:2:1 ratio (thus avoiding widths and heights that are not powers of two). For partitioning conflict prevention, additional partitioning constraints are typically imposed on the MTT for QT / BT / TT block partitioning in VVC with respect to blocks 1103 (quadrant), 1104 (bisection, JEM), and 1105 (ternary) to avoid overlapping partitions (e.g., prohibit vertical / horizontal bisection at intermediate partition resulting from vertical / horizontal 3-way partitioning), as shown in simplified diagram 1002 of Fig. 10C. Further restrictions may be set on the maximum depth of BT and TT.

[0084] The main advantage of such a ternary tree partitioning, pointed out above as ternary block 1105, as a complement to quadtree and binary tree partitioning, is that ternary tree partitioning can capture objects located at the block center, while quadtree and binary trees always partition along the block center, and the width and height of the proposed ternary tree partitions are always powers of two, so no additional transformations are required.

[0085] The design of two-level trees is primarily motivated by reduced complexity. In theory, the complexity of traversing a tree is T D where T represents the number of split types and D is the depth of the tree.

[0086] FIG. 11 shows an example of block partitioning in VP9 and AVI 1100, where an exemplary coding tree unit (CTU) 1111 for VP9 shows that VP9 uses a 4-way partition tree starting from the 64x64 level 1112 to the 4x4 level 1113 with some additional restrictions on blocks 8x8 and below as shown in the top half of level 1113. Note that partitions designated as R are referred to as recursive in that the same partition tree is repeated at lower scales until the lowest 4x4 level is reached. An exemplary CTU 1104 for AV1 not only expands the partition tree to a 10-way structure 1116, but also increases the maximum size (called a superblock in VP9 / AV1 terminology) to start at the 128x128 level 1115. Note that this includes 4:1 / 1:4 rectangular partitions that did not exist in VP9. And none of the rectangular partitions can be further subdivided. In addition, AV1 adds more flexibility in the use of decompositions below the 8x8 level, in the sense that 2x2 chroma inter prediction is allowed in certain cases.

[0087] And in HEVC, coding tree units (CTUs) may be divided into coding units (CUs) using a quadtree structure represented as a coding tree to adapt to various local characteristics. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four prediction units (PUs) according to the PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder on a PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU may be divided into transform units (TUs) according to another quadtree structure, such as the coding tree of the CU. One of the important features of the HEVC structure is to have multiple partition concepts, including CUs, PUs, and TUs. In HEVC, a CU or TU can only be square in shape, while a PU can be square or rectangular in shape for inter-predicted blocks. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform is performed on each sub-block, i.e., TU. Each TU can be further split recursively (using quadtree partitioning) into smaller TUs called residual quadtrees (RQTs), and at picture boundaries, HEVC uses implicit quadtree partitioning such that a block continues to be quadtree partitioned until its size fits within the picture boundary.

[0088] FIG. 12 also illustrates an example 1200 related to a merge mode with motion vector difference (MMVD) according to an exemplary embodiment. For example, in addition to the merge mode in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU, a merge mode with motion vector difference (MMVD) is introduced into VVC. Then, the MMVD flag may be signaled immediately after sending the skip flag and the merge flag to specify whether the MMVD mode is used for the CU. And in the MMVD, after a merge candidate is selected, further information is further refined by the signaled motion vector difference (MVD) information to include a merge candidate flag, an index to specify the magnitude of the motion, and an index to indicate the direction of the motion. In the MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV basis. The merge candidate flag may be signaled to specify which one is used.

[0089] The distance index specifies the magnitude information of the motion and indicates a predefined offset from the starting point. And FIG. 12 shows L0 reference 1201 and L1 reference 1202 where the offset is added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1.

[0090] [Table 1]

[0091] According to an exemplary embodiment, the direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 2 below. The meaning of the MVD code may differ according to the information of the starting MV. For example, if the starting MV is a uni-predictive MV or a bi-predictive MV where both lists point to the same side of the current picture (i.e., the Picture Order Counts (POC) of the two references are both greater than or both less than the POC of the current picture), the code in Table 2 specifies the code of the MV offset added to the starting MV. And / or if the starting MV is a bi-predictive MV where two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture), the code in Table 2 specifies the code of the MV offset added to the starting MV. OC and the other reference's POC is smaller than the POC of the current picture), if the difference in POC in list 0 is greater than the difference in POC in list 1, then the signs in Table 2 specify the signs of the MV offsets added to the MV components of list 0 of the starting MV and the signs of the MV in list 1 have opposite values. Otherwise, if the difference in POC in list 1 is greater than the difference in POC in list 0, then the signs in Table 2 specify the signs of the MV offsets added to the MV components of list 1 of the starting MV and the signs of the MV in list 0 have opposite values.

[0092] According to an example embodiment, the MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, the MVD in list 1 is scaled. If the POC difference in L1 is greater than L0, the MVD in list 0 is scaled as well. If the starting MV is uni-predicted, the MVD is added to the available MV.

[0093] [Table 2]

[0094] According to an exemplary embodiment, there may be symmetric MVD coding, and the MVD may be scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, the MVD in list 1 is scaled. If the POC difference in L1 is greater than L0, the MVD in list 0 is scaled as well. If the starting MV is uni-predicted, the MVD is added to the available MV.

[0095] And according to an exemplary embodiment, in VVC, in addition to the normal unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional MVD signaling may be applied. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not signaled but is derived. The decoding process of the symmetric MVD mode is as follows: 1. At the slice level, the variables BiDirPredFlag, RefIdxSymL0 and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is 1, then BiDirPredFlag is set equal to 0. Otherwise, if the closest reference picture in list 0 and the closest reference picture in list 1 form a backward-forward pair of reference pictures of a forward-backward pair of reference pictures, then BiDirPredFlag is set to 1 and both the reference pictures in list 0 and list 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2. At the CU level, if a CU is bi-predictively coded and BiDirPredFlag is equal to 1, a symmetric mode flag is explicitly signaled indicating whether symmetric mode is used or not.

[0096] And if the symmetric mode flag is true, then only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set equal to the pair of reference pictures, respectively. MVD1 is set equal to (-MVD0).

[0097] According to an example embodiment, CWG-B018 may have inter mode coding, and in AV1, for each coding block in an inter frame, if the mode of the current block is an inter coding mode rather than a skip mode, another flag is signaled to indicate whether a single reference mode or a mixed reference mode is used for the current block, and the predictive block is generated by one motion vector in the single reference mode, whereas the predictive block is generated by a weighted average of two predictive blocks derived from two motion vectors in the mixed reference mode.

[0098] For example, in the case of a single reference, the following modes can be signaled: Use one of the motion vector predictors (MVPs) in the list pointed to by the NEARMV-DRL (Dynamic Reference List) index NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and apply the delta to the MVP. GLOBALMV - Use motion vectors based on frame-level global motion parameters

[0099] Also, for the mixed reference mode, the following modes can be signaled: NEAR_NEARMV-Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index. NEAR_NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and transmit the delta MV of the second MV. NEW_NEARMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and transmit the delta MV of the first MV. NEW_NEWMV - Use one of the motion vector predictors (MVPs) in the list signaled by the DRL index as reference and send the delta MV of both MVs. GLOBAL_GLOBALMV - Use MV from each reference based on frame-level global motion parameters.

[0100] And according to an exemplary embodiment, there may also be motion vector difference coding in AV1, where AV1 allows 1 / 8 pixel motion vector accuracy (or precision), and the following syntax is used to signal the motion vector difference of reference frame list 0 or list 1, i.e. mv_joint specifies which components of the motion vector difference are non-zero, i.e. 0 indicates that there is no nonzero MVD along either the horizontal or vertical directions, 1 indicates that there is nonzero MVD only along the horizontal direction, 2 indicates that there is non-zero MVD only along the vertical direction, 3 indicates the presence of non-zero MVD along both the horizontal and vertical directions, mv_sign specifies whether the motion vector difference is positive or negative. mv_class specifies the class of the motion vector difference (as shown in Table 3, a higher class means that the motion vector difference has a larger magnitude).

[0101] [Table 3]

[0102] mv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude for each MV class, mv_fr specifies the first two fractional bits of the motion vector difference, mv_hp specifies the third fractional bit of the motion vector difference.

[0103] Also, according to an example embodiment, CWG-B092 may have adaptive MVD resolution, where for NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD depends on the associated class and the magnitude of the MVD.

[0104] First, fractional MVD may only be allowed if the magnitude of the MVD is one pixel or less. Second, only one MVD value may be allowed when the value of the associated MV class is MV_CLASS_1 or greater, and the MVD value for each MV class is derived as 4, 8, 16, 32, 64 for MV class 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). The MVD values ​​allowed for each MV class are shown in Table 4.

[0105] [Table 4]

[0106] In addition, if the current block is coded as NEW_NEARMV or NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class, otherwise the other context is used to signal mv_joint or mv_class.

[0107] According to an example embodiment, there may also be a joint MVD coding (JMVD) in CWG-B092, in which a new inter-coding mode named JOINT_NEWMV may be applied to indicate whether the MVDs for two reference lists are signaled together. When the inter-prediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are signaled together. Therefore, only one MVD named joint_mvd may be signaled and sent to the decoder, and the delta MVs of reference list 0 and reference list 1 are derived from joint_mvd.

[0108] The JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. According to an example embodiment, no additional context is added.

[0109] Also, when JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distances. Specifically, the distance between reference frame list 0 and the current frame is denoted as td0, and the distance between reference frame list 1 and the current frame is denoted as td1. If td0 is greater than or equal to td1, joint_mvd is used as is for reference list 0, and the mvd of reference list 1 is derived from joint_mvd based on equation (1).

number

[0110] Otherwise, if td1 is greater than or equal to td0, then joint_mvd is used as is for reference list 1, and the mvd for reference list 0 is derived from joint_mvd according to equation (2).

number

[0111] According to an exemplary embodiment, there is also an improvement of adaptive MVD resolution in CWG-C011, in which a new inter-coding mode named AMVDMV may be added to the single reference case. When the AMVDMV mode is selected, the selection indicates that AMVD is applied to the signal MVD. One flag named amvd_flag may be added under the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. And when the adaptive MVD resolution is applied to the joint MVD coding mode, the MVDs of the two reference frames are signaled together, and the accuracy of the MVD is implicitly determined by the size of the MVD. Otherwise, the MVDs of the two (or more) reference frames are signaled together, and other MVD coding may be applied.

[0112] According to an example embodiment, adaptive motion vector resolution (AMVR) may be used in CWG-C012 and CWG-C020. AMVR in CWG-C012 supports a total of seven motion vector (MV) precisions (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8). For each prediction block, the adaptive AMVR encoder may search all supported precision values ​​and signal the best precision to the decoder. To reduce the encoder runtime, two precision sets are supported. Each precision set includes four predefined precisions. The precision set is adaptively selected at the frame level based on the maximum precision value of the frame. Similar to AV1, the maximum precision is signaled in the frame header. Table 5 summarizes the supported precision values ​​based on the frame-level maximum precision.

[0113] [Table 5]

[0114] In the current AMVR software (similar to AV1), there is a frame-level flag that indicates whether the MVs of a frame contain sub-pel precision. AMVR can only be enabled if the cur_frame_force_integer_mv flag has a value of 0. In AMVR, if the precision of a block is less than the maximum precision, the motion model and the interpolation filter are not signaled. If the precision of a block is less than the maximum precision, the motion mode is inferred to translational motion and the interpolation filter is inferred to REGULAR interpolation filter. Similarly, if the precision of a block is either 4 pixels or 8 pixels, the inter-intra mode is not signaled and is inferred to be 0.

[0115] Also, in JMVD, it is assumed that there is linear motion between the backward and forward reference frames, but additional improvements can be made because when the JMVD coding mode is selected for a block, one joint MVD is signaled for both reference frames, and the MVDs of the two reference frames are derived from the joint MVD based on the distance between the reference frame and the current frame. However, the motion between the two reference frames may not always be linear, for example, because the motion may be slower or faster from the backward reference frame to the forward reference frame.

[0116] According to an exemplary embodiment, the orientation of the reference frame is determined by whether the reference frame is before the current frame in display order or after the current frame in display order. Also, although the terms x-axis and y-axis refer to the horizontal and vertical components of the 2D value, these terms can be replaced by two other axes along two predefined directions perpendicular to each other, and the same embodiments described herein also apply. For example, the x-axis and y-axis can be replaced by 45 degree and 135 degree axes.

[0117] 13 shows a flowchart 1300 in which there is acquisition of a coding block of video data in S130 and in which it may be determined which mode is selected for that block in S131. For example, according to an example embodiment, if it is determined that the JMVD mode is selected for a block, flexible scaling factors may be used in S133 and the MVD of reference frame list 0 and / or 1 may be derived from the signaled joint MVD in S135 based on the signaled scaling factors in the bitstream or the coding information of the current block (or neighboring blocks). The flexible scaling factors may be applied to one or more components of the MVD along one or more predefined directions, either together or separately.

[0118] Also, according to an example embodiment, the predefined direction of the MVD refers to the MVD components along the x-axis and / or the y-axis.

[0119] According to an exemplary embodiment, when it is determined in S131 that JMVD mode is selected for a block, then in S135, the MVD of reference frame list 0 or 1 is derived from the signaled integrated MVD based on the distance between the reference frame and the current frame and / or the scaling factor of the JMVD mode.

[0120] According to an example embodiment, in S133, it may be determined that the signaling / derivation flag indicates that the scaling factor is not equal to a first predefined default value (e.g., 1). Then, in S135, this associated scaling factor of the signaled flag may be used to derive an MVD from the aggregate MVD of one of the reference frames. The scaling factor used to derive the MVD of the other reference frame is set to a second predefined default value (e.g., 1).

[0121] According to an example embodiment, when it is determined in S131 that the current block is coded as a JMVD mode, at least one flag may be signaled in the bitstream in S136 to indicate scaling factors for components of the MVD along one or more predefined directions, where the components of the MVD along one or more predefined directions refer to MVD components along the x-axis and / or y-axis.

[0122] Then, when it is determined in S131 that the current block is coded in a joint AMVD (or AMVR) coding mode, it may be determined in S133 and S134 that the scaling factors for both the x-axis and the y-axis are equal.

[0123] According to an example embodiment, once it is determined that the current block is coded as a JMVD mode, a flag such as jmvd_scale_factor_flag may be signaled in S136 to indicate the scaling factor for deriving the MVD. Also, if jmvd_scale_factor_flag indicates that the value of the scaling factor is not equal to 1, another flag named scale_factor_dir may be signaled to indicate whether the scaling factor is applied to the x-axis, or the y-axis, or both the x-axis and the y-axis of the MVD.

[0124] According to an exemplary embodiment, when it is determined that the scaling factor is to be applied to both the x-axis and the y-axis, the scaling factors for both the x-axis and the y-axis are equal in S133 and S134.

[0125] According to an example embodiment, once it is determined in S131 that the current block is coded as JMVD mode, a flag such as scale_factor_dir may be signaled in S136 to indicate whether the scaling factor is applied to the x-axis, or the y-axis, or both the x-axis and the y-axis of the MVD. Then, another flag named jmvd_scale_factor_flag is signaled to indicate the value of the scaling factor for deriving the MVD.

[0126] According to an example embodiment, when it is determined in S131 that the current block is coded as JMVD mode, a flag such as jmvd_scale_factor_flag may be signaled to indicate an index in a scaling factor lookup table for deriving MVD. Each entry in this scaling factor lookup table then specifies a value of a scaling factor for the x-axis or y-axis. The order of the entries in the lookup table may be fixed and predefined according to statistics or other rules, or according to a decoder-side search algorithm such as template matching or bilateral matching (i.e., template matching-based / two-sided matching-based reordering). An example of a scaling factor lookup table is shown in Table 6.

[0127] [Table 6]

[0128] According to an example embodiment, when it is determined in S131 that the current block is coded as a JMVD mode, a flag such as jmvd_scale_factor_equal_to_one may be coded to indicate the scaling factor and direction of the combined MVD. That is, if the flag jvmd_scale_factor_equal_to_one is equal to one value (e.g., 1), the scaling factor is set equal to 1 in S133, and this factor is applied to both the horizontal (x-axis) and vertical (y-axis) in S134. Otherwise, if the flag jmvd_scale_factor_equal_to_one is equal to the other value (e.g., 0), an additional jmvd_scaling_dir syntax element is signaled. Based on a lookup table (e.g., Table 7), the scaling factors of different directions (axes) can be obtained. It should be noted that the order of entries in the lookup table can be fixed and predefined according to statistics or other rules, or according to a decoder-side search algorithm such as template matching or bilateral matching (i.e., template matching-based / two-sided matching-based reordering).

[0129] [Table 7]

[0130] According to an exemplary embodiment, the context for signaling jmvd_scale_factor_flag and / or scale_factor_dir depends on the block size of the current block, or the coding information of the current block or neighboring blocks, such as the MVP of the current block, or whether the coding mode of the current block is a joint AMVD (or AMVR) coding mode, or the MVD of the neighboring blocks, or the coding mode of the neighboring blocks. Also, the value of the scaling factor may be limited to n powers of 2, where n may be 0 or a positive integer or a negative integer. For example, the value of the scaling factor may be limited to {1, 2}, or the value of the scaling factor may be limited to {1, 1 / 2, 2}.

[0131] According to an example embodiment, the value of the scaling factor may be m / M, where M is n squared and m is an integer, and the values ​​of the scaling factor are restricted to {1 / 8, 2 / 8, 3 / 8, 4 / 8, ..., 15 / 8.16 / 8}.

[0132] According to an exemplary embodiment, the MVD scaling process is the same as the method described in U.S. Patent No. 63 / 328,062, filed April 6, 2022, the entirety of which is incorporated herein by reference.

[0133] According to an example embodiment, the scaling factor may be derived in S133 based on coding information including but not limited to the quantization step size or quantization parameter, block size, difference / relationship between MVP0 and MVP1, motion vector prediction block of the current block, MVD class, reference picture, MVD scaling factor of neighboring blocks. Also, one syntax may be signaled in S136 in high level syntax to indicate whether jmvd_scale_factor_flag and / or scale_factor_dir need to be signaled in the bitstream. The high level syntax includes but is not limited to sequence header, frame header, and slice header.

[0134] And if no particular mode is determined in S131, then any other processing described above may occur with other modes in S132, and additional processing may occur in S137 as well as any processing mode described herein. Additionally, obtaining a coding block in S130 after reaching processing S137 may depend in an iterative manner on previous processing of a previous coding block.

[0135] In the above description of FIG. 13, it should be understood that one or more of the operations may occur in a different order than that shown and / or in parallel with each other and other operations.

[0136] The techniques described above can be implemented using computer readable instructions, as computer software physically stored on one or more computer readable media, or by one or more tangibly configured hardware processors. For example, FIG. 14 illustrates a computer system 1400 suitable for implementing certain embodiments of the disclosed subject matter.

[0137] The computer software can be coded using any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., directly, or via interpretation, microcode execution, etc.

[0138] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0139] 14 with respect to computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. Neither the arrangement of components should be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of computer system 1400.

[0140] The computer system 1400 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, via tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (speech, music, ambient sounds, etc.), images (scanned images, photographic images obtained from still image cameras, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0141] The input human interface devices may include one or more of a keyboard 1401, a mouse 1402, a trackpad 1403, a touch screen 1410, a joystick 1405, a microphone 1406, a scanner 1408, and a camera 1407 (only one of each is shown).

[0142] The computer system 1400 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the senses of a human user, for example, through haptic output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen 1410, or a joystick 1405, although there may be haptic feedback devices that do not function as input devices), audio output devices (such as speakers 1209, headphones (not shown)), visual output devices (such as screens 1410, including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions via means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0143] The computer system 1400 may also include human accessible storage devices and their associated media, such as CD / DVD 1411 or CD / DVD ROM / RW 1420 with similar media, thumb drives 1422, removable hard drives or solid state drives 1423, legacy magnetic media such as tapes and floppy disks (not shown), optical media including dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), and the like.

[0144] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0145] The computer system 1400 may also include an interface 1499 to one or more communication networks 1498. The network 1498 may be, for example, wireless, wired, optical. Additionally, the network 1498 may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks 1498 include local area networks such as Ethernet, cellular networks including WLAN, GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable television, satellite television and terrestrial broadcast television, vehicular and industrial including CANBus, etc. Particular networks 1498 typically require an external network interface adapter coupled to a particular general purpose data port or peripheral bus (1450 and 1451) (e.g., USB port of the computer system 1400, etc.), while others are typically integrated into the core of the computer system 1400 by coupling to a system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 1498, computer system 1400 can communicate with other entities. Such communications can be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to a particular CANbus device), or two-way, such as to other computer systems using local area or wide area digital networks. Specific protocols and protocol stacks can be used with each of these networks and network interfaces, as described above.

[0146] The aforementioned human interface devices, human accessible storage devices, and network interfaces can be connected to the core 1440 of the computer system 1400 .

[0147] The core 1440 may include one or more central processing units (CPUs) 1441, graphics processing units (GPUs) 1442, graphics adapters 1417, dedicated programmable processing units in the form of field programmable gate areas (FPGAs) 1443, hardware accelerators 1444 for specific tasks, etc. These devices may be connected via a system bus 1448, along with read only memory (ROM) 1445, random access memory 1446, internal mass storage devices 1447 such as internal hard drives, SSDs, etc. that are not accessible to the user. In some computer systems, the system bus 1448 may also be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripherals may be coupled directly to the core's system bus 1448 or via a peripheral bus 1451. Architectures for peripheral buses include PCI, USB, etc.

[0148] The CPU 1441, GPU 1442, FPGA 1443, and accelerator 1444 can execute certain instructions that can combine to constitute the aforementioned computer code. That computer code can be stored in ROM 1445 or RAM 1446. Persistent data can be stored in, for example, internal mass storage device 1447, while transitory data can also be stored in RAM 1446. Fast storage and fast retrieval to any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU 1441, GPU 1442, mass storage device 1447, ROM 1445, RAM 1446, etc.

[0149] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0150] As an example, and not by way of limitation, the architecture, and in particular the computer system 1400 having the core 1440, can provide functionality as a result of a processor (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer readable media. Such computer readable media can be mass storage accessible to a user as introduced above, as well as media associated with a particular storage of the core 1440 of a non-transitory nature, such as the core internal mass storage 1447 or the ROM 1445. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core 1440. The computer readable media can include one or more memory devices or chips, depending on the particular need. The software can cause the core 1440, and in particular the processor therein (including CPU, GPU, FPGA, etc.) to perform certain processes or certain parts of certain processes described herein, including defining data structures stored in the RAM 1446 and modifying such data structures according to the processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 1444) that may operate in place of or together with software to perform certain processes or certain portions of certain processes described herein. References to software may encompass logic, and vice versa, as appropriate. References to computer-readable media may encompass circuitry (such as integrated circuits (ICs)) that store software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0151] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that are within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope. [Explanation of symbols]

[0152] 100 Communication Systems 101 Terminal 102 Terminals 103 Terminal 104 Terminals 105 Network 201 Video Sources 202 Encoder 203 Capture Subsystem 204 encoded video bitstream 205 Streaming Server 206 Copy 207 Streaming Client 208 Copy 209 Display 210 outgoing video sample streams that can be rendered 211 Video Decoder 212 Streaming Client 213 Sample Stream 300 Decoder 301 Channel 302 Receiver 303 Buffer Memory 304 Entropy Decoder / Parser 305 Scaler / Descaler Unit 306 Motion Compensation Prediction Unit 307 Intra Prediction Units 308 Reference Picture Memory 309 Current (partially reconstructed) picture 310 Aggregator 311 Loop Filter Unit 312 Display 313 Symbols 400 Encoder 401 Video Source 402 Controller 403 Source Coder 404 Predictor 405 Reference Picture Memory 406 Local Video Decoder 407 Coding Engine 408 Entropy Coder 409 Transmitter 411 Communication Channels 601 Segment A 602 Segment B 603 Segment C 604 Segment D 605 Segment E 606 Segment F 607 The Fourth Reference Hierarchy 608 Third Reference Hierarchy 609 Second Reference Hierarchy 610 First Reference Hierarchy 700 DC mode PDPC weighting diagram 800 Local Lighting Compensation (LIC) Diagram 900 Intra Prediction Modes Used in HEVC An example of block division using 1000 QTBT 1001 Corresponding Tree Representation 1002 Contained Multi-Type Tree (MTT) Structures 1100 An example of block division in VP9 and AVI 1103 Block (Quarter) 1104 Block (Bisection, JEM) 1105 Block (Triple) 1111 An example coding tree unit (CTU) for VP9 1112 64x64 level 1113 4×4 Level 1115 128x128 level 1116 10-way structure 1200 An example related to merge mode with motion vector difference 1201 See L0 1202 See L1 1300 Flowchart 1400 Computer Systems 1401 Keyboard 1402 Mouse 1403 Trackpad 1405 Joystick 1406 Mike 1407 Camera 1408 Scanner 1410 Touch Screen 1411 CD / DVD or similar media 1417 Graphics Adapter 1420 CD / DVD ROM / RW 1422 Thumb Drive 1423 Removable Hard Drive or Solid State Drive 1440 cores 1441 One or more Central Processing Units (CPUs) 1442 Graphics Processing Unit (GPU) 1443 Field Programmable Gate Area (FPGA) 1444 Hardware accelerators for specific tasks 1445 Read-Only Memory (ROM) 1446 Random Access Memory 1447 Internal hard drives, SSDs and other internal mass storage devices 1448 System Bus 1450 Surrounding Bus 1451 Surrounding Bus 1498 Communication Network 1499 Network Interface

Claims

1. 1. A method for video coding performed by at least one processor, the method comprising: obtaining a coding block of a video bitstream; determining whether joint motion vector difference coding (JMVD) is used to predict the coding block; obtaining a plurality of scaling coefficients and a joint motion vector difference (joint MVD) from the video bitstream based on the determination that the joint MVD is used to predict the coding block; deriving a motion vector difference (MVD) of at least one reference frame list based on application of the plurality of scaling factors to one or more components of the integrated MVD along one or more predefined directions; and reconstructing the coding block based on at least the derived MVD.

2. The MVD is derived further based on any one of a distance between a reference frame and a current frame.

2. A method for video coding according to claim 1.

3. deriving the MVD includes determining whether a flag indicates that at least one of the scaling factors is not equal to a first predefined default value; the at least one of the scaling factors is used to derive an MVD from the joint MVD for one of the reference frames; another one of the scaling factors used to derive an MVD from the integrated MVD for a second one of the reference frames is set to a second predefined default value; 3. A method for video coding according to claim 2.

4. the step of obtaining the scaling factor is based on obtaining at least one flag signaled in the video bitstream; the at least one flag indicates the scaling factor for at least one of the components along the one or more predefined directions.

2. A method for video coding according to claim 1.

5. The step of deriving the MVD comprises: determining whether a first flag indicates that at least one of the scaling factors is not equal to a first predefined default value; and based on determining that the first flag indicates that the at least one of the scaling factors is not equal to the first predefined default value, determining whether the scaling factor is to be applied to at least one direction of the MVD and at least one direction of the predefined directions based on a value of a second flag.

2. A method for video coding according to claim 1.

6. deriving the MVD includes applying the one or more scaling factors equally to both of the predefined directions based on determining that the second flag indicates both of the predefined directions; A method for video coding according to claim 5.

7. obtaining the scaling factor includes obtaining an index of the scaling factor in a lookup table; the lookup table indicates that at least one pair in a first one of the indices of the scaling factors has the same scaling factor value in both of the predefined directions; the lookup table indicates that at least a second pair in a second one of the indices of the scaling factors has a different scaling factor value between one of the predefined directions; 2. A method for video coding according to claim 1.

8. at least one of the same scaling factor value and the different scaling factor value is a fractional scaling factor value; at least one other of the same scaling factor value and the different scaling factor value is m / M, where M is a power of 2 and m and n are integers; A method for video coding according to claim 7.

9. The scaling factor is derived based on at least one coding information among a quantization step size, a quantization parameter, a block size, a difference between a motion vector prediction block of a current block, an MVD class, a reference picture, and an MVD scaling factor of a neighboring block in the vicinity of the coding block.

2. A method for video coding according to claim 1.

10. At least one of a frame header, a slice header, and a sequence header indicates whether to signal the scaling factor of the video bitstream.

2. A method for video coding according to claim 1.

11. An apparatus for video coding configured to perform a method according to any one of claims 1 to 10.

12. A computer program which, when executed by a computer, causes the computer to perform the method according to any one of claims 1 to 10.