Video coding and decoding

By optimizing the use of affine motion modes based on inter prediction and adjacent block analysis, the method improves video encoding efficiency and reduces complexity, addressing the challenges of complex motion patterns in existing standards.

JP2025111774AActive Publication Date: 2025-07-30CANON KK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025076374
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-09-21
Filing Date
2025-05-01
Publication Date
2025-07-30
Estimated Expiration
2039-09-18

AI Technical Summary

Technical Problem

The existing video coding standards, such as HEVC, face challenges in efficiently handling complex motion patterns like zoom, rotation, and perspective motions due to the increased complexity and signal overhead of affine motion modes.

Method used

The method involves determining the inter prediction mode and enabling affine motion mode based on skip flags, adjacent block usage, and context encoding flags, optimizing the transmission of affine motion modes to improve coding efficiency and reduce complexity.

Benefits of technology

This approach enhances encoding efficiency and reduces complexity by strategically employing affine motion modes, leading to more efficient and faster video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111774000001_ABST
    Figure 2025111774000001_ABST
Patent Text Reader

Abstract

To provide a coding method, decoding method, and device for improving the coding efficiency of the affine mode.SOLUTION: The decoding method includes transmitting the affine mode in an encoded video stream, specifically determining a merge candidate list corresponding to adjacent blocks A1, B1, B0, A0, and B2 of the current block, encoding a context-coded flag from the data stream, meaning transmitting the affine mode of the current block, and determining the context variable of the flag on the basis of whether the adjacent blocks use the affine mode.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video encoding and decoding.

Background Art

[0002] Recently, the Joint Video Experts Team (JVET), which is a joint team formed by MPEG and VCEG of ITU-T Study Group 16, started research on a new video coding standard called Versatile Video Coding (VVC). The goal of VVC is to provide a significant improvement in compression performance beyond the existing HEVC standard (i.e., typically twice that of the previous one) and to be completed by 2020. The main target applications and services include, but are not limited to, 360-degree and high dynamic range (HDR) video. Overall, JVET evaluated responses from 32 groups using formal subjective tests conducted by independent test houses. Some proposals demonstrated a compression efficiency gain of typically over 40% compared to using HEVC. Specific effectiveness was shown for test materials of ultra-high definition (UHD) video. Therefore, we can expect a compression efficiency gain far exceeding 50% which is the goal of the final standard.

[0003] The JVET Exploration Model (JEM) uses all HEVC tools. An additional tool not present in HEVC is to use the "affine motion mode" when applying motion compensation. Motion compensation in HEVC is limited to translation, but in reality, there are many types of motions such as, for example, zoom in / out, rotation, perspective motion, and other irregular motions. When using the affine motion mode, more complex transforms are applied to the block in order to more accurately predict such motion patterns. However, the use of the affine motion mode may increase the complexity of the encoding / decoding process and may also increase the signal overhead. Therefore, at least one solution to the aforementioned problems is desirable.

Summary of the Invention

[0004] In a first aspect of the present invention, a method for transmitting a motion prediction mode of a part of a bitstream, the method comprising: determining an inter prediction mode used for the part of the bitstream; and transmitting an affine motion mode depending on the inter prediction mode used for the part of the bitstream. Optionally, the inter prediction mode used is determined based on a state of a skip flag in the part of the bitstream. Optionally, the affine mode is not enabled when the skip flag is present. Optionally, the method further comprises enabling a merge mode when the affine mode is enabled. Optionally, the affine mode is enabled when the inter prediction mode is an Advanced Motion Vector Predictor (AMVP). Optionally, the determination is performed based on a high-level syntax flag, and the high-level syntax flag indicates at least one of processing at a slice level, a frame level, a sequence level, and a Coding Tree Unit (CTU) level. Optionally, determining the inter prediction mode includes determining a mode of one or more blocks adjacent to a current block.

[0005] In a second aspect of the present invention, a method for transmitting a motion prediction mode within a bitstream is provided, which includes determining the mode of one or more adjacent blocks for a current block and transmitting an affine motion mode of the current block depending on the mode. Optionally, the adjacent blocks are composed of only block A1 and B1. Alternatively, the adjacent blocks include block A2 and B3, preferably composed of only block A2 and B3. Optionally, the method includes enabling the affine motion mode when one or both of the adjacent blocks use the affine motion mode. Optionally, the adjacent blocks further include B0, A0, and B2. Optionally, the use of the affine mode in the adjacent blocks is continuously determined, and when one of the adjacent blocks uses the affine mode, the affine mode is enabled for the current block. Preferably, a series of adjacent blocks are A2, B3, B0, A0, B2.

[0006] In a third aspect of the present invention, a method for transmitting a motion prediction mode of a part of a bitstream is provided, which includes determining a list of merge candidates corresponding to blocks adjacent to a current block and enabling an affine mode of the current block when one or more of the merge candidates use the affine mode. Optionally, the list starts with a block used to determine a context variable related to the block. Optionally, the list starts with block A2 and B3 in its order. Optionally, the list is A2, B3, B0, or A0, or B2 in its order. Optionally, the affine mode is enabled for the current block when the adjacent blocks do not use the merge mode. Optionally, the affine mode is enabled for the current block when the adjacent blocks do not use the merge skip mode. Optionally, transmitting the affine mode includes inserting a context coding flag into the data stream, and the context variable of the flag is determined based on whether the adjacent blocks use the affine mode.

[0007] In a further aspect of the present invention, a method for transmitting a motion prediction mode of an encoded block in a bitstream, comprising: determining whether a block adjacent to the encoded block in the bitstream uses an affine mode; and inserting a context encoding flag into the bitstream, wherein a context variable of the context encoding flag depends on the determination of whether a block adjacent to the encoded block in the bitstream uses an affine mode. Optionally, the adjacent blocks include blocks A1 and B1. Optionally, when the mode of a block for which the motion prediction mode is enabled is a merge mode, the adjacent blocks include blocks A1 and B1. Optionally, the context of the affine flag is obtained according to the following formula Ctx = IsAffine(A1)+ IsAffine(B1), where Ctx is a context variable of the affine flag, and IsAffine is a function that returns 0 when the block is not an affine block and 1 when the block is affine.

[0008] In a fourth aspect of the present invention, a method for transmitting a motion prediction mode of an encoded block in a bitstream is provided depending on whether adjacent blocks use a merge mode and / or a merge skip mode. In a fifth aspect of the present invention, a method for transmitting a motion prediction mode in a bitstream is provided, which includes compiling a list of motion predictor candidates and inserting an affine merge mode as a merge candidate. Optionally, the affine merge mode candidate is after the adjacent block motion vectors in the list of merge candidates. Optionally, the affine merge mode candidate is before the alternative temporal motion vector predictor (ATMVP) candidates in the list of merge candidates. Optionally, the position (merge index) of the affine merge mode candidate in the list of candidates is fixed. Optionally, the position of the affine merge mode candidate in the candidate list is variable. Optionally, the position of the affine merge mode candidate is determined based on one or more of a) the state of the skip flag, b) the motion information of adjacent blocks, c) the alternative temporal motion vector predictor (ATMVP) candidates, and d) whether adjacent blocks use an affine mode.

[0009] Optionally, if one or more of the following conditions are met, i.e., a) a skip flag exists, b) the motion information of adjacent blocks is equal, c) the ATMVP candidate contains only one motion information, d) one or more adjacent blocks use the affine mode, the affine merge mode is placed lower in the list of candidates (to which a higher merge index is assigned). Optionally, the adjacent blocks include block A1 and B1. Optionally, the affine merge mode is placed lower in the list of candidates (to which a higher merge index is assigned) than the spatial motion vector candidate if one or more of the above conditions a) to d) are met. Optionally, the affine merge mode is placed lower (to which a higher merge index is assigned) than the temporal motion vector candidate if one or more of the above conditions a) to d) are met. Optionally, the affine merge mode is assigned a merge index related to the number of adjacent blocks using the affine mode. Optionally, the affine merge mode is assigned a merge index equal to five minus the amount of adjacent blocks using the affine mode from among five of A1, B1, B0, A0, B2.

[0010] According to another aspect of the present invention, there is provided a method for transmitting an affine motion mode in a video stream, the method comprising determining whether the likelihood of an affine mode is being used for a current block, compiling a motion candidate predictor list, and inserting an affine merge mode as a merge candidate depending on determining the likelihood of the affine mode of the current block. Optionally, the likelihood is determined based on at least one of a) the state of a skip flag, b) motion information of adjacent blocks, and c) an ATMVP candidate. Optionally, the affine merge mode is not inserted as a merge candidate if one or more of the following conditions are met, namely a) the state of the skip flag, b) the motion information of adjacent blocks is equal, and c) the ATMVP candidate includes only one motion information. Optionally, the adjacent blocks include block A1 and B1. Optionally, the affine mode is transmitted depending on the characteristics of a device used to record a video corresponding to an encoded bitstream.

[0011] Aspects of the present invention provide an improvement in encoding efficiency and / or a reduction in encoding complexity compared to existing encoding standards or proposals. In this way, more efficient and faster video encoding and / or decoding methods and systems are provided. Further aspects of the present invention relate to encoding and decoding methods using any of the above aspects. Yet another aspect of the present invention relates to an apparatus for transmitting the use of an affine mode in a bitstream representing an encoded video as defined by claim 14. Yet another aspect of the present invention relates to an encoding unit and a decoding unit as defined by claims 17 and 18 respectively. Yet another aspect of the present invention relates to a program as defined by claim 19. The program may be provided by itself or may be carried on, by, or in a carrier medium. The carrier medium may be non-transitory, for example, a storage medium, particularly a computer-readable storage medium. The carrier medium may also be transitory, for example, a signal or another transmission medium. The transmission may be via any suitable network including the Internet.

[0012] Furthermore, a further aspect of the present invention relates to a peripheral device such as a camera or a mobile device as defined by claims 15 and 16. Optionally, the camera further includes zoom means adapted to indicate that the zoom means is operating when the zoom means is operating and in a signal affinity mode that depends on the instruction. Optionally, the camera further includes pan means adapted to indicate that the pan means is operating when the pan means is operating and in a signal affinity mode that depends on the instruction. Optionally, the mobile device further includes at least one position sensor adapted to sense a change in the orientation of the mobile device and adapted to communicate an affinity mode depending on sensing the change in the orientation of the mobile device. Further features of the present invention are characterized by other independent and dependent claims.

[0013] Any feature in one aspect of the present invention may be applied, in any suitable combination, to other aspects of the present invention. In particular, method aspects may be applied to apparatus aspects, and vice versa. Further, features implemented in hardware may be implemented in software, and vice versa. Any reference to software and hardware features herein should be construed accordingly. Any apparatus feature as described herein may be provided as a method feature, and vice versa. As used herein, means-plus-function features may alternatively be expressed in terms of their corresponding structures, such as a properly programmed processor and associated memory. Also, it should be understood that certain combinations of the various features described and defined in any aspect of the present invention may be implemented and / or supplied and / or used independently.

Brief Description of the Drawings

[0014] Here, by way of example, the following attached drawings are referred to for explanation.

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6a

Figure 6b

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Embodiments for Carrying Out the Invention

[0016] The present invention relates to improved transmission of affine motion modes, in particular to determining when an affine mode may result in an improvement in coding efficiency, and ensuring that the affine mode is used and / or priority is assigned accordingly. FIG. 1 relates to the coding structure used in the High Efficiency Video Coding (HEVC) video standard. A video sequence 1 consists of a series of digital images i. Each such digital image is represented by one or more matrices. The matrix coefficients represent pixels. The images 2 of the sequence can be divided into slices 3. In some examples, a slice can constitute the entire image. These slices are divided into non-overlapping Coding Tree Units (CTUs). A Coding Tree Unit (CTU) is the basic processing unit of the High Efficiency Video Coding (HEVC) video standard and conceptually corresponds to the macroblock units used in some previous video standards. A CTU is sometimes also called a Largest Coding Unit (LCU). A CTU has luminance and chrominance component parts, and each of the component parts is called a Coding Tree Block (CTB). These different color components are not shown in FIG. 1.

[0017] A CTU is generally of size 64 pixels x 64 pixels for HEVC and may be of size 128 pixels x 128 pixels for VVC. Each CTU may be iteratively and sequentially divided into smaller variable-size Coding Units (CUs) 5 using quadtree decomposition. A coding unit is the basic coding element and is composed of two types of subunits called a Prediction Unit (PU) and a Transform Unit (TU). The maximum size of a PU or TU is equal to the CU size. A prediction unit corresponds to the partitioning of a CU for prediction of pixel values. As shown by 606, it includes partitioning into four square PUs and two different partitionings into two rectangular PUs, and various different partitionings of a CU into PUs are possible. A transform unit is the basic unit subject to spatial transformation using DCT. A CU can be divided into TUs based on the quadtree representation 607.

[0018] Each slice is embedded in one network abstraction layer (NAL) unit. Further, the encoding parameters of the video sequence are stored in a dedicated NAL unit called a parameter set. In HEVC and H.264 / AVC, two types of parameter set NAL units are used, namely, first, a sequence parameter set (SPS) NAL unit that collects all parameters that do not change during the entire video sequence. Typically, it processes the encoding profile, the size of the video frame, and other parameters. Second, a picture parameter set (PPS) NAL unit contains parameters that can change from one picture (or frame) of the sequence to another. HEVC also includes a video parameter set (VPS) NAL unit that contains parameters describing the overall structure of the bitstream. The VPS is a new type of parameter set defined in HEVC and is applied to all layers of the bitstream. A layer may include multiple temporal sublayers, and all version 1 bitstreams are restricted to a single layer. HEVC has specific hierarchical extensions for scalability and multi-view, which enable multiple layers with a backward-compatible version 1 base layer.

[0019] FIG. 2 shows a data communication system in which one or more embodiments of the present invention are implemented. The data communication system includes a transmitting device, in this case a server 201, operable to transmit data packets of a data stream via a data communication network 200 to a receiving device, in this case a client terminal 202. The data communication network 200 may be a wide area network (WAN) or a local area network (LAN). Such a network may be, for example, a wireless network (Wifi / 802.11a or b or g), an Ethernet network, an Internet network, or a hybrid network composed of a plurality of different networks. In a particular embodiment of the present invention, the data communication system may be a digital television broadcast system in which the server 201 transmits the same data content to a plurality of clients. The data stream 204 provided by the server 201 may be composed of multimedia data representing video and audio data. The audio and video data streams may, in some embodiments of the present invention, be captured by the server 201 using a microphone and a camera, respectively. In some embodiments, the data stream may be stored on the server 201, or received by the server 201 from another data provider, or generated by the server 201. The server 201 particularly includes an encoding unit for encoding video and audio streams in order to provide a compressed bit stream for transmission, which is a more compact representation of the data presented as input to the encoding unit.

[0020] To obtain a better ratio of the quality of transmitted data to the amount of transmitted data, the compression of video data may follow, for example, the HEVC format or the H.264 / AVC format. The client 202 receives the transmitted bitstream and decodes the reconstructed bitstream in order to reproduce the video image on the display device and the audio data by the speaker. In the example of FIG. 2, a streaming scenario is considered, but it will be understood that in some embodiments of the present invention, the data communication between the encoding unit and the decoding unit may be performed using a medium storage device such as an optical disk. In one or more embodiments of the present invention, the video image is transmitted together with data representing a compensation offset in order to provide filtered pixels in the final image for application to the reconstructed pixels of the image.

[0021] FIG. 3 schematically shows a processing device 300 configured to implement at least one embodiment of the present invention. The processing device 300 may be a device such as a microcomputer, a workstation, or a light portable device. The device 300 includes a communication bus 313 connected thereto, that is, a central processing unit 311 such as a microprocessor represented by a CPU, a read-only memory 306 denoted as ROM for storing a computer program for implementing the present invention, a random access memory 312 denoted as RAM for storing the executable code of the method of the embodiment of the present invention and for storing registers adapted to record variables and parameters necessary for implementing a method of encoding a sequence of digital images and / or a method of decoding a bitstream according to an embodiment of the present invention, and a communication interface 302 connected to a communication network 303 through which the digital data to be processed is transmitted or received.

[0022] Optionally, the apparatus 300 can also include the following components, namely, a computer program for implementing the method of one or more embodiments of the present invention, and data storage means 304 such as a hard disk for storing data used or generated during the implementation of one or more embodiments of the present invention, a disk drive 305 for a disk 306, the disk drive being adapted to read data from or write data to the disk 306, and a screen 309 that functions as a graphical interface with the user for displaying data and / or by means of a keyboard 310 or any other indicating means.

[0023] The apparatus 300 may be connected to various peripheral devices such as, for example, a digital camera 320 or a microphone 308, each of which is connected to an input / output card (not shown) for supplying multimedia data to the apparatus 300. The communication bus provides communication and interconnectivity between the various elements included in or connected to the apparatus 300. The representation of the bus is not limited, and in particular, the central processing unit is operable to communicate instructions to any element of the apparatus 300 either directly or by means of another element of the apparatus 300.

[0024] The disk 306 can be replaced by any information medium such as, for example, a compact disk (CD-ROM), rewritable or not, ZIP disk or memory card, and generally is readable by a microcomputer or microprocessor, integrated with the apparatus or not, removable if possible, and configured to store one or more programs enabling a method of encoding a sequence of digital images and / or a method of decoding a bitstream implemented according to the invention. The executable code may be stored in any of the read-only memory 306, hard disk 304, or removable digital media such as disk 306 as described above. According to a variant, the executable code of the program may be received by means of the communication network 303 via the interface 302 in order to be stored in one of the storage means of the apparatus 300 before being executed, such as the hard disk 304.

[0025] The central processing unit 311 is adapted to control and direct the execution of a part of the software code of an instruction or program or the program described in the present invention, an instruction stored in one of the aforementioned storage means. At power-on, a program or programs stored in non-volatile memory, for example on the hard disk 304 or read-only memory 306, are transferred to the random access memory 312, which then contains the program or programs' executable code and registers for storing variables and parameters necessary for implementing the present invention. In this embodiment, the apparatus is a programmable apparatus using software for implementing the present invention. However, alternatively, the present invention may be implemented in hardware (for example, in the form of an application specific integrated circuit or ASIC).

[0026] Figure 4 shows a block diagram of an encoding unit according to at least one embodiment of the present invention. The encoding unit is represented by connected modules, and each module implements at least one corresponding step of a method for encoding an image of an image sequence according to one or more embodiments of the present invention, in the form of program instructions to be executed, for example, by the CPU 311 of the device 300. The original sequence of digital images i0 to in401 is received as input by the encoding unit 400. Each digital image is represented by a set of samples known as pixels. The bitstream 410 is the output by the encoding unit 400 after the implementation of the encoding process. The bitstream 410 comprises a plurality of encoding units or slices, and each slice comprises a slice header for transmitting the encoded video data and the encoded value of the encoding parameters used for encoding the slice body.

[0027] The input digital images i0 to in401 are divided into pixel blocks by the module 402. The blocks correspond to image portions and may be of variable size (e.g., 4×4, 8×8, 16×16, 32×32, 64×64, 128×128 pixels, and some rectangular block sizes may be considered). An encoding mode is selected for each input block. Two families of encoding modes are provided, namely, an encoding mode based on spatial prediction encoding (intra prediction) and an encoding mode based on temporal prediction (inter encoding, merge skip). The possible encoding modes are tested. The module 403 performs an intra prediction process in which a given block to be encoded is predicted by a predictor calculated from pixels in the vicinity of the block to be encoded. The indication of the selected intra predictor and the difference between the given block and its predictor are encoded to provide a residual if intra encoding is selected.

[0028] Temporal prediction is performed by the motion estimation module 404 and the motion compensation module 405. First, a reference image is selected from the set of reference images 416, and a portion of the reference image, also called a reference region or image portion, which is the region closest to a given block to be encoded, is selected by the motion estimation module 404. Next, the motion compensation module 405 uses the selected region to predict the block to be encoded. The difference between the selected reference region and the given block, also called the residual block, is calculated by the motion compensation module 405. The selected reference region is indicated by a motion vector. Thus, in both cases (spatial and temporal prediction), the residual is calculated by subtracting the prediction from the original block. In intra prediction performed by module 403, the prediction direction is encoded. In temporal prediction, at least one motion vector is encoded. When inter prediction is selected, information for the motion vector and the residual block is encoded. To further reduce the bitrate, assuming that the motions are homogeneous, the motion vectors are encoded by the difference with respect to the motion vector predictor. The motion vector predictors of the set of motion information predictors are obtained from the motion vector field 418 by the motion vector prediction and encoding module 417.

[0029] The symbolization unit 400 further includes a selection module 406 for encoding mode selection by applying an encoding cost criterion such as a rate distortion criterion. To further reduce redundancy, a transformation (such as DCT) is applied to the residual block by the transformation module 407, the obtained transformed data is quantized by the quantization module 408, and entropy encoded by the entropy encoding module 409. Finally, the encoded residual block of the currently encoded block is inserted into the bitstream 410. Also, the symbolization unit 400 decodes the encoded image to generate a reference image for motion estimation of subsequent images. This enables the symbolization unit and the multiplexing unit that receive the bitstream to have the same reference frame. The inverse quantization module 411 performs inverse quantization of the quantized data, followed by inverse transformation by the inverse transformation module 412. The backward intra prediction module 413 uses prediction information to determine which predictor to use for a given block, and the backward motion compensation module 414 actually adds the residual obtained by the module 412 to the reference region obtained from the set of reference images 416.

[0030] Next, post-filtering is applied by module 415 to filter the frame of the reconstructed pixels. In an embodiment of the present invention, an SAO loop filter is used in which a compensation offset is added to the pixel value of the reconstructed pixels of the reconstructed image. FIG. 5 shows a block diagram of a decoding unit 60 that can be used to receive data from an encoding unit according to an embodiment of the present invention. The decoding unit is represented by connected modules, and each module is adapted to perform the corresponding steps of the method implemented by the decoding unit 60, for example, in the form of program instructions executed by the CPU 311 of the device 300. The decoding unit 60 receives a bitstream 61 including an encoding unit, each of which is composed of a header including information regarding encoding parameters and a body including encoded video data. As described with respect to FIG. 4, the encoded video data is entropy encoded, and the index of the motion vector predictor is encoded with a predetermined number of bits for a predetermined block. The received encoded video data is entropy decoded by module 62. Next, the residual data is inverse quantized by module 63, and then an inverse transform is applied by module 64 to obtain pixel values.

[0031] Mode data indicating the encoding mode is also entropy decoded, and based on the mode, intra-type decoding or inter-type decoding is performed on the encoded block of the image data. In the case of the intra mode, the intra predictor is determined by the intra inverse prediction module 65 based on the intra prediction mode specified in the bitstream. When the mode is inter, motion prediction information is extracted from the bitstream to find the reference region used by the encoding unit. The motion prediction information is composed of a reference frame index and a motion vector residual. A motion vector predictor is added to the motion vector residual to obtain a motion vector by the motion vector decoding module 70.

[0032] The motion vector decoding module 70 applies motion vector decoding to each current block encoded by motion prediction. For the current block, once the actual value of the motion vector related to the current block that can be decoded is obtained, the index of the motion vector predictor is used by module 66 to apply backward motion compensation. The reference image portion indicated by the decoded motion vector is extracted from the reference image 68 to apply backward motion compensation 66. The motion vector field data 71 is updated with the decoded motion vectors for use in inverse prediction of subsequent decoded motion vectors. Finally, the decoded block is obtained. Post-filtering is applied by the post-filtering module 67. The decoded video signal 69 is finally provided by the decoding unit 60.

[0033] (CABAC) HEVC uses multiple types of entropy coding, such as context adaptive binary arithmetic coding (CABAC), Golomb-Rice coding, or a simple binary representation called fixed-length coding. In most cases, binary coding processes are performed to represent different syntax elements. This binary coding process is also very specific and depends on different syntax elements. Arithmetic coding represents syntax elements according to their current probabilities. CABAC is an extension of arithmetic coding that separates the probabilities of syntax elements according to the "context" defined by context variables. This corresponds to conditional probability. The context variables can be derived from the current syntax values of the upper left block that has already been coded (A2 in FIG. 6b, as described in more detail below) and the upper left block (B3 in FIG. 6b).

[0034] (Inter-coding) HEVC uses three different inter modes, namely, the inter mode, the merge mode, and the merge skip mode. The main difference between these modes is the data transmission in the bitstream. For motion vector coding, the current HEVC standard includes a competition-based method for motion vector prediction that did not exist in previous versions of the standard. For each of the inter or merge modes, it means that several candidates are competing with the distortion rate criterion on the encoder side to find the best motion vector predictor or the best motion information. The index corresponding to the best predictor or the best candidate of the motion information is inserted into the bitstream. The decoder can derive the same set of predictors or candidates and use the best one according to the multiplexed index. In the screen content extension of HEVC, a new coding tool called intra block copy is transmitted as one of those three inter modes, and the difference between the inter mode equivalent to IBC is done by checking whether the reference frame is the current one. This can be implemented, for example, by checking the reference index of list L0 and, if this is the last frame in that list, assuming that this is an intra block copy. Another way to implement this is to compare the current picture order count with the reference frame, i.e., if they are equal, this is an intra block copy.

[0035] The design of predictors and candidate derivation is important in achieving the best coding efficiency without unevenly affecting complexity. In HEVC, two motion vector derivations are used, namely, one for the inter mode (Advanced Motion Vector Prediction (AMVP)) and one for the merge mode (merge derivation process). These processes are described below. FIGS. 6a and 6b show the spatial and temporal blocks that can be used to generate motion vector predictors in the Advanced Motion Vector Prediction (AMVP) and merge mode of the HEVC encoding and decoding system, and FIG. 7 shows a simplified step of the process for deriving the AMVP predictor set. Two predictors, namely the two spatial motion vectors in the AMVP mode, are selected from the upper block (indicated by the letter "B") and the left block (indicated by the letter "A") that include the upper corner block (block B2) and the left corner block (block A0), and one predictor is selected from the lower right block (H) and the central block (center) of the arranged blocks as shown in FIG. 6a. Table 1 below shows an overview of the nomenclature used when referring to the blocks in terms of the current block as shown in FIGS. 6a and 6b. It should be understood that this nomenclature is used as a shorthand notation, but other labeling systems may be used, especially in future versions of the standard.

[0036]

Table 1

[0037] It should be noted that the "current block" may be variable in size, for example, 4x4, 16x16, 32x32, 64x64, 128x128, or any size in between. The dimensions of the block are preferably two factors (i.e., 2^n×2^m, where n and m are positive integers) so as to result in a more efficient use of bits when using binary coding. The current block need not be square, but this is often a preferred embodiment for the complexity of coding. Referring to FIG. 7, the first step aims to select, from among the lower left blocks A0 and A1, the first spatial predictor (candidate 1, 706) whose spatial position is shown in FIG. 6. For this purpose, these blocks are selected one after another in a predetermined order (700, 702), and for each selected block, the following conditions are evaluated in a predetermined order (704). The first block that satisfies the conditions is set as the predictor, that is, the motion vector from the same reference list and the same reference image, the motion vector from the same reference image of another reference list, the scaled motion vector from a different reference image of the same reference list, or the scaled motion vector from a different reference image of another reference list.

[0038] If the value cannot be found, the left predictor is considered unavailable. In this case, it indicates that the related blocks are intra-coded or that those blocks do not exist. The next step aims to select a second spatial predictor (Candidate 2, 716) from among the upper-right block B0, upper block B1, and upper-left block B2 whose spatial positions are shown in FIG. 6. For this purpose, these blocks are selected one by one in a predetermined order (708, 710, 712), and for each selected block, the above conditions are evaluated in a predetermined order (714), and the first block that satisfies the above conditions is set as the predictor. Again, if the value cannot be found, the upper predictor is considered unavailable. In this case, it indicates that the related blocks are intra-coded or that those blocks do not exist. In the next step (718), if both predictors are available, they are compared with each other to remove one of them if they are equal (i.e., the same motion vector value, the same reference list, the same reference index, and the same direction type). If only one spatial predictor is available, the algorithm is looking for a temporal predictor in the next step.

[0039] The temporal motion predictor (Candidate 3, 726) is derived as follows. That is, the bottom right (H, 720) position of the block placed in the previous frame is first considered in the availability check module 722. If it does not exist, or if the motion vector predictor is not available, the center (Center, 724) of the placed block is selected to be checked. These temporal positions (Center and H) are shown in FIG. 6. In either case, the scaling 723 is applied to those candidates so that the temporal distance between the current frame and the first frame matches that in the reference list. Next, the motion predictor value is added to the set of predictors. Next, the number of predictors (Nb_Cand) is compared with the maximum number of predictors (Max_Cand) (728). As described above, the maximum number of motion vector predictors (Max_Cand) that need to be generated by the AMVP derivation process is 2 in the current version of the HEVC standard. When this maximum number is reached, the final list or set (732) of AMVP predictors is constructed. Otherwise, a zero predictor is added to the list (730). The zero predictor is a motion vector equal to (0, 0).

[0040] As shown in FIG. 7, the final list or set (732) of AMVP predictors is constructed from a subset (700 to 712) of spatial motion predictors and a subset (720, 724) of temporal motion predictors. As described above, motion predictor candidates in the merge mode or merge skip mode represent all the necessary motion information of direction, list, reference frame index, and motion vector. An indexed list of multiple candidates is generated by the merge derivation process. In the current HEVC design, the maximum number of candidates for both merge modes is equal to 5 (4 spatial candidates and 1 temporal candidate). FIG. 8 is a schematic diagram of the motion vector derivation process in the merge mode. In the first step of the derivation process, five block positions are considered (800 to 808). These positions are the spatial positions shown in FIG. 3 at reference A1, B1, B0, A0, and B2. In the next step, the availability of spatial motion vectors is checked and at most five motion vectors are selected (810). The predictor is considered available if it exists and the block is not intra-coded.

[0041] Therefore, the selection of motion vectors corresponding to the five blocks as candidates is performed according to the following conditions, that is, when the A1 motion vector (800) of "left" is available (810), that is, when it exists and this block is not intra-coded, the motion vector of the "left" block is selected and used as the first candidate in the candidate list (814). When the B1 motion vector (802) of "above" is available (810), the motion vector of the candidate "above" block is compared with the A1 motion vector of "left" if it exists (812). When the B1 motion vector is equal to the A1 motion vector, B1 is not added to the list of spatial candidates (814). Conversely, when the B1 motion vector is not equal to the A1 motion vector, B1 is added to the list of spatial candidates (814). When the B0 motion vector (804) of "top right" is available (810), the motion vector of "top right" is compared with the B1 motion vector (812). When the B0 motion vector is equal to the B1 motion vector, the B0 motion vector is not added to the list of spatial candidates (814). Conversely, when the B0 motion vector is not equal to the B1 motion vector, the B0 motion vector is added to the list of spatial candidates (814).

[0042] When the A0 motion vector (806) of "bottom left" is available (810), the motion vector of "bottom left" is compared with the A1 motion vector (812). When the A0 motion vector is equal to the A1 motion vector, the A0 motion vector is not added to the list of spatial candidates (814). Conversely, when the A0 motion vector is not equal to the A1 motion vector, the A0 motion vector is added to the list of spatial candidates (814). When the list of spatial candidates does not contain four candidates, the availability of the B2 motion vector (808) of "top left" is checked (810). If it is available, it is compared with the A1 motion vector and the B1 motion vector. When the B2 motion vector is equal to the A1 motion vector or the B1 motion vector, the B2 motion vector is not added to the list of spatial candidates (814). Conversely, when the B2 motion vector is not equal to the A1 motion vector or the B1 motion vector, the B2 motion vector is added to the list of spatial candidates (814).

[0043] At the end of this stage, the list of spatial candidates contains up to four candidates. For temporal candidates, two positions can be used, namely, the lower right position of the placed block (816 indicated by H in FIG. 6) and the center of the placed block (818). These positions are shown in FIG. 6. For the AMVP motion vector derivation process, the first step aims to check the availability of the block at the H position (820). Next, if it is not available, the availability of the block at the center position is checked (820). If the motion vector of at least one of these positions is available, the temporal motion vector is scaled to the reference frame with index 0 for both lists L0 and L1 if necessary to create a temporal candidate (824) that is added to the list of merge motion vector predictor candidates (822). It is placed after the spatial candidates in the list. Lists L0 and L1 are two reference frame lists containing 0, one or more reference frames.

[0044] If the number of candidates (Nb_Cand) is strictly less than the maximum number of candidates (826) (the value is transmitted in the bitstream slice header and is equal to 5 in the current HEVC design, Max_Cand), and the current frame is of B type, combined candidates are generated (828). The combined candidates are generated based on the available candidates in the list of merge motion vector predictor candidates. It mainly consists of combining the motion vector of one candidate in list L0 with the motion vector of one candidate in list L1. If the number of candidates (Nb_Cand) remains strictly less than the maximum number of candidates (Max_Cand) (830), zero motion candidates are generated (832) until the number of candidates in the merge motion vector predictor candidate list reaches the maximum number of candidates. At the end of this process, a list or set of merge motion vector predictor candidates is constructed (834). As shown in FIG. 8, the list or set of merge motion vector predictor candidates is constructed (834) from a subset of spatial candidates (800~808) and a subset of temporal candidates (816, 818).

[0045] (Alternative Temporal Motion Vector Prediction (ATMVP)) Alternative Temporal Motion Vector Prediction (ATMVP) is a specific motion compensation. Instead of considering only one motion information for the current block from the temporal reference frame, each motion information of the respectively arranged blocks is considered. Therefore, this temporal motion vector prediction gives the division of the current block, together with the relevant motion information of each sub-block, as shown in FIG. 9. In the current VTM reference software, ATMVP is transmitted as a merge candidate inserted into the list of merge candidates. When ATMVP is enabled at the SPS level, the maximum number of merge candidates is increased by one. Therefore, when this mode is disabled, 6 candidates are considered instead of 5.

[0046] Furthermore, when this prediction is enabled at the SPS level, all bins of the merge index become contexts encoded by CABAC. While inside HEVC, or when ATMVP is not enabled at the SPS level, only the first bin is the context encoded, and the remaining bins are bypass encoding contexts. FIG. 10(a) shows the encoding of the merge index for HEVC or when ATMVP is not enabled at the SPS level. This corresponds to unary maximum encoding. Furthermore, the first bit is CABAC encoded and the other bits are bypass CABAC encoded. FIG. 10(b) shows the encoding of the merge index when ATMVP is enabled at the SPS level. Furthermore, all bits are CABAC encoded (from the first bit to the fifth bit). Note that each index has its own context, in other words, their probabilities are separated.

[0047] (Affine Mode) In HEVC, only translational motion models are applied for motion compensated prediction (MCP). On the other hand, in the real world, there are many types of motions such as zoom in / zoom out, rotation, perspective motion, and other irregular motions. In JEM, simplified affine transform motion compensated prediction is applied, and based on the extraction of document JVET - G1001 presented at the JVET meeting in Turin from July 13th to 21st, 2017, the general principle of the affine mode is described below. The whole of this document is incorporated herein by reference as it describes other algorithms used in JEM. As shown in Fig. 11(a), the affine motion field of a block is described by two control point motion vectors. The motion vector field (MVF) of the block is described by Equation 1 below.

[0048]

Equation

[0049] Here, (v 0x , v 0y ) is the motion vector of the control point at the upper left corner, and (v 1x , v 1y ) is the motion vector of the control point at the upper right corner. To further simplify the motion compensated prediction, sub - block - based affine transform prediction is applied. The sub - block size is derived as in Equation 2, where MvPre is the motion vector fractional precision (1 / 16 in JEM), and (v2x, v2y) is the motion vector of the lower left control point calculated according to Equation 1.

[0050]

Equation

[0051] After being derived by Equation 2, M and N may be adjusted downward to be divisors of w and h, respectively, if necessary. To derive the motion vector of each M×N sub-block, as shown in FIG. 6a, the motion vector of the central sample of each sub-block is calculated according to Equation 1 and rounded up to 1 / 16 fractional precision. Next, a motion compensation interpolation filter is applied to generate a prediction for each sub-block having the derived motion vector.

[0052] The affine mode is a motion compensation mode as an inter mode (AMVP, merge, merge skip). Its principle is to generate one motion information for each pixel according to two or three adjacent motion information. In the current VTM reference software, as shown in FIG. 11(a), the affine mode derives one motion information for each 4x4 block. This mode can be used for AMVP, and both merge modes are enabled by a flag. This flag is CABAC encoded. In one embodiment, the context depends on the sum of the affine flags of the left block (position A2 in FIG. 6b) and the upper left block (position B3 in FIG. 6b).

[0053] Therefore, in JEM, three context variables (0, 1, or 2) can be taken for the affine flag given by the following equation. Ctx = IsAffine(A2) + IsAffine(B3) Here, IsAffine(block) is a function that returns 0 when the block is not an affine block and 1 when the block is affine.

[0054] (Affine Merge Candidate Derivation) In JEM, the affine merge mode (merge or merge skip) is derived from the first adjacent block that is affine among the blocks at positions A1, B1, B0, A0, B2. These positions are shown in FIGS. 6a and 6b. However, how the affine parameters are derived is not fully defined, and the present invention aims to improve at least this point.

[0055] (Affine merge signaling) FIG. 12 is a flowchart of a partial decoding process of some syntax elements related to the coding mode. In this figure, a skip flag (1201), a prediction mode (1211), a merge flag (1203), a merge index (1208), and an affine flag (1207) can be decoded. For all CUs within an inter-slice, the skip flag is decoded (1201). If the CU is not skipped (1202), the pred mode (prediction mode) is decoded (1211). This syntax element indicates whether the current CU is in an inter or intra mode. Note that if the CU is skipped (1202), its current mode is inter mode. For a CU (1212), the CU is coded within the AMVP or in the merge mode. If the CU is inter (1212), the merge flag is decoded (1203). If the CU is a merge (1204) or the CU is skipped (1202), it is verified whether the affine flag (1206) needs to be decoded (1205). If the current CU is a 2N×2N CU, which means that in the current VVC, the height and width of the CU must be equal, this flag is decoded.

[0056] Furthermore, at least one adjacent CU A1 or B1 or B0 or A0 or B2 must be encoded in affine mode (merge or AMVP). Finally, the current CU must not be a 4x4 CU, and by default the CU 4x4 is disabled in the VTM reference software. If this condition (1205) is false, it is certain that the current CU is encoded in classical merge mode or merge skip mode and the merge index is decoded (1208). If the affine flag (1206) is set equal to 1 (1207), the CU is a merge affine CU or a merge skip affine CU, and the merge index (1208) need not be decoded. Otherwise, the current CU is a classical (basic) merge or merge skip CU, and the merge index candidates (1208) are decoded. Otherwise, the current CU is a classical (basic) merge or merge skip CU, and the merge index candidates (1208) are decoded. In this specification, "transmission" can mean the insertion into or extraction from a bitstream of one or more syntactic elements representing the enabling or disabling of other mode information.

[0057] (Merge candidate derivation) Figure 13 is a flowchart showing merge candidate derivation. This derivation is built on top of the HEVC merge list derivation shown in FIG. 8. The main changes compared to HEVC are the addition of ATMVP candidates (1319, 1321, 1323), the complete duplicate check of candidates (1320, 1325), and the new order of candidates. The ATMVP prediction is set as a special candidate since it represents some motion information of the current CU. The values of the first sub-block (upper left) are compared with the temporal candidates, and the temporal candidates are not added to the merge list if they are equal (1320). The ATMVP candidates are not compared with the other spatial candidates. Contrary to the temporal candidates that are compared with each spatial candidate already in the list (1325), if it is a duplicate candidate, it is not added to the merge candidate list. <U+ <U+

[0058] <U+ When a spatial candidate is added to the list, it is compared to other spatial candidates in the list which is not the case for the final version of HEVC (1310). In the current VTM version, the list of merge candidates is set as the following order since it is determined to provide the best results across the encoding test conditions. · A1 · B1 · B0 · A0 · ATMVP · B2 · Temporal · Combined · Zero_MV

[0059] It is important to note that the spatial candidate B2 is set after the ATMVP candidate. Furthermore, when ATMVP is enabled at the slice level, the maximum number of the candidate list is 6 instead of 5. The object of the present invention is to transmit the affine mode to a part of the bitstream in an efficient manner considering the encoding efficiency and complexity. Also, the object of the present invention is to transmit the affine mode in a way that requires the minimum amount of structural change to the existing video encoding framework. Here, exemplary embodiments of the present invention will be described with reference to FIGS. 13 to 21. The embodiments may be combined unless otherwise specified. For example, a specific combination of the embodiments may improve the encoding efficiency with increased complexity, which should be noted may be acceptable in specific use cases. Generally, in order to use the affine mode which has a higher possibility of providing improved motion compensation, it is possible to improve the encoding efficiency with an acceptable increase in encoding complexity by modifying the syntax for transmitting the motion predictor mode.

[0060] (First Embodiment) In the first embodiment, the affine motion prediction mode can be transmitted (e.g., enabled or disabled) for a part of the bitstream for at least one inter-mode. The inter-prediction mode used for a part of the bitstream is determined, and the affine motion mode is transmitted (enabled or disabled) depending on the inter-prediction mode used for a part of the bitstream. The advantage of this embodiment is that better coding efficiency is achieved by removing unused syntax. Further, it reduces the complexity of the coding unit by avoiding some inter-coding possibilities that do not need to be evaluated. Finally, on the coding unit side, some affine flags to be CABAC-coded do not need to be extracted from the bitstream that improves the efficiency of the coding process. In the example of the first embodiment, the skip mode is not enabled for the affine mode. If the CU is a skipped CU (based on the state or presence of the skip flag in the data stream), it means that the affine flag does not need to be extracted from the bitstream. FIG. 14 (which shares the same structure as FIG. 12 and the corresponding description applies here) shows this example. In FIG. 14, when the CU is a skip (1402), the affine flag (1406) is not decoded and the condition at 1405 is not evaluated. When the CU is a skip, the merge index is decoded (1406).

[0061] The advantage of this example is the improvement in coding efficiency for sequences with little motion, rather than the reduction in coding efficiency for sequences with more motion. This is because the skip mode is typically used when there is little or no motion, and as such, it is less likely that the affine mode is appropriate. As described above, the complexity of the encoding and decoding processes is also reduced. In an additional example, the affine merge skip mode can be enabled or disabled at a high level, for example, at the slice, frame, sequence, or CTU level. This may be determined based on a high-level syntax flag. In such a case, the affine merge skip may be disabled for sequences or frames with little motion and enabled when the amount of motion increases. The advantage of this additional example is the flexibility regarding the use of affine merge skip.

[0062] In one embodiment, the affine merge skip mode is not evaluated on the encoding side, and as a result, the bitstream does not include the affine merge skip mode. The advantage is that coding efficiency is observed, but it is smaller than in the first embodiment. For the affine, the merge and skip modes may not be enabled (or the affine is only enabled in AMVP). In a further example, the merge and merge skip modes are not enabled for the affine mode. When a CU is skipped or merged, it means that the affine flag does not need to be extracted from the bitstream. Comparing with FIG. 14, in this embodiment, modules 1405, 1406, and 1407 are removed. The advantage of this example is the same as the example immediately above. The advantage is the improvement in coding efficiency for sequences with little motion and the same coding efficiency for sequences with more motion. As described above, the complexity of the encoding and decoding processes is reduced.

[0063] The high-level syntax element conveys that affine merge can be enabled. In yet another example, the affine merge mode and the merge skip mode may be enabled or disabled at a high level as a slice, frame, sequence, or CTU level having one flag. In this case, affine merge may be disabled for a sequence or a frame having little motion, and may be enabled when the amount of motion increases. The advantage of this additional embodiment is the flexibility regarding the use of affine skip. In another example, one flag is transmitted for both the merge skip mode and the merge mode. In another example, the affine merge skip mode and the merge mode are not evaluated at the encoding side. As a result, the bitstream does not include the affine merge skip mode. The advantage is that encoding efficiency can be observed.

[0064] (Second Embodiment) In the second embodiment, transmitting the affine mode of the current block depends on the modes of one or more adjacent blocks. There may be a correlation with how the adjacent blocks that can be used to improve encoding efficiency are encoded. In particular, when one or more specific adjacent blocks use the affine mode, the affine mode is more likely to be appropriate for the current mode. In one embodiment, the number of candidates for affine merge or affine merge skip mode is reduced to only two candidates. The advantage of this embodiment is the reduction of complexity on the decoder side because fewer affine flags are decoded for the merge mode and fewer comparisons and memory buffer accesses are required for the affine merge check condition (1205). Fewer affine merge modes on the encoder side need to be evaluated.

[0065] In an example of the second embodiment, one adjacent block and one adjacent block to the left above the current block (e.g., candidate A1 and B1 as shown in FIG. 6) are evaluated to know whether the affine flag needs to be decoded and for the derivation of affine merge candidates. The advantage of using only these two positions for affine merge is that, as a current VTM implementation with reduced complexity, it has the same coding efficiency as maintaining five candidates. FIG. 15 shows this embodiment. In this figure compared to FIG. 12, module 1505 has been changed by checking only the positions of A1 and B1. In a further example of the second embodiment, only candidates A2 and B3 as shown in FIG. 6b are evaluated to determine whether the affine flag needs to be decoded and for the derivation of affine merge candidates. The advantage of this example is the same as the previous example, but it also reduces the "worst case" memory access compared to the previous example. In fact, at positions A2 and B3, the positions are the same as those used for affine flag context derivation. In fact, for the affine flag, context derivation depends on the adjacent blocks at positions A2 and B3 in FIG. 6b. As a result, if the affine flag needs to be decoded, the affine flag values of A2 and B3 are already in memory for the current affine flag context derivation, and thus no further memory access is required.

[0066] (Third Embodiment) In the third embodiment, communicating the affinity mode of the current block depends on a list of merge candidates corresponding to blocks adjacent to the current block. In one example of the third embodiment, the list starts from the block used to determine the context variable associated with the block because the affinity flag value of such a block is already in memory and no further memory access is required for context derivation of the current affinity flag. For example, possible affinity merge candidates are, as shown in FIG. 6(b), in the order of A2 or B3 or BO or AO or B2 (instead of A1 or B1 or BO or AO or B2). This gives an improvement in coding efficiency compared to the current VTM. And it also limits the amount of affinity flags that need to be accessed for decoding of the affinity flags for the worst-case scenario. In the current version, only 5 for module 1205 and 2 for affinity flag context derivation and in the current embodiment, only 5 as the affinity flag values of A2 and B3 are already present in memory for context derivation of the current affinity flag, so no such further memory access is required.

[0067] The variation of the third embodiment relates to context alignment. Transmitting the affine mode may include inserting a context encoding flag into the data stream, and the context variable for the flag is determined based on whether adjacent blocks use the affine mode. In an alternative example of the third embodiment, the positions considered for context derivation of the affine flag are positions A1 and B1 instead of positions A2 and B3, as shown in FIG. 6b. In that case, the same advantages as the previous example are obtained. This is another alignment between the context and the affine merge derivation. In that case, the context variable of the affine flag is obtained according to the following formula, that is, Ctx = IsAffine(A1)+ IsAffine(B1), where Ctx is the context of the affine flag, and IsAffine is a function that returns 0 when the block is not an affine block and 1 when the block is affine. In this example, for the current context derivation of the affine flag, the affine flag values of A1 and B1 are stored in memory, and as such, no further memory access is required.

[0068] In a further alternative example, the positions considered for context derivation of the affine flag are positions A1 and B1 instead of positions A2 and B3 when the current block is in the merge mode (both merge modes). An additional advantage compared to the previous example is better coding efficiency. In fact, for AMVP, since affine blocks are not considered for this derivation, there is no need for context derivation to be aligned with the motion vector derivation for AMVP.

[0069] (Fourth Embodiment) In the fourth embodiment, the transmission of the affine mode is performed depending on whether the adjacent blocks are in the merge mode or not. In an example of the fourth embodiment, the candidate for affine merge (merge and skip) may be only the affine AMVP candidate. FIG. 17 shows this embodiment. The advantage of this embodiment is that since only a few affine flags need to be decoded without affecting the coding efficiency, the coding complexity is reduced. In a further example of the fourth embodiment, the candidate for affine merge (merge and skip) may be only the affine AMVP candidate or the merge affine candidate, but not necessarily the affine merge skip. Similar to the previous example, the advantage of this example is that since only a few affine flags need to be decoded without affecting the coding efficiency, the coding complexity is reduced.

[0070] (Fifth Embodiment) In the fifth embodiment, transmitting the affine mode includes inserting the affine mode as a candidate motion predictor. In an example of the fifth embodiment, the affine merge (and merge skip) is transmitted as a merge candidate. In this case, the modules 1205, 1206, and 1207 in FIG. 12 are removed. In addition, the maximum possible number of merge candidates is incremented so as not to affect the coding efficiency of the merge mode. For example, in the current VTM version, this value is set equal to 6, and thus, when applying this embodiment to the current version of VTM, the value becomes 7. The present advantage is that since only a few syntax elements need to be decoded, the design of the syntax elements of the merge mode is simplified. In some situations, coding efficiency can be observed.

[0071] Here, two possibilities for implementing this example are described below. The position of the candidate motion predictor indicates the likelihood of its selection, and as such, if it is placed higher in the list (lower index value), it indicates a higher likelihood that the motion vector predictor will be selected. In the first example, the affine merge index always has the same position within the list of merge candidates. This means having a fixed merge idx value. For example, since the affine merge mode should represent complex motions that are not the most likely content, this value could be set equal to 5. An additional advantage of this embodiment is that when the current block not only decodes the data itself but also performs syntax analysis / decoding / reading of the syntax elements, the current block can be set as an affine block. Consequently, the value can be used to determine the CABAC context of the affine flag used for the AMVP. Therefore, the conditional probability should be improved for this affine flag, and the coding efficiency will be better.

[0072] In the second example, the affine merge candidates are derived along with the other merge candidates. In this example, new affine merge candidates are added to the list of merge candidates. FIG. 18 shows this example. Comparing with FIG. 13, the affine candidates are the first affine adjacent blocks A1, B1, BO, AO, B2 (1917). When the same conditions as 1205 in FIG. 12 are valid (1927), a motion vector field generated using the affine parameters is generated to obtain the affine candidates (1929). The initial list of candidates can have 4, 5, 6, or 7 candidates according to the use of ATMVP, temporal, and affine candidates. The order among all these candidates is important because the more likely candidates should be processed first to ensure that they are more likely to perform the cut of the motion vector candidates, and the preferred order is as follows. A1 B1 B0 A0 Affine merge ATMVP B2 temporal combination Zero_MV

[0073] It is important to note that the affine merge is before the ATMVP mode but after the four main adjacent blocks. The advantage of setting the affine merge before the ATMVP candidates is an improvement in coding efficiency as compared to setting it after the ATMVP and the temporal predictors. This improvement in coding efficiency depends on the GOP (Group of Pictures) structure and the quantization parameter (QP) setting of each picture within the GOP. However, for the most commonly used GOP and QP settings, this order gives an improvement in coding efficiency. A further advantage of this solution is a clean design of merge and merge skip for both syntax and derivation. Additionally, the affine candidate merge index can be changed according to the availability or value (duplicate check) of the previous candidates in the list. As a result, efficient transmission is obtained. In a further example, the affine merge index is variable according to one or several conditions.

[0074] For example, the merge index or position within the list associated with the affine candidate changes according to a criterion. The principle is to set a low value for the merge index corresponding to the affine merge when there is a high probability that the affine merge is selected (and a higher value in case of a low probability of being selected). The advantage of this example is an improvement in coding efficiency thanks to an optimal adaptation of the merge index when it is most likely to be used.

[0075] The criteria for selecting the position of the affine mode within the list of merge candidates include the following. a) When the skip mode is enabled (status of the skip flag) In an example of applying this criterion, if the affine merge index has a value set equal to a high value (e.g., 5), or if the current merge is in the merge skip mode, it is set after the spatial and temporal MVs. As described for the first embodiment, since there is a low possibility of any large (or complex) motion, the possibility that the affine mode is selected for the skip mode is low. b) Motion information of adjacent blocks In an example of applying this criterion, if the affine merge index has a value set equal to a high value, or if the motion information of one block to the left and one block above (e.g., blocks A1 and B1) is similar or equal, it is set after the spatial and temporal MVs. If A1 has the same motion information as B1, the motion information has a high probability of being constant for the current block. Therefore, the affine merge has a low probability of being selected.

[0076] c) ATMVP candidates In an example of applying this criterion, if the affine merge index has a value set equal to a high value, or if the ATMVP candidate includes only one motion information, it is set after the spatial and temporal MVs. In that case, there is no subdivision in the frame before the placed block. Therefore, since there is a slight possibility that the current block content is within non-constant motion, it is preferable not to set the affine at a high position in the merge list. d) When adjacent blocks use the affine mode In an example of applying this criterion, if the affine merge index is set to an equal low value or two or more adjacent blocks are affine, it is set before temporal prediction and far from the spatial predictor. In an additional example of applying this criterion, the affine merge index or affine position (idx) is set equal to idx = P - N, where P is the lowest possible position for the affine merge index and N is the number of affine adjacent blocks. In one example, P is 5, N is 5, and the adjacent blocks are A1, B1, B0, A0, B2. It should be noted that in this notation, the highest position has an index value of 0.

[0077] In this example, the affine merge index of the merge candidate position is set according to the probability associated with its adjacent blocks. Thus, at the first position when all adjacent positions are affine and at the fourth position when only one adjacent block is affine. It should be understood that the exemplary value "5" can be set to 6 or 7 to obtain similar coding efficiency. Also, it should be understood that combinations of these criteria are possible.

[0078] In another example of the fifth embodiment, the affine mode is transmitted depending on the likelihood of the affine mode of the current block for the determination. In a specific example, the affine merge candidate is not added to the list of candidates or there is no index corresponding to the affine merge according to the criterion. The principle of this example is to invalidate the affine mode that is unlikely to be useful. The advantage of this example is an improvement in coding efficiency thanks to the optimal use of the merge index bits.

[0079] The criteria for determining the possibility of a useful affine mode include the following. a) The status of the skip flag In an example of applying this criterion, if the current merge is in the merge skip mode, the affine merge candidate is not added. As explained in the first embodiment, there is a low possibility that the affine mode is selected in the skip mode. b) Motion information of adjacent blocks In one embodiment of applying this criterion, if the motion information of one block on the left and one block above (e.g., blocks A1 and B1) is similar or equal, the affine merge candidate is not added. If one block on the left and one block above (e.g., blocks A1 and B1) have the same motion information, there is a high probability that the motion information is constant with respect to the current block. Therefore, the affine merge will be invalidated. c) ATMVP candidate In one embodiment of applying this criterion, if the ATMVP candidate contains only one motion information, the affine merge candidate is not added. In such an example, since there is a slight possibility that the current block content is inside non-constant motion, it is preferable to invalidate the affine at a high position in the merge list. It should be understood that combinations of these criteria are possible.

[0080] (Implementation of embodiments of the present invention) FIG. 20 is a schematic block diagram of a computing device 1300 for implementing one or more embodiments of the present invention. The computing device 1300 may be a device such as a microcomputer, a workstation, or a light portable device. The computing device 1300 includes a central processing unit (CPU) 1301 such as a microprocessor, a random access memory (RAM) 1302 for storing executable code of the method of one or more embodiments of the present invention, and registers adapted to record variables and parameters necessary for implementing a method for encoding or decoding at least a part of an image according to an embodiment of the present invention. Their memory capacity may be expanded, for example, by optional RAM connected to an expansion port. A read-only memory (ROM) 1303 for storing a computer program for implementing an embodiment of the present invention, and a network interface (NET) 1304 is typically connected to a communication network through which digital data to be processed is transmitted or received. The network interface (NET) 1304 may be a single network interface or may be composed of a set of different network interfaces (for example, wired and wireless interfaces, or different types of wired or wireless interfaces).

[0081] Data packets are written to the network interface for transmission or read from the network interface for reception under the control of a software application executed by the CPU 1301. The user interface (UI) 1305 may be used to receive input from the user or display information to the user. The hard disk (HD) 1306 may be provided as a mass storage device. The input / output module (IO) 1307 may be used to receive / send data from / to an external device such as a video source or a display. The executable code may be stored in any of the ROM 1303, HD 1306, or a removable digital medium such as a disk. According to a variant, the executable code of the program can be received by means of a communication network via NET 1304 in order to be stored in one of the storage means of the communication device 1300 such as HD 1306 before being executed. The CPU 1301 is adapted to control and direct the execution of the instructions or portions of a program or program software code according to an embodiment of the invention, where the instructions are stored in one of the aforementioned storage means. After power-on, the CPU 1301 can execute instructions related to a software application from the main RAM memory 1302, for example, after these instructions are loaded from the program ROM 1303 or HD 1306. When such a software application is executed by the CPU 1301, it causes the steps of the method according to the invention to be executed.

[0082] Also, according to another embodiment of the present invention, it is understood that a decoding unit according to the above-described embodiment is provided in a user terminal such as a computer, a mobile phone (cellular phone), a tablet, or any other type of device (e.g., a display device) capable of providing / displaying content to a user. According to yet another embodiment, the encoding unit according to the above-described embodiment is provided in an image capture device that also includes a camera, a video camera, or a network camera (e.g., a closed-circuit television or a video surveillance camera) that captures and provides the content to be encoded by the encoding unit. Two such examples are provided below with reference to FIGS. 21 and 22.

[0083] FIG. 21 is a diagram showing a network camera system 2100 including a network camera 2102 and a client device 2104. The network camera 2102 includes an imaging unit 2106, an encoding unit 2108, a communication unit 2110, and a control unit 2112. The network camera 2102 and the client device 2104 are interconnected so as to be communicable with each other via a network 200. The imaging unit 2106 includes a lens and an imaging element (e.g., a charge-coupled device (CCD) or a complementary metal-oxide semiconductor (CMOS)), captures an image of a subject, and generates image data based on the image. This image may be a still image or a video image. Also, the imaging unit may include zoom means and / or pan means adapted to zoom or pan (either optically or digitally). The encoding unit 2108 encodes the image data using the encoding method described in the first to fifth embodiments. The encoding unit 2108 uses at least one of the encoding methods described in the first to fifth embodiments. For other examples, the encoding unit 2108 can use a combination of the encoding methods described in the first to fifth embodiments.

[0084] The communication unit 2110 of the network camera 2102 transmits the encoded image data encoded by the encoding unit 2108 to the client device 2104. Further, the communication unit 2110 receives a command from the client device 2104. The command includes a command for setting parameters for encoding by the encoding unit 2108. The control unit 2112 controls other units within the network camera 2102 according to the command received by the communication unit 2110. The client device 2104 includes a communication unit 2114, a decoding unit 2116, and a control unit 2118. The communication unit 2118 of the client device 2104 transmits a command to the network camera 2102. Further, the communication unit 2118 of the client device 2104 receives the encoded image data from the network camera 2102. The decoding unit 2116 decodes the encoded image data by using the decoding method described in any of the first to fifth embodiments. As another example, the decoding unit 2116 can use a combination of the decoding methods described in the first to fifth embodiments.

[0085] The control unit 2118 of the client device 2104 controls other parts within the client device 2104 according to the user operations or commands received by the communication unit 2114. The control unit 2118 of the client device 2104 controls the display device 2120 to display the image decoded by the decoding unit 2116. Also, the control unit 2118 of the client device 2104 controls the display device 2120 to display a GUI (Graphical User Interface) that specifies the values of the parameters of the network camera 2102, including the parameters for encoding by the encoding unit 2108. Further, the control unit 2118 of the client device 2104 controls other parts within the client device 2104 in response to user operation inputs to the GUI displayed by the display device 2120. The control unit 2118 of the client device 2104 controls the communication unit 2114 of the client device 2104 to transmit a command that specifies the values of the parameters of the network camera 2102 to the network camera 2102 in response to user operation inputs to the GUI displayed by the display device 2120. The network camera system 2100 can determine whether the camera 2102 uses zoom or pan during video recording, and such information can be used while encoding the video stream as zoom or pan, benefiting from the use of the affine mode, which is well-suited for encoding complex movements such as zooming, rotating, and / or stretching (which can be a side effect of panning, especially when the lens is a "fisheye" lens).

[0086] Figure 22 is a diagram showing a smartphone 2200. The smartphone 2200 includes a communication unit 2202, a decoding / encoding unit 2204, a control unit 2206, and a display unit 2208. The communication unit 2202 receives encoded image data via a network. The decoding unit 2204 decodes the encoded image data received by the communication unit 2202. The decoding unit 2204 decodes the encoded image data by using the decoding method described in the first to fifth embodiments. The decoding unit 2204 can use at least one of the decoding methods described in the first to fifth embodiments. For example, the encoding unit 2202 can use a combination of the decoding methods described in the first to fifth embodiments. The control unit 2206 controls other units within the smartphone 2200 in response to a user operation or command received by the communication unit 2202.

[0087] For example, the control unit 2206 controls the display device 2208 to display the image decoded by the decoding unit 2204. The smartphone can further include an image recording device 2210 (for example, a digital camera associated with a circuit) for recording an image or video. Such a recorded image or video may be encoded by the decoding / encoding unit 2204 under the instruction of the control unit 2206. The smartphone may further include a sensor 2212 adapted to sense the orientation of the mobile device. Such a sensor can include an accelerometer, a gyroscope, a compass, a global positioning (GPS) unit, or a similar position sensor. Such a sensor 2212 can determine whether the smartphone changes its orientation, and such information may be used when encoding a video stream as a change in orientation while shooting can benefit from the use of the affine mode which is well-suited for encoding complex movements such as rotation.

[0088] (Alternative Examples and Modification Examples) The object of the present invention is to ensure that the affine mode is utilized in the most efficient way, and it will be understood that the above specific examples relate to communicating the use of the affine mode depending on the likelihood that the affine mode is perceived as useful. Further examples of this may be applied to the encoding unit when it is known that complex motion (for which an affine transformation is particularly efficient) is being encoded. Examples of such cases include the following. a) Camera zoom in / zoom out b) A portable camera (e.g., a mobile phone) that changes orientation during shooting (i.e., rotational motion) c) Panning of a "fisheye" lens camera (e.g., stretching / distorting of a part of the image)

[0089] As such, the display of complex motion may be raised during the recording process so that the affine mode is highly likely to be used for slices, frame sequences, or the entire actual video stream. In a further example, a high likelihood of use may be given depending on the characteristics or functionality of the device for which the affine mode is used to record video. For example, since a mobile device is more likely to change orientation than a fixed surveillance camera (say), the affine mode may be more suitable for encoding video from the former. Examples of characteristics or functions include the presence / use of zoom means, the presence / use of a position sensor, the presence / use of panning means, whether the device is portable, or user selection on the device.

[0090] Although the present invention has been described with reference to embodiments, it should be understood that the present invention is not limited to the disclosed embodiments. It will be understood by those skilled in the art that various changes and modifications can be made without departing from the scope of the present invention as defined in the appended claims. All features disclosed in this specification (including the appended claims, abstract and drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. Each feature disclosed in this specification (including the appended claims, abstract and drawings) may be replaced by alternative features serving the same, equivalent, or similar purpose, unless otherwise stated. Thus, unless otherwise stated, each feature disclosed is only an example of a general series of equivalent or similar features.

[0091] Also, it should be understood that results shown or determined / inferred can be used in a process, for example, during decoding processing, instead of actually performing comparison, determination, evaluation, selection, execution, implementation, or consideration, such that any result of the above-described comparison, determination, evaluation, selection, execution, implementation, or consideration can be used, for example, a selection made during encoding or filtering processing can be indicated from data in a bitstream, for example, by a flag or data indicating the result, or may be determinable / inferable. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. The mere fact that different features are recited in mutually different dependent claims does not indicate that a combination of these features cannot be advantageously used. The reference signs appearing in the claims are for illustrative purposes only and do not have a limiting effect on the claims.

Claims

1. A method for decoding an image from a bitstream using motion prediction, comprising: determining whether to use an intra mode or an inter mode; when it is determined to use the inter mode, generating a list of motion predictor candidates that can include candidates for sub-block collocated temporal prediction and candidates for sub-block affine prediction; wherein the sub-block collocated temporal prediction can use the motion information of each of a plurality of sub-blocks included in a block at the same position as the target block in the reference image for the target block in the image; the order of the candidates for sub-block affine prediction in the list is variable based on whether a candidate for sub-block collocated temporal prediction is included in the list; A method characterized by the above.

2. A method for encoding an image into a bitstream using motion prediction, comprising: determining whether to use an intra mode or an inter mode; when it is determined to use the inter mode, generating a list of motion predictor candidates that can include candidates for sub-block collocated temporal prediction and candidates for sub-block affine prediction; wherein the sub-block collocated temporal prediction can use the motion information of each of a plurality of sub-blocks included in a block at the same position as the target block in the reference image for the target block in the image; the order of the candidates for sub-block affine prediction in the list is variable based on whether a candidate for sub-block collocated temporal prediction is included in the list; A method characterized by the above.

3. The method according to claim 1 or 2, characterized in that the maximum number of candidates in the list depends on whether the sub-block affine prediction is effective.

4. The method according to any one of claims 1 to 3, characterized in that the sub-block affine prediction derives motion information for each of a plurality of sub-blocks in a block according to two or three pieces of motion information.

5. A method for decoding an image from a bitstream, comprising: determining a prediction mode to be used for decoding the current block of the image from a plurality of prediction modes including an intra mode and an inter mode; When the inter mode is determined as the prediction mode for the current block, generating a list of a plurality of motion predictor candidates including candidates for sub-block affine prediction, The candidates for sub-block affine prediction are arranged as merge candidates below the temporal motion vector candidates in the list, The sub-block affine prediction is characterized by deriving at least one motion vector for each sub-block of the current block using two or three motion vectors.

6. Further comprising selecting a sub-block merge mode that uses sub-block affine prediction for the current block, Selecting the sub-block merge mode that uses sub-block affine prediction includes decoding a flag from the bitstream using CABAC decoding, The context variable for the flag is determined based on whether a first block adjacent to the current block uses sub-block affine prediction and whether a second block adjacent to the current block uses sub-block affine prediction. The method according to claim 5, characterized in that.

7. The method according to claim 6, characterized in that the first block is located to the left of the current block and the second block is located above the current block.

8. In a state where the current block has a size of 16×16, the number of sub-blocks in the current block is 16, and in the sub-block affine prediction, one motion vector is derived for each sub-block of the current block using the two or three motion vectors. The method according to claim 5, characterized in that.

9. The method according to claim 5, characterized in that the temporal motion vector candidates use motion vectors in blocks of an image different from the image including the current block.

10. A method for encoding an image in a bitstream, Determining a prediction mode to be used for encoding the current block of the image from a plurality of prediction modes including an intra mode and an inter mode; When the inter mode is determined as the prediction mode for the current block, generating a list of a plurality of motion predictor candidates including candidates for sub-block affine prediction, The candidates for the sub-block affine prediction are arranged as merge candidates below the temporal motion vector candidates in the list. The sub-block affine prediction uses two or three motion vectors to derive at least one motion vector for each sub-block of the current block. Claim 11 Further comprising selecting a sub-block merge mode that uses sub-block affine prediction for the current block. Selecting the sub-block merge mode that uses the sub-block affine prediction includes encoding a flag from the bitstream using CABAC encoding. The context variable for the flag is determined based on whether the first block adjacent to the current block uses sub-block affine prediction and whether the second block adjacent to the current block uses sub-block affine prediction. The method according to claim 10, wherein Claim 12 The method according to claim 11, wherein the first block is located to the left of the current block and the second block is located above the current block. Claim 13 In a state where the current block has a size of 16×16, the number of sub-blocks in the current block is 16. In the sub-block affine prediction, one motion vector is derived for each sub-block of the current block using the two or three motion vectors. The method according to claim 10, wherein Claim 14 The method according to claim 10, wherein the temporal motion vector candidates use motion vectors in blocks of an image different from the image including the current block. Claim 15 An encoding apparatus for encoding an image into a bitstream using motion prediction, means for determining whether to use the intra mode or the inter mode, means for generating a list of motion predictor candidates including candidates for sub-block collocated temporal prediction and candidates for sub-block affine prediction when it is determined to use the inter mode. The sub-block collocated temporal prediction can use the motion information of each of a plurality of sub-blocks included in a block at the same position as the target block in the reference image for the target block in the image. The order of candidates for the sub-block affine prediction in the list is variable based on whether candidates for sub-block collocated temporal prediction are included in the list. An encoding apparatus characterized by this. **Claim 16** A decoding apparatus that decodes an image from a bit stream using motion prediction, means for determining whether to use the intra mode or the inter mode, means for generating a list of motion predictor candidates that can include candidates for sub-block collocated temporal prediction and candidates for sub-block affine prediction when it is determined that the inter mode is used, the sub-block collocated temporal prediction can use the motion information of each of a plurality of sub-blocks included in a block at the same position as the target block in the reference image for the target block in the image, the order of candidates for the sub-block affine prediction in the list is variable based on whether candidates for sub-block collocated temporal prediction are included in the list, A decoding apparatus characterized by this. **Claim 17** A decoding apparatus that decodes an image from a bit stream, means for determining a prediction mode to be used for decoding the current block of the image from a plurality of prediction modes including the intra mode and the inter mode, means for generating a list of a plurality of motion predictor candidates including candidates for sub-block affine prediction when the inter mode is determined as the prediction mode for the current block, the candidates for the sub-block affine prediction are arranged as merge candidates below the temporal motion vector candidates in the list, the sub-block affine prediction is characterized by deriving at least one motion vector for each sub-block of the current block using two or three motion vectors. A decoding apparatus. **Claim 18** An encoding apparatus that encodes an image into a bit stream, means for determining a prediction mode to be used for encoding the current block of the image from a plurality of prediction modes including the intra mode and the inter mode, means for generating a list of a plurality of motion predictor candidates including candidates for sub-block affine prediction when the inter mode is determined as the prediction mode for the current block, The candidates for the sub-block affine prediction are arranged as merge candidates below the temporal motion vector candidates in the list. The sub-block affine prediction is an encoding device characterized by deriving at least one motion vector for each sub-block of the current block using two or three motion vectors. A computer program for causing a computer to execute the method according to claim 1 or 5. A computer program for causing a computer to execute the method according to claim 2 or 10. ​ ​

Citation Information

Patent Citations

  • Prediction image generation device, moving image decoding device, and moving image encoding device

    WO2017130696A1

  • Affine motion prediction for video coding

    WO2017200771A1

  • Image processing device and image processing method

    WO2018131523A1