Method for predicting time-based motion information, method for constructing a list of candidate motion information, and video decoding method
By predicting temporal motion information and constructing a candidate motion information list, the method enhances video coding efficiency by utilizing temporal redundancy between frames.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2022-01-04
- Publication Date
- 2026-05-26
AI Technical Summary
Existing video coding methods fail to effectively utilize temporal redundancy between frames, leading to suboptimal coding efficiency.
The method involves predicting temporal motion information for non-adjacent positions in a current block and constructing a candidate motion information list using spatial and temporal motion information, which is then used for video encoding and decoding.
Improves coding efficiency by leveraging temporal correlations between frames, resulting in more effective video compression and transmission.
Smart Images

Figure 0007866062000001 
Figure 0007866062000002 
Figure 0007866062000003
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to video technology, but are not limited thereto. More specifically, they relate to a method for predicting temporal motion information, a method for constructing a candidate motion information list, a corresponding apparatus, and a system.
Background Art
[0002] Currently, a block-based hybrid coding framework is used in general-purpose video coding (including encoding and decoding) standards. Each image, sub-image, or frame in video is divided into a largest coding unit (LCU) or coding tree unit (CTU) of the same size square (e.g., 128x128, 64x64, etc.). Each largest coding unit or coding tree unit can be divided into rectangular coding units (CU) based on rules. Coding units can also be divided into prediction units (PU), transform units (TU), etc. The hybrid coding framework includes modules such as prediction, transform, quantization, entropy coding, and in-loop filtering. The prediction module includes intra-prediction and inter-prediction. Inter-prediction includes motion estimation and motion compensation. Because there is a strong correlation between adjacent samples within a single frame of a video, video coding techniques utilize intra-prediction methods to eliminate spatial redundancy between adjacent samples. Similarly, because there is a strong similarity between adjacent frames in a video, video coding techniques can improve coding efficiency by utilizing inter-prediction methods to eliminate temporal redundancy between adjacent frames.
[0003] However, the coding efficiency of existing interpretation methods still needs improvement. [Overview of the Initiative]
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of protection of the claims.
[0005] One embodiment of the present disclosure provides a time motion information prediction method. The time motion information prediction method includes the following: determining at least one non-adjacent position in the current block; and determining first time motion information of the current block based on motion information of at least one non-adjacent position in the coded image.
[0006] One embodiment of this disclosure further provides a method for constructing a candidate motion information list. This method for constructing a candidate motion information list includes the following: spatial motion information prediction and temporal motion information prediction are performed on the current block to determine the spatial motion information and temporal motion information of the current block. The spatial motion information and temporal motion information are added to the candidate motion information list of the current block in a set order. Temporal motion information prediction utilizes the temporal motion information prediction method described in any one embodiment of this disclosure. Temporal motion information includes first temporal motion information.
[0007] One embodiment of the present disclosure further provides a video encoding method. The video encoding method includes the following: constructing a candidate motion information list for the current block based on the method for constructing a candidate motion information list described in any one embodiment of the present disclosure; selecting one or more candidate motion information from the candidate motion information list and recording the index of the selected candidate motion information; determining a predicted block for the current block based on the selected candidate motion information; encoding the current block based on the predicted block; and encoding the index of the candidate motion information.
[0008] One embodiment of the present disclosure further provides a video decoding method. The video decoding method constructs a candidate motion information list for the current block based on a method for constructing a candidate motion information list described in any one embodiment of the present disclosure. One or more candidate motion information are selected from the candidate motion information list based on the index of the candidate motion information for the current block obtained by decoding. A predicted block for the current block is determined based on the selected candidate motion information, and the current block is reconstructed based on the predicted block.
[0009] One embodiment of the present disclosure further provides a time motion information prediction device. The time motion information prediction device comprises a processor and a memory in which a computer program is stored. When the processor executes the computer program, it executes a time motion information prediction method described in any one embodiment of the present disclosure.
[0010] One embodiment of the present disclosure further provides an apparatus for constructing a candidate motion information list. The apparatus for constructing a candidate motion information list comprises a processor and memory in which a computer program is stored. When the processor executes the computer program, it performs a method for constructing a candidate motion information list as described in any one embodiment of the present disclosure.
[0011] One embodiment of the present disclosure further provides a video encoding apparatus, which comprises a processor and memory storing a computer program. When the processor executes the computer program, it executes a video encoding method described in any one embodiment of the present disclosure.
[0012] One embodiment of the present disclosure further provides a video decoding apparatus. The video decoding apparatus comprises a processor and memory in which a computer program is stored. When the processor executes the computer program, it executes a video decoding method described in any one embodiment of the present disclosure.
[0013] One embodiment of the present disclosure further provides a video coding system, which includes a video encoding device described in any one embodiment of the present disclosure and a video decoding device described in any one embodiment of the present disclosure.
[0014] One embodiment of the present disclosure further provides a non-temporary computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, a method described in any one embodiment of the present disclosure is performed.
[0015] One embodiment of the present disclosure further provides a bitstream, which is generated by a video encoding method described in any one embodiment of the present disclosure.
[0016] After reading and understanding the drawings and detailed explanations, other aspects can be understood. [Brief explanation of the drawing]
[0017] The drawings are provided to illustrate embodiments of the disclosure and are part of the specification. Both the drawings and embodiments of the disclosure are used to illustrate the technical proposal of the disclosure and are not intended to limit the technical proposal of the disclosure. [Figure 1A] Figure 1A is a schematic diagram showing a coding system according to one embodiment of the present disclosure. [Figure 1B] Figure 1B shows the structure of a video encoding device according to one embodiment of the present disclosure. [Figure 1C] Figure 1C shows the structure of a video decoding device according to one embodiment of the present disclosure. [Figure 2] Figure 2 is a schematic diagram showing the reference relationships between the current block and the bidirectional reference blocks. [Figure 3] Figure 3 is a schematic diagram showing the adjacent positions of the current block, which are used to derive spatial motion information and temporal motion information. [Figure 4] FIG. 4 is a schematic diagram showing the calculation of a motion vector from a current block to a reference block based on a motion vector from a collocated block to a collocated reference block. [Figure 5] FIG. 5 is an image in which there is motion from right to left. [Figure 6] FIG. 6 is an image in which there is motion from right to left. [Figure 7] FIG. 7 is an image in which there is motion from right to left. [Figure 8] FIG. 8 is a flowchart showing a method for predicting temporal motion information according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a flowchart showing a method for constructing a candidate motion information list according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a schematic diagram showing an example of non-adjacent positions used to derive temporal motion information according to an embodiment of the present disclosure. [Figure 11] FIG. 11 is a schematic diagram showing another example of non-adjacent positions used to derive temporal motion information according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a schematic diagram showing various positions and their order used to derive spatial motion information and temporal motion information according to an embodiment of the present disclosure. [Figure 13] FIG. 13 is a schematic diagram showing another example of non-adjacent positions used to derive temporal motion information according to an embodiment of the present disclosure. [Figure 14] FIG. 14 is a flowchart showing a video encoding method according to an embodiment of the present disclosure. [Figure 15] FIG. 15 is a flowchart showing a video decoding method according to an embodiment of the present disclosure. [Figure 16] FIG. 16 is a schematic diagram showing a temporal motion information prediction apparatus according to an embodiment of the present disclosure.
MODE FOR CARRYING OUT THE INVENTION
[0018] While this disclosure describes several embodiments, the description is illustrative and not limiting, and it will be apparent to those skilled in the art that more examples and embodiments may be included within the scope of the embodiments described herein.
[0019] In the descriptions of this disclosure, terms such as “exemplary” or “for example” mean “example, illustration, or explanation.” No embodiment described “for example” or “exemplary” in this disclosure should be construed as superior to other embodiments. The terms “and / or” in this specification describe a relationship between related subjects and indicate that there are three types of relationships. For example, A and / or B indicates three situations: A alone exists, A and B exist simultaneously, or B alone exists. “Multiple” means two or more things. In addition, terms such as “first” and “second” are used to distinguish those whose function and role are substantially the same or similar in order to clearly describe the technical ideas of the embodiments of this disclosure. A person skilled in the art will understand that terms such as “first” and “second” do not limit the number or order of execution, and terms such as “first” and “second” do not necessarily limit them to being different.
[0020] In describing representative exemplary embodiments, methods and / or processes may be presented as a specific sequence of steps. However, such methods or processes do not depend on a specific sequence of steps described herein, and should not be limited to the steps in a specific sequence described. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, a specific sequence of steps described in the specification should not be construed as a limitation of the claims. Furthermore, the claims corresponding to methods and / or processes should not be limited to being performed in the order described. Those skilled in the art will readily understand that these sequences may be changed and still remain within the spirit and scope of the embodiments of this disclosure.
[0021] In this specification, a video image is abbreviated as "image," and such image includes the video image and a portion of the video image. The portion of the video image may be, for example, a subimage separated from the video image, or a slice, slice segment, etc., separated from the video image.
[0022] In this specification, the motion information derived by motion information prediction used to represent a prediction operation includes reference picture information and motion vector (MV) information. In related video standards, the same prediction operation may be represented by a "motion vector predictor," and the derived information includes not only motion vector information but also reference picture information. A "motion vector predictor" can also be understood as the motion information prediction used to represent a prediction operation in this specification.
[0023] In this specification, "temporal motion information" refers to motion information obtained by a temporal motion information predictor (or temporal motion vector predictor). Temporal motion information predictor is sometimes referred to as "temporal motion vector prediction" in standards. In this specification, "spatial motion information" refers to motion information obtained by a spatial motion information predictor (or spatial motion vector predictor). Spatial motion information predictor is sometimes referred to as "spatial motion vector prediction" in standards. In this specification, "motion information" as a prediction result is also referred to as "motion vector prediction" in some standards.
[0024] In this specification, a non-adjacent location of the current block is a location whose coordinates are not adjacent to any sample within the current block's range. An adjacent location of the current block is a location whose coordinates are adjacent to at least one sample within the current block's range.
[0025] In this specification, the current block may be the current coding unit (CU) or the current prediction unit (PU), etc. The current image is the image in which the current block is located, and the current image sequence is the sequence of images in which the current block is located.
[0026] Figure 1A is a block diagram illustrating a video coding system applicable to embodiments of the present disclosure. As shown in the figure, the system is divided into an encoding device 1 and a decoding device 2. The encoding device 1 generates a bitstream. The decoding device 2 can decode the bitstream. The encoding device 1 and the decoding device 2 may comprise one or more processors and memory coupled to one or more processors. Examples of memory include random access memory, electrically erasable programmable read-only memory, flash memory, or other media. The encoding device 1 and the decoding device 2 can be implemented in a variety of devices, such as desktop computers, mobile computing devices, notebook computers, tablet computers, set-top boxes (STBs), televisions, cameras, display devices, digital media players, in-vehicle computers, or other similar devices.
[0027] The decoding device 2 can receive a bitstream from the encoding device 1 via link 3. Link 3 includes one or more media or devices that can move the bitstream from the encoding device 1 to the decoding device 2. In one example, link 3 includes one or more communication media that enable the encoding device 1 to directly transmit the bitstream to the decoding device 2. The encoding device 1 can modulate the bitstream according to a communication standard (e.g., a wireless communication protocol) and transmit the modulated bitstream to the decoding device 2. The one or more communication media include wireless communication media and / or wired communication media, and include, for example, a radio frequency (RF) spectrum or one or more physical transmission lines. The one or more communication media can form part of a packet-based network, which is, for example, a local area network (LAN), a wide area network (WAN), or a global network (e.g., the Internet). The one or more of the above communication media may include a router, switch, base station, or other device that facilitates communication from the encoding device 1 to the decoding device 2. In another example, the bitstream may also be output to a storage device from the output interface 15. The decoding device 2 can read the data stored from the storage device via streaming or download.
[0028] In the example shown in Figure 1A, the encoding device 1 is the data source 11, Encoding device13, and an output interface 15 are included. In some examples, the data source 11 includes a video capture device (e.g., a camera), an archive containing previously captured data, a feeding interface for receiving data from a content provider, a computer graphics system for generating data, or a combination of these sources. Encoding device 13 can encode data from data source 11 and output it to output interface 15. Output interface 15 may include at least one of a tuner, a modem, and a transmitter.
[0029] In the example shown in Figure 1A, the decoding device 2 has an input interface 21, Decoding device This includes 23 and a display device 25. In some examples, the input interface 21 includes at least one of a receiver and a modem. The input interface 21 can receive a bitstream via link 3 or from a storage device. Decoding device 23 decodes the received bitstream. The display device 25 is used to display the decoded data. The display device 25 may be integrated with other components of the decoding device 2 or it may be provided separately. The display device 25 may be, for example, a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device. In other examples, the decoding device 2 may not include the display device 25, or it may include other devices or equipment to which the decoded data can be applied.
[0030] Based on the video coding system shown in Figure 1A, video compression can be achieved using various video coding methods. International video coding standards include H.264 / Advanced Video Coding (AVC), H.265 / High Efficiency Video Coding (HEVC), H.266 / Versatile Video Coding (VVC), Moving Picture Experts Group (MPEG), AOM (Alliance for Open Media), Audio Video Standard (AVS), and other arbitrary standards that are extensions or customized versions of these standards. These standards use video compression technology to reduce the amount of data transmitted and stored, thereby achieving more efficient video coding and transmission storage. All of the above video coding standards use a block-based mixed coding scheme. Embodiments of this disclosure apply to, but are not limited to, the basic flow of video coding in the block-based mixed coding framework.
[0031] Figure 1B is a block diagram of an exemplary video encoding device that may be used in embodiments of the present disclosure.
[0032] As shown in the figure, the video encoding device 1000 includes a prediction processing unit 1100, a splitting unit 1101, a residual generation unit 1102, a conversion processing unit 1104, a quantization unit 1106, an inverse quantization unit 1108, an inverse conversion processing unit 1110, a reconstruction unit 1112, a filter unit 1113, a decoded image buffer 1114, an image resolution adjustment unit 1115, and an entropy encoding unit 1116. The prediction processing unit 1100 is an interpretation processing unit 1121 This includes an intra-predictive processing unit 1126. Video encoding device 1000It may also include more, fewer, or different functional components compared to this example.
[0033] The splitting unit 1101 works in cooperation with the prediction processing unit 1100 to split the received video data into slices, CTUs, or other relatively large units. The video data received by the splitting unit 1101 may be a video sequence containing video frames such as I-frames, P-frames, and B-frames.
[0034] The prediction processing unit 1100 can divide the CTU into CUs and perform intra-predictive coding or inter-predictive coding on the CUs. When performing intra-predictive and inter-predictive coding on a CU, the CU can be divided into one or more prediction units (PUs).
[0035] The interpretation processing unit 1121 can perform interpretation on the PU and generate prediction data for the PU. This prediction data includes the prediction block of the PU, the motion information of the PU, and various syntax elements.
[0036] The intra-prediction processing unit 1126 can perform intra-prediction on the PU and generate prediction data for the PU. This prediction data for the PU can include prediction blocks and various syntax elements for the PU.
[0037] The residual generation unit 1102 can generate residual blocks of the CU by subtracting the predicted blocks of the PU obtained by dividing the CU from the original blocks of the CU.
[0038] The transformation processing unit 1104 can divide the CU into one or more Transform Units (TUs). The residual blocks associated with the TUs are subblocks obtained by dividing the residual blocks of the CU. By applying one or more transformations to the residual blocks associated with the TUs, coefficient blocks associated with the TUs are generated.
[0039] The quantization unit 1106 can quantize the coefficients within a coefficient block based on a selected quantization parameter (QP). The degree of quantization of the coefficient block can be adjusted by adjusting the QP value.
[0040] The inverse quantization unit 1108 and the inverse transformation processing unit 1110 can apply inverse quantization and inverse transformation, respectively, to the coefficient block to obtain a reconstructed residual block related to the TU.
[0041] The reconstruction unit 1112 can generate a reconstruction block of the CU by adding the reconstruction residual block and the prediction block generated by the prediction processing unit 1100.
[0042] The filter unit 1113 performs in-loop filtering on the reconstructed block and then stores it as a reference image in the decoded image buffer 1114. The intra-prediction processing unit 1126 can extract reference images of blocks adjacent to the PU from the decoded image buffer 1114 and perform intra-prediction. The inter-prediction processing unit 1121 can perform inter-prediction on the PU of the current image using the reference images of the image before they were cached in the decoded image buffer 1114.
[0043] The image resolution adjustment unit 1115 resamples the reference image stored in the decoded image buffer 1114. Resampling may include upsampling and / or downsampling. The resulting reference images with multiple resolutions are stored in the decoded image buffer 1114.
[0044] The entropy encoding unit 1116 processes the received data (e.g., syntax elements, quantized). coefficient The entropy encoding operation is performed on blocks, motion information, etc.
[0045] Figure 1C is a block diagram of an exemplary video decoding device that may be used in embodiments of the present disclosure.
[0046] As shown in the figure, the video decoding device 101 includes an entropy decoding unit 150, a prediction processing unit 152, an inverse quantization unit 154, and an inverse transformation processing unit. 155 This includes a reconstruction unit 158 (shown in the figure as a circle with a +), a filter unit 159, and an image buffer 160. In other embodiments, Video decoding device 101 It may include more, fewer, or different functional components.
[0047] The entropy decoding unit 150 performs entropy decoding on the received bitstream and can extract information such as syntax elements, quantized coefficient blocks, and PU motion information. Prediction processing unit 152, inverse quantization unit 154, inverse transformation processing unit 155 The reconstruction unit 158 and the filter unit 159 can all perform corresponding operations based on the syntax elements extracted from the bitstream.
[0048] As a functional component that performs reconstruction operations, the inverse quantization unit 154 can inverse quantize the coefficient blocks associated with the quantized TU. Inverse Transform Processing Unit 155To generate the reconstructed residual block of TU, one or more inverse transforms can be applied to the inversely quantized coefficient block.
[0049] The prediction processing unit 152 includes an inter-prediction processing unit 162 and an intra-prediction processing unit 164. When intra-prediction coding is used in the PU, the intra-prediction processing unit 164 determines the intra-prediction mode of the PU based on the syntax elements analyzed from the bitstream, and the determined intra-prediction mode and the image buffer 160 Based on the reconfigured reference information of the adjacent PUs obtained from the interprediction unit 162, intraprediction can be performed to generate predicted blocks for the PUs. When interprediction coding is used for the PUs, the interprediction processing unit 162 can determine one or more reference blocks of the PUs based on the motion information of the PUs and the corresponding syntax elements, and generate predicted blocks for the PUs based on the reference blocks.
[0050] The reconstruction unit 158 can obtain a reconstruction block of the CU based on the reconstruction residual block associated with the TU and the prediction block of the PU generated by the prediction processing unit 152 (i.e., intra-prediction data or inter-prediction data).
[0051] The filter unit 159 can perform in-loop filtering on the reconstruction block of the CU to obtain a reconstructed image. The reconstructed image is stored in the image buffer 160. The image buffer 160 can provide a reference image for subsequent motion compensation, intra-prediction, inter-prediction, etc., and can also output the reconstructed video data as decoded video data and display it on a display device.
[0052] Display device 25 This could be, for example, a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or other types of display devices. In another example, the decoding side is: Display device 25It does not have to include, and instead includes other devices to which the decoded data can be applied.
[0053] The basic flow of video coding is as follows: On the encoding side, one image is divided into blocks, intra-prediction or inter-prediction is performed on the current block to generate a predicted block of the current block, the predicted block is obtained by subtracting the predicted block from the original block of the current block, the residual block is obtained by transforming and quantizing the residual block to obtain a quantization coefficient matrix, and the quantization coefficient matrix is entropy coded and output to a bitstream. On the decoding side, intra-prediction or inter-prediction is performed on the current block to generate a predicted block of the current block, while the bitstream is decoded to obtain a quantization coefficient matrix, the quantization coefficient matrix is inversely quantized and inversely transformed to obtain a residual block, and the predicted block and residual block are added to obtain a reconstructed block. The reconstructed block forms a reconstructed image. The reconstructed image is in-loop filtered based on the image or blocks to obtain the decoded image. On the encoding side, a similar process to that on the decoding side is required to obtain the decoded image. The decoded image can be a reference image for inter-prediction of subsequent images. Block partitioning information determined on the encoding side, as well as mode information or parameter information such as prediction, transformation, quantization, entropy coding, and in-loop filtering, are output to the bitstream as needed. The decoding side analyzes the existing information to determine the same block partitioning information, prediction, transformation, quantization, entropy coding, and in-loop filtering mode or parameter information as the encoding side. This ensures that the decoded image obtained on the encoding side and the decoded image obtained on the decoding side are the same. The decoded image obtained on the encoding side is usually also called the reconstructed image. During prediction, the current block can be divided into prediction units. The above is the basic flow of a video codec (including encoder and decoder) in a block-based mixed coding framework. As technology advances, some modules or steps in this framework or flow may be optimized.
[0054] Movement and movement information
[0055] Video is composed of images. To make a video appear smooth, it contains tens to hundreds of frames per second. For example, it may contain 24 frames per second, 30 frames per second, 50 frames per second, 60 frames per second, 120 frames per second, and so on. Thus, video has a very significant temporal redundancy. In other words, there is a lot of temporal correlation. Therefore, interpretation typically uses temporal correlation in "motion". One very simple "motion" model is that an object is at a certain position in the image corresponding to a certain time, and after a certain amount of time, it is translated to a different position in the image corresponding to that time. This is the most basic and commonly used translation motion in video coding. Interpretation uses motion information to represent "motion". Basic motion information includes reference picture (e.g., reference frame) information and motion vector (MV) information. The codec determines the reference image based on the reference image information and determines the coordinates of the reference block based on the motion vector information and the coordinates of the current block. In the reference image, the coordinates of the reference block are used to determine the reference block. Motion in video is not always this simple. Even motion considered as translation has subtle changes over time, such as subtle deformations, changes in brightness, and changes in noise. To achieve better prediction effects, multiple reference blocks can be used to predict the current block. For example, in currently commonly used bidirectional prediction, two reference blocks are used to predict the current block. The two reference blocks can be one forward reference block and one backward reference block. It is also permitted that both reference blocks are forward reference blocks, or both are backward reference blocks."Forward" means that the time corresponding to the reference image is before the current image, and "backward" means that the time corresponding to the reference image is after the current image. In other words, "forward" means that the position of the reference image in the video is before the current image, and "backward" means that the position of the reference image in the video is after the current image. In other words, "forward" means that the POC (picture order count) of the reference image is smaller than the POC of the current image, and "backward" means that the POC of the reference image is larger than the POC of the current image. Future video coding standards may support prediction with more reference blocks. One simple way to generate a prediction block using two reference blocks is to obtain the prediction block by averaging the sample values of the positions corresponding to the two reference blocks. To obtain better prediction results, weighted averaging such as BCW (Bi-prediction with CU-level weight), currently used in VVC, can also be used. Geometric Partitioning Mode (GPM) within VVC can also be understood as a special type of bidirectional prediction. To utilize bidirectional prediction, it is naturally necessary to find two reference blocks. Thus, two sets of reference image information and motion vector information are required. Each set can be understood as a single unidirectional motion sequence.
[0056] Motion within a video can include not only simple translation but also scaling, rotation, distortion, and various other forms. Combining the two sets mentioned above results in a single bidirectional motion information. In practical implementation, the same data structure can be used for both unidirectional and bidirectional motion information. In bidirectional motion information, both sets of reference image information and motion vector information are valid, while in unidirectional motion information, one set of reference image information and motion vector information is invalid. Motion can be complex. VVC simulates some simple motions using an affine model. The affine model in VVC uses two or three control points, and a linear model is employed based on these control points to derive the motion vectors of each subblock within the current block. The reason only motion vectors, and not motion information, are mentioned here is that all motion vectors point to the same reference image. In normal translational motion, the "whole block" is found from the reference image, whereas in affine, it can be understood as finding a set of non-adjacent "subblocks" from the reference image. All of the above fall within the realm of unidirectional prediction. Affine motion can also enable bidirectional prediction and prediction using more "reference blocks." Here, a reference block is a combination of subblocks. Specifically, in the data structure of affine motion information, one unidirectional motion information can include one reference image and two to three motion vector information, or it can include two to three sets of reference image information and motion vector information, but these reference image information are the same.
[0057] Of course, motion information can include not only basic reference image information and motion vector information, but also additional information such as whether or not BCW is used, and the BCW index.
[0058] Reference image (reference frame)
[0059] When parallel processing is not considered, video is processed image by image. Images that have already been coded (including encoding and decoding) are stored in a buffer and serve as reference images for subsequently coded images. Current coding standards include a reference image management method. This method is used to manage which images can be reference images for the current image, the indices of these reference images, which images need to be stored in the buffer, and which images can be removed from the buffer when they are no longer reference images.
[0060] Based on the order in which images are coded, currently commonly used scenarios can be divided into two types: random access (RA) and low delay (LD). In low delay scenarios, the order of image display and the order of coding are the same, while in random access scenarios, the order of image display and the order of coding can be different. Generally speaking, in low delay scenarios, images are coded one by one according to the order of the video itself. In random access scenarios, the order of the video itself can be disrupted, with some images not being coded initially, some images being coded first, and then the skipped images being coded. An advantage of random access is that some images can reference both the previous and subsequent reference images, allowing for improved compression efficiency by effectively utilizing "motion."
[0061] Figure 2 shows a typical Group of Picture (GOP) structure for RA. In the figure, P-frames (Predictive Frames) are frames that can only perform unidirectional (forward) prediction, while B-frames (Bi-predictive Frames) are frames that can perform bidirectional prediction. This restriction on reference relationships can also be applied to levels below the image level, for example, P-slices and B-slices that are divided at the slice level.
[0062] The arrows in Figure 2 indicate reference relationships. I-frames do not require a reference image. After encoding an I-frame with a POC of 0, we encode a P-frame with a POC of 4. When encoding the P-frame with a POC of 4, we can reference the I-frame with a POC of 0. Next, we encode a B-frame with a POC of 2. When encoding the B-frame with a POC of 2, we can reference the I-frame with a POC of 0 and the P-frame with a POC of 4.
[0063] The codec manages reference images using a reference image list. VVC supports two reference image lists, denoted as RPL0 and RPL1, where RPL is an abbreviation for Reference Picture List. In VVC, P slices can only use RPL0, while B slices can use both RPL0 and RPL1. For each slice, each reference image list contains several reference images, and the codec uses the reference image index to find a specific reference image. In VVC, motion information is represented by the reference image index and motion vector. For example, for the above bidirectional motion information, VVC uses the reference image index refIdxL0 and the motion vector mvL0 corresponding to RPL0, and the reference image index refIdxL1 and the motion vector corresponding to RPL1. mvL1 The following are used. The reference image index corresponding to RPL0 and the reference image index corresponding to RPL1 can be understood as the above reference image information. In VVC, two flags indicate whether or not motion information corresponding to RPL0 is used, and RPL1 These two flags, predFlagL0 and predFlagL1, respectively, indicate whether or not the corresponding motion information is used. Furthermore, predFlagL0 and predFlagL1 can be understood as indicating whether or not the above unidirectional motion information is "valid."
[0064] The accuracy of motion vectors is not limited to integer pixels; VVC allows for predictions with accuracy of 1 / 2, 1 / 4, 1 / 8, and 1 / 16 pixels. Prediction at fractional pixels requires interpolation to integer pixels. This results in finer motion vectors and improved prediction quality.
[0065] Motion information prediction (motion vector prediction or motion information prediction)
[0066] For the current block, motion information can be used to find the reference block from the reference image, and based on the reference block, the predicted block for the current block can be determined.
[0067] The motion information used for a block can usually be predicted using several pieces of related information, which can be called motion information prediction or motion vector prediction. For example, motion information used for adjacent coded blocks around the current block in the current image (e.g., frame or slice) can be used because there is a strong correlation between adjacent blocks. It is also possible to use motion information used for non-adjacent coded blocks around the current block in the current image because there is some correlation between blocks in a relatively close area and the current block, even if they are not adjacent. This method of predicting motion information using motion information used for coded blocks around the current block in the current image is generally called spatial motion information prediction. In addition to coded blocks around the current block, motion information can also be used for blocks related to the current block's position in the coded image to predict motion information for the current block; this is generally called temporal motion information prediction. Furthermore, the motion information of coded blocks can also be maintained in a list according to the coding order. Generally, the list contains multiple different pieces of recently coded motion information, and motion information in this list is used to predict motion information for the current block. This is generally called history-based motion information prediction. It can be easily understood that spatial motion information is derived using motion information from the same image as the current block, while temporal motion information is derived using motion information from a different image than the current block.
[0068] To use spatial motion information prediction, it is necessary to store the motion information of the currently coded blocks in the image (or slice). Generally, a minimum storage unit is set, for example, a 4x4 size, but it may be 8x8 or other sizes. Each time the codec codes a block, it stores motion information for all the minimum storage units corresponding to that block. To find motion information for blocks around the current block, the minimum storage unit can be found based on its coordinate position and its motion information can be obtained. Similarly, to use temporal motion information prediction, it is necessary to store the motion information of the coded image (or slice). Generally, a minimum storage unit is also set. It may be the same size as the spatial motion information storage unit, or it may be a different size, as determined by the standard. To find motion information for a block in the above image (or slice) for the current block, the minimum storage unit can be found based on its coordinate position and its motion information can be obtained. Due to limitations in storage space or implementation complexity, it may only be possible to obtain temporal or spatial motion information for the current block within a certain range of coordinate positions.
[0069] Motion information prediction can yield one or more motion information. When obtaining multiple motion information, it is usually necessary to specify which one or more motion information to select, or to specify which one or more motion information to select according to some predetermined rules. For example, two motion information may be needed for a GPM within a VVC, and for example, in a sub-block based TMVP (abbreviated as SbTMVP), one motion information may be needed for each sub-block. TMVP refers to a temporal motion vector predictor. In this specification, TMVP is used to represent a temporal motion prediction based on adjacent positions.
[0070] To use motion information obtained through prediction, specific motion information obtained through prediction can be directly used as the motion information for the current block. One example is the merge function within HEVC. Alternatively, new motion information can be obtained by adding a motion vector difference (MVD) to specific motion information obtained through prediction. Ideally, predicted motion information should be as close to actual motion information as possible, but since it is not always possible to guarantee that motion information prediction is accurate, motion vector differences can be used to obtain more accurate motion information. A new representation method for motion vector differences has been added to VVC, and a combination of motion vector prediction and this new representation method for motion vector differences can be used in merge mode, which is called MMVD (merge with MVD). In other words, motion information prediction can be used directly, or it can be used in combination with other methods.
[0071] The following example illustrates how to construct a list of merge movement information candidates within a VVC.
[0072] The merge motion information candidate list is denoted as mergeCandList. When constructing mergeCandList, first, the spatial motion information prediction based on positions 1-5 in Figure 3 is checked, and then the temporal motion information prediction is checked. When checking the temporal motion information prediction, if position 6 is available, the temporal motion information prediction is derived (drive) based on position 6. Derivation is also called deriving. That is, the temporal motion information is obtained by calculating step by step based on position 6. Position 6 being available means that position 6 does not cross the boundary of an image or sub-image and is in the same CTU row as the current block; please refer to the standard text for details. If position 6 is not available, the temporal motion information prediction is derived based on position 7. Position 6 can be represented by coordinates (xColBr, yColBr), and position 7 can be represented by coordinates (xColCtr, yColCtr).
[0073] One specific method for determining spatial motion information is as follows. As an example, let's derive the spatial motion information for position 2 in Figure 3. Position 2 is denoted as B1, or the block in which position 2 is located is denoted as B1. The available flag availableFlagB1, the reference image index refIdxLXB1, the reference image list usage flag predFlagLXB1, and the motion vector mvLXB1 are derived. The above availableFlagB1 can be used to indicate whether or not the spatial motion information for B1 is available. The reference image index refIdxLXB1, the reference image list usage flag predFlagLXB1, and the motion vector mvLXB1 together constitute the motion information. X = 0..1.
[0074] (xCb, yCb) is the coordinate of the top-left corner of the current block relative to the top-left corner of the current image, cbWidth is the width of the current block, and cbHeight is the height of the current block.
[0075] Set the coordinates (xNbB1, yNbB1) in the adjacent block B1 to (xCb + cbWidth - 1, yCb - 1). Determine whether the block where (xNbB1, yNbB1) is located is available. One way to determine this is as follows: If the block where (xNbB1, yNbB1) is located is already coded and intercoded, then the block is available; otherwise, the block is unavailable. Of course, there may be additional conditions for determination. For example, if xCb >> Log2ParMrgLevel is equal to xNbB1>>Log2ParMrgLevel and yCb >> Log2ParMrgLevel is equal to yNbB1>>Log2ParMrgLevel, then the block is unavailable. Log2ParMrgLevel is a variable determined by the sequence level parameter. ">>" indicates a right shift operation.
[0076] --If the block in which (xNbB1, yNbB1) is located is unavailable, the value of availableFlagB1 is set to 0, the horizontal and vertical components of mvLXB1 are both set to 0, the value of refIdxLXB1 is set to -1, the value of predFlagLXB1 is set to 0, and X = 0..1. --If not, the value of availableFlagB1 is set to 1 and the following values are assigned. mvLXB1 = MvLX[ xNbB1][ yNbB1] refIdxLXB1= RefIdxLX[ xNbB1][ yNbB1] predFlagLXB1= PredFlagLX[ xNbB1][ yNbB1]
[0077] One possible understanding is that motion information at (xNbB1, yNbB1) can be used to predict spatial motion information based on position B1, and spatial motion information predictions based on other positions can be derived using a similar method.
[0078] Derivation of time-motion information
[0079] As shown in Figure 4, currPic is the current image, currCb is the current block, currRef is the reference image for the time motion information of the current block, colCb is the collated block, colPic is the image in which the collated block is located, and colRef is the reference image for the motion information used for the collated block. In deriving the time motion information, first, colCb is found from colPic based on one positional piece of information, and the motion information of colCb is found. The motion information of colCb shown in Figure 4 is the motion vector from colPic to colRef (shown by the solid line). The time motion information is obtained by scaling the motion vector shown by the solid line to the motion vector from currPic to currRef (i.e., the motion vector shown by the dotted line). tb is a variable determined based on the difference between the POC of currPic and the POC of currRef, and td is a variable determined based on the difference between the POC of colPic and the POC of colRef. The motion vector is scaled based on tb and td. The diagram illustrates unidirectional motion information, or the scaling of a single motion vector. It can be understood that time-motion information can utilize bidirectional motion information.
[0080] One specific method for deriving time-motion information is as follows. As an example, let's derive the time-motion information for position 6 in Figure 3. Position 6 is denoted as Col, or the block in the reference image where position 6 is located is denoted as Col. Col is an abbreviation for collocated, and this block is called a collocated block, and this reference image is called a collocated image. Derive the available flag availableFlagCol, the reference image index refIdxLXCol, the reference image list usage flag predFlagLXCol, and the motion vector mvLXCol. The above availableFlagCol can be used to indicate whether or not the time-motion information of Col is available. The reference image index refIdxLXCol, the reference image list usage flag predFlagLXCol, and the motion vector mvLXCol together constitute the motion information. X = 0..1.
[0081] The coordinates (xColBr, yColBr) at position 6 are denoted as (xCb + cbWidth, yCb + cbHeight). If the coordinates (xColBr, yColBr) satisfy the requirements, for example, if they do not exceed the range of the current image or sub-image, or the range of the CTU row in which the current block is located, the coordinates (xColCb, yColCb) of the collated block are denoted as ((xColBr>>3)<<3, (yColBr>>3)<<3). The current block is denoted as currCb, and the collated block on the collated image ColPic is colCb, and colCb is the block that covers (xColCb, yColCb). currPic is the current image. When calculating coordinates, the reason for first right-shifting (>>) by 3 bits and then left-shifting (<<) by 3 bits is that the motion information within the collated image in this example is stored in a minimum memory unit of 8x8 (to save buffer space, the granularity of caching the reference image motion information can be relatively coarse). By first right-shifting by 3 bits and then left-shifting by 3 bits, the last 3 bits become 0. For example, the binary value of 10 is 1010, and by first right-shifting by 3 bits and then left-shifting by 3 bits, it becomes 8, and the binary value is 1000. In different embodiments, the coordinate conversion may differ.
[0082] Here, for the sake of explanation, we make several simplified assumptions. In this example, both the motion information of the collated block and the time motion information to be derived utilize only one reference image list L0. In this example, all X's that appear below are 0. However, it should be noted that in practice, these two motion information may utilize two reference image lists in some scenarios. Also, the forward motion information or backward motion information of the time motion information derived based on the forward motion information or backward motion information of the collated block can have multiple combinations, and these are not limited in this disclosure, so only the simplest example will be described. For simplicity, refIdxLX is also set to 0. refIdxLX can have multiple possible values, and these are not limited in this disclosure, so only the simplest example will be described.
[0083] mvLXCol and availableFlagCol are derived based on the following method. --If colCb is encoded in intra-mode, intra-block copy (IBC) mode, or palette mode, and both the horizontal and vertical components of mvLXCol are set to 0, and availableFlagCol is set to 0, then the collated block is considered unavailable. --Otherwise, availableFlagCol is set to 1, and the motion vector mvCol, reference image index refIdxCol, and reference image list indication listCol are derived based on the following method: mvCol, refIdxCol, and listCol constitute the motion information of the collated block used for scaling. mvCol, refIdxCol, and listCol are set to mvL0Col[ xColCb ][ yColCb ], refIdxL0Col[ xColCb ][ yColCb ], and L0, respectively.
[0084] predFlagLXCol is set to 1.
[0085] mvLXCol is derived based on the following method. --refPicList[listCol][refIdxCol] is a reference image for the motion information of the collated block colCb, i.e., colRef in the figure. RefPicList[X][refIdxLX] is a reference image for the time motion information, i.e., currRef in the figure. The POC distance colPocDiff between ColPic and refPicList[listCol][refIdxCol] is calculated, and the POC distance currPocDiff between currPic and RefPicList[X][refIdxLX] is calculated. colPocDiff = DiffPicOrderCnt( ColPic, refPicList[ listCol ][ refIdxCol ] ) currPocDiff = DiffPicOrderCnt( currPic, RefPicList[ X ][ refIdxLX ] ) If colPocDiff is equal to currPocDiff, mvLXCol = Clip3( -131072, 131071, mvCol ) Otherwise, mvLXCol is scaled based on mvCol. tx = ( 16384 + ( Abs( td ) >> 1 ) ) / td distScaleFactor = Clip3( -4096, 4095, ( tb * tx + 32 ) >> 6 ) mvLXCol = Clip3( -131072, 131071, (distScaleFactor * mvCol + 128-( distScaleFactor * mvCol >= 0 ) ) >> 8 ) Here, td = Clip3( -128, 127, colPocDiff ). tb = Clip3( -128, 127, currPocDiff ).
[0086] The `clip3` function here is a clipping function, and the values within `clip3` such as -131072, 131071, -4096, and 4095 relate to data precision; these values may differ depending on the precision specification. For details on the above calculations, please refer to the relevant standards.
[0087] When adding each candidate motion information to mergeCandList, identity and similarity checks can be performed to prevent the same or very similar motion information from being added to mergeCandList. This helps mergeCandList actually provide more candidates.
[0088] First, let's explain how to determine if two sets of motion information are the same, using motion information within a VVC as an example. Motion information within a VVC includes predFlagLX, refIdxLX, and mvLX. X = 0..1. Let's denote the two sets of motion information as mi0 and mi1, respectively. One way to determine if the two sets of motion information are the same is as follows: Whether X is 0 or 1, if mi0.predFlagLX is equal to mi1.predFlagLX, mi0.refIdxLX is equal to mi1.refIdxLX, and mi0.mvLX is equal to mi1.mvLX, then mi0 and mi1 are the same. If predFlagLX is 0, we can assume that the corresponding refIdxLX and mvLX are pre-set values, that is, they are not random values.
[0089] Alternatively, if predFlagLX is 1, a comparison is performed between refIdxLX and between mvLX. If predFlagLX is 0, there is no need to compare between refIdxLX and between mvLX. If the same reference image exists in two reference image lists, they may point to the same reference image even if each predFlagLX and each refIdxLX are different. If the corresponding mvLX is the same, the motion information may actually be the same. A reference image whose index in reference image list LX is refIdxLX is written as refPicList[LX][refIdxLX]. Generally, since there are no two identical reference images in the same reference image list, we can take the example of using two reference image lists with different unidirectional motion information. That is, one of mi0.predFlagL0 and mi0.predFlagL1 is 1 and the other is 0, and the one that is 1 is written as corresponding to list A. If one of mi1.predFlagL0 and mi1.predFlagL1 is 1 and the other is 0, then the one that is 1 corresponds to List B. If mi0.refPicList[LA][refIdxLA] is equal to mi1.refPicList[LB][refIdxLB] and mi0.mvLA is equal to mi1.mvLB, then the two motion information sets are the same. The two methods described above are a method of directly comparing each parameter and a method of comparing the derived reference images, respectively.
[0090] The following introduces that two motion information are similar. The motion information in VVC includes predFlagLX, refIdxLX, and mvLX. X = 0..1. Denote the two motion information as mi0 and mi1 respectively. One of the methods to determine whether the two motion information are similar is as follows. Whether X is 0 or 1, if mi0.predFlagLX is equal to mi1.predFlagLX, mi0.refIdxLX is equal to mi1.refIdxLX, and diff(mi0.mvLX, mi1.mvLX) < diffTh, then mi0 and mi1 are similar. If predFlagLX is 0, assume that the corresponding refIdxLX and mvLX are preset values, that is, they are not random values.
[0091] Alternatively, when predFlagLX is 1, perform the comparison between refIdxLX and the comparison between mvLX. When predFlagLX is 0, there is no need to perform the comparison between refIdxLX and the comparison between mvLX. If there are the same reference images in the two reference image lists, even if each predFlagLX is different and each refIdxLX is different, they may point to the same reference image. If the corresponding mvLX are similar, the motion information may actually be similar. The reference image with the index refIdxLX in the reference image list LX is denoted as refPicList[LX][refIdxLX]. Generally, since there are no two identical reference images in the same reference image list, take as an example the case where two unidirectional motion information use different reference image lists respectively. That is, one of mi0.predFlagL0 and mi0.predFlagL1 is 1 and the other is 0, and the one that is 1 is denoted as corresponding to list A. One of mi1.predFlagL0 and mi1.predFlagL1 is 1 and the other is 0, and the one that is 1 is denoted as corresponding to list B. When mi0.refPicList[LA][refIdxLA] is equal to mi1.refPicList[LB][refIdxLB] and diff(mi0.mvLX, mi1.mvLX) < diffTh, the two motion information are similar. diff(mi0.mvLX, mi1.mvLX) can be used to represent the difference between the two motion vectors. For example, it can be the sum of the absolute value of the difference between the horizontal components of the two motion vectors and the absolute value of the difference between the vertical components of the two motion vectors, or the maximum value of the absolute value of the difference between the horizontal components of the two motion vectors and the absolute value of the difference between the vertical components of the two motion vectors. diffTh is a threshold for determining whether they are similar. If the difference is smaller than diffTh, it is determined that the two motion vectors (information) are similar; otherwise, it is determined that the two motion vectors (information) are not similar. The above two methods are respectively the method of directly comparing each parameter and the method of comparing the derived reference images.
[0092] The order of video coding is generally from left to right and from top to bottom. For example, within a slice, coding begins with the first CTU in the top left corner, then proceeds to the right, processing one row of CTUs, and then continues from the first CTU on the left of the second row. Within a single CTU, coding is also performed from left to right and from top to bottom. This processing order ensures that relevant information is easily obtained from the left and top, and more difficult to obtain from the right and bottom. One important piece of this relevant information is motion information. Motion can refer to any one direction; it can be from left to right or from right to left. As an example of a simple scenario, in a low-latency coding arrangement, all images are coded in chronological order. Assuming an object moves from right to left, if the collated block at position 6, used to derive time-motion information, does not contain the content of this object, then it is highly likely that motion information cannot be predicted for the block at present. Furthermore, if some movements cannot be obtained from the currently coded portion of the image, and cannot be obtained from the position of the bottom-right corner adjacent to the current block, it is highly likely that this movement information cannot be predicted. Examples include the jockey moving from right to left in the image RaceHorses in Figure 5, the person moving from bottom-right to top-left in the image BasketballDrill in Figure 6, and the car moving from top-right to bottom-left in the image BQTerrace in Figure 7. In these scenarios, existing motion information prediction methods cannot accurately predict such movement information.
[0093] If an object does not appear in the coded portion, it is difficult to obtain its movement information by performing spatial motion prediction on the current block. In this case, we can rely on temporal motion prediction. Temporal motion prediction uses the position of the lower right corner adjacent to the current block. However, even under random access scenarios, motion information at a single location is often insufficient because the object in question may not be included in the collated block at that location.
[0094] This is especially true in GPM mode. In GPM mode, in some divisions, the edges of two objects are simulated. If there is one object in the upper left corner and another object in the lower right corner, it is easy to obtain relational information for the object in the upper left corner, but it is difficult to obtain relational information for the object in the lower right corner. Because the motion information of the object in the lower right corner is not obtained, it may not be possible to use GPM.
[0095] To solve this problem, time-motion information can be predicted using one or more locations that are not currently adjacent to the block.
[0096] One embodiment of this disclosure provides a method for predicting time-motion information. The method includes the following:
[0097] Step 110: Determine at least one non-adjacent location in the current block.
[0098] Step 120: Determine the first time motion information of the current block based on the motion information of at least one non-adjacent location in the coded image.
[0099] In this specification , non Time motion information prediction based on adjacent positions is called non-adjacent temporal motion information prediction, or non-adjacent temporal motion vector prediction, abbreviated as NATMVP. The use of "motion vector" rather than "motion information" is consistent with common terminology. The motion vector here still includes reference image information such as the reference image index and the reference image list usage flag. In this specification, the first time motion information is also called "time motion information determined by NATMVP."
[0100] In this specification, the motion information of the current block determined by time motion information prediction is referred to as the time motion information of the current block, and the motion information of the current block determined by spatial motion information prediction is referred to as the spatial motion information of the current block.
[0101] In one exemplary embodiment of the present disclosure, the time-motion information prediction method further includes time-motion information prediction based on adjacent locations, i.e., time-motion information prediction based on adjacent locations of the current block (abbreviated as TMVP). If time-motion information cannot be obtained from adjacent locations of the current block, time-motion information can also be obtained from within the current block.
[0102] When predicting time motion information based on a non-adjacent position of the current block, the collated block corresponding to the current block in the reference image can be determined based on the coordinates of the non-adjacent position (the collated block is a coding unit or minimum memory unit in the reference image that includes the non-adjacent position, and the reference image is considered a collated image), motion information of the collated block is obtained, and the motion vector in the motion information is scaled to obtain the time motion information of the current block. Since the time motion information is derived based on the non-adjacent position, in this specification, the non-adjacent position is also referred to as the position from which the time motion information is derived.
[0103] In this specification, predicting spatial motion information based on positions adjacent to the current block (i.e., adjacent positions to the current block) is abbreviated as SMVP (spatial motion vector predictor). Of course, it is also possible to predict spatial motion information based on positions not adjacent to the current block (i.e., non-adjacent positions to the current block), and in this specification, this prediction is abbreviated as NASMVP (non-adjacent spatial motion vector predictor).
[0104] In one exemplary embodiment of the present disclosure, at least one non-adjacent location includes one or more non-adjacent locations in the following directions: Currently, the non-adjacent position to the right of the block, Currently, the non-adjacent positions below the block, and Currently, it's in a non-adjacent position on the lower right side of the block.
[0105] As shown in Figure 10, the large square in the figure represents a 16x16 block, and the block is currently located in the center. There is a section line in the figure. (That is, intersecting lines were drawn.) The small squares are 4x4 blocks and correspond to the smallest memory unit. For ease of representation, Figure 10 uses one small square to represent non-adjacent positions whose coordinates fall within that small square.
[0106] For example, the non-adjacent position to the right of the current block refers to a non-adjacent position whose horizontal coordinate is greater than xCb + cbWidth and whose vertical coordinate is in the interval [yCb, yCb + cbHeight - 1]. The non-adjacent position below the current block refers to a non-adjacent position whose vertical coordinate is greater than yCb + cbHeight and whose horizontal coordinate is in the interval [xCb, xCb + cbWidth - 1]. The non-adjacent position to the lower right of the current block refers to a non-adjacent position whose horizontal coordinate is greater than xCb + cbWidth and whose vertical coordinate is greater than yCb + cbHeight. These can be seen by referring to the three boxes indicated by thick lines in Figure 10.
[0107] In Figure 10, the position of the lower right corner adjacent to the current block is the position used for TMVP. If the position of the lower right corner adjacent to the current block is unavailable, the position of the center of the current block is used. It is considered that the upper right and lower left positions used for NASMVP are often unavailable. In another embodiment, as shown in Figure 11, the upper right and / or lower left positions of the current block may also be used for NATMVP. In yet another embodiment, all positions around the current block may be used for NATMVP, including non-adjacent positions in at least one of the left, top, and upper left directions, and these non-adjacent positions are already coded in the reference image. In actual applications, only a few positions may be used. Note that the specific coordinates of the non-adjacent positions actually used in embodiments of this disclosure may change based on those shown, for example, by performing an operation of "+1" or "-1" on at least one of the horizontal and vertical coordinates. Reflected in the figure, a small square corresponding to one non-adjacent position may change from the lower left corner of one large square to the upper left corner of the next large square.
[0108] In one exemplary embodiment of the present disclosure, the horizontal and / or vertical distances between a non-adjacent location and a setting sample within the current block are preset fixed values, or The horizontal and / or vertical distances between a non-adjacent location and a setting sample within the current block are variable and determined based on the following parameters, or any combination thereof. Current block size, Current image sequence parameters, The current sequence level flag for the image sequence, The current image level flag, and The slice level flag for the current image.
[0109] In one example, the horizontal distance (i.e., horizontal distance) and vertical distance (i.e., vertical distance) between a non-adjacent position and a set sample within the current block can be preset. For example, the horizontal and vertical distances from position 10 in Figure 12 to the lower right corner of the current block are both 16 pixels, and the horizontal and vertical distances from position 16 in Figure 12 to the lower right corner of the current block are both 16*2 pixels, or 32 pixels. This can be inferred in this way going forward.
[0110] In one example, a non-adjacent position can be determined based on the current block's position. (xCb, yCb) is the coordinate of the top-left corner of the current block relative to the top-left corner of the current image, cbWidth is the width of the current block, and cbHeight is the height of the current block. In this example, the horizontal and / or vertical distance between the non-adjacent position and the set sample within the current block is determined according to the size of the current block. For example, the horizontal distance from position 10 in Figure 12 to the bottom-right corner of the current block is cbWidth, the vertical distance is cbHeight, and the coordinates of position 10 are (xCb + 2*cbWidth-1, yCb + 2*cbHeight-1). The horizontal distance from position 16 in Figure 12 to the bottom-right corner of the current block is 2*cbWidth, the vertical distance is 2*cbHeight, and the coordinates of position 16 are (xCb + 3*cbWidth-1, yCb + 3*cbHeight-1). The following can be inferred in this way.
[0111] In one example, the horizontal and / or vertical distances between a non-adjacent position and a set sample within the current block are determined based on the parameters of the current image sequence (e.g., resolution). For example, in a 1920×1080 current image sequence, the horizontal and vertical distances from position 10 in Figure 12 to the lower right corner of the current block are both 32 pixels, and in a 1280×720 current image sequence, 12 The horizontal and vertical distances from position 10 within the block to the bottom right corner are both 16 pixels.
[0112] For example, the horizontal and / or vertical distances between a non-adjacent position and a set sample in the current block are determined based on flags, such as the sequence level (sequence parameter set) flag, image level flag, or slice level (slice header) flag of the current image sequence. For example, a sequence level flag, image level flag, or slice level flag (which can be a 1-bit flag) being 0 means that the horizontal and vertical distances from position 10 in Figure 12 to the lower right corner of the current block are both 32 pixels, while a flag being 1 means that the horizontal and vertical distances from position 10 in Figure 12 to the lower right corner of the current block are both 16 pixels. The sequence level flag, image level flag, and slice level flag may be newly added flags or reused existing flags, and indicate values to be selected from a plurality of pre-set values that represent the horizontal and / or vertical distances between a non-adjacent position and a set sample in the current block.
[0113] In one exemplary embodiment of the present disclosure, the coded image is a reference image of the current block. Determining first temporal motion information of the current block based on motion information of at least one non-adjacent location in the coded image includes: Determining the collated block in the reference image corresponding to the non-adjacent location; Obtaining motion information of the collated block; Scale the motion vectors in the obtained motion information to obtain first temporal motion information of the current block. In this embodiment, the collated block in the reference image corresponding to the non-adjacent location is a coding unit or minimum storage unit in the reference image at that non-adjacent location. Or, if the coordinate range of one coding unit or minimum storage unit in the reference image includes the coordinates of that non-adjacent location, then that coding unit or minimum storage unit is the collated block in the reference image corresponding to that non-adjacent location.
[0114] In one exemplary embodiment of the present disclosure, any non-adjacent position lies within the range of the largest coding unit (LCU) or coding tree unit (CTU) where the current block is located, or any non-adjacent position lies within the range of the LCU row or CTU row where the current block is located. Considering storage costs, a decoding device can generally cache only a portion of the motion information stored in a collated image. Therefore, the range of motion information in a collated image that can be read for the current block may be limited. For example, for the current block, only the motion information stored in the collated image at the position corresponding to the current CTU may be read, or for the current block, only the motion information stored in the collated image at the position corresponding to the current CTU row may be read. This limitation may also apply to all of the above methods. If a position used to derive time-motion information exceeds the range available to the current block, for example, beyond the current CTU or current CTU row, the collated block at this position is unavailable, and the derivation of time-motion information at this position can be terminated.
[0115] In one exemplary embodiment of the present disclosure, the horizontal and / or vertical distances between a non-adjacent position and a set sample in the current block can also be determined based on a combination of several parameters. For example, if the sequence level flag of the current image sequence is 1 and the size of the current block is 16 × 16, then the horizontal and vertical distances are set to 16. Different combinations correspond to different horizontal and vertical distances.
[0116] In one exemplary embodiment of the present disclosure, the time-motion information prediction method further includes time-motion information prediction based on adjacent positions, i.e., including both NATMVP and TMVP.
[0117] In one exemplary embodiment of the present disclosure, if there are at least three positions in the same direction for deriving first time-motion information, the distances between adjacent positions among the at least three positions are all the same, or the distances between adjacent positions among the at least three positions are variable, with the distance between two adjacent positions increasing as they move away from the current block. That is, the distance between positions in the same direction may be fixed or variable. Referring to Figure 12, there are three positions numbered 17, 25, and 33 to the right of the current block. The distance between position 17 and position 25 is equal to the distance between position 25 and position 33, i.e., positions 17, 25, and 33 are set to be equidistant. In another example, the distance between position 25 and position 33 may be greater than the distance between position 17 and position 25.
[0118] In one exemplary embodiment of the present disclosure, at least one non-adjacent location includes one or more of the following locations: The coordinates are (xCb+2*k_cbWidth-1, yCb+2*k_cbHeight-1), (xCb+2*k_cbWidth, yCb+2*k_cbHeight-1), (xCb+2*k_cbWidth-1, yCb+2*k_cbHeight), or (xCb+2*k_cbWidth, the first position, which is yCb+2*k_cbHeight), The coordinates are (xCb+3*k_cbWidth-1, yCb+3*k_cbHeight-1), (xCb+3*k_cbWidth, yCb+3*k_cbHeight-1), (xCb+3*k_cbWidth-1, yCb+3*k_cbHeight), or (xCb+3*k_cbWidth, the second position, which is yCb+3*k_cbHeight), The coordinates are (xCb+3*k_cbWidth-1, yCb+k_cbHeight / 2-1), (xCb+3*k_cbWidth, yCb+k_cbHeight / 2-1), (xCb+3*k_cbWidth-1, yCb+k_cbHeight / 2), or (xCb+3*k_cbWidth, the third position, which is yCb+k_cbHeight / 2); The fourth position where the coordinates are (xCb + k_cbWidth / 2, yCb + 3*k_cbHeight - 1), (xCb + k_cbWidth / 2, yCb + 3*k_cbHeight), (xCb + k_cbWidth / 2 - 1, yCb + 3*k_cbHeight - 1), or (xCb + k_cbWidth / 2 - 1, yCb + 3*k_cbHeight), The fifth position where the coordinates are (xCb + 2*k_cbWidth - 1, yCb + k_cbHeight / 2 - 1), (xCb + 2*k_cbWidth, yCb + k_cbHeight / 2 - 1), (xCb + 2*k_cbWidth - 1, yCb + k_cbHeight / 2), or (xCb + 2*k_cbWidth, yCb + k_cbHeight / 2), The sixth position where the coordinates are (xCb + k_cbWidth / 2, yCb + 2*k_cbHeight - 1), (xCb + k_cbWidth / 2, yCb + 2*k_cbHeight), (xCb + k_cbWidth / 2 - 1, yCb + 2*k_cbHeight - 1), or (xCb + k_cbWidth / 2 - 1, yCb + 2*k_cbHeight), The seventh position where the coordinates are (xCb + k_cbWidth, yCb + 2*k_cbHeight), (xCb + k_cbWidth, yCb + 2*k_cbHeight - 1), (xCb + k_cbWidth - 1, yCb + 2*k_cbHeight), or (xCb + k_cbWidth - 1, yCb + 2*k_cbHeight - 1), and The eighth position where the coordinates are (xCb + 2*k_cbWidth, yCb + k_cbHeight), (xCb + 2*k_cbWidth - 1, yCb + k_cbHeight), (xCb + 2*k_cbWidth, yCb + k_cbHeight - 1), or (xCb + 2*k_cbWidth - 1, yCb + k_cbHeight - 1).
[0119] xCb is the horizontal coordinate of the top-left corner of the block, yCb is the vertical coordinate of the top-left corner of the block, k_cbWidth is the current width of the block, or half the width, or a quarter of the width, or twice the width, and k_cbHeight is the current height of the block, or half the height, or a quarter of the height, or twice the height, and "*" indicates a multiplication operation.
[0120] Referring to Figures 12 and 13, the position with coordinates (xCb+2*cbWidth-1, yCb+2*cbHeight-1) is represented as position 10 in Figure 12, and the position with coordinates (xCb+3*cbWidth-1, yCb+3*cbHeight-1) is represented as position 16 in Figure 12. Other positions are not described in detail.
[0121] In one exemplary embodiment of the present disclosure, the time motion information prediction method further includes time motion information prediction based on adjacent positions. The adjacent positions include the adjacent position of the lower right corner of the current block, and the non-adjacent positions of the current block and the adjacent position of the lower right corner of the current block are distributed in an array, the upper left corner of the array is the adjacent position of the lower right corner of the current block, and different positions in the array are distributed in different minimum memory units in the coded image.
[0122] Each position in the above array can be obtained by scanning within a set range. For example, by scanning in a specific order within the rectangular area shown in the lower right corner of Figure 13, non-adjacent positions used for predicting time motion information are determined, and time motion information is determined based on the scanned positions.
[0123] The width and height of the rectangular region in Figure 13 can be preset, for example, both width and height are 64 pixels. Alternatively, the width and height of the rectangular region can be determined based on the size of the current block, for example, the width of the rectangular region is four times the width of the current block, and the height of the rectangular region is four times the height of the current block. Alternatively, the width and height of the rectangular region can be determined based on the parameters of the current image sequence, for example, in a 1920×1080 sequence, both the width and height of this region are 64 pixels, and in a 1280×720 sequence, both the width and height of this region are 32 pixels. Alternatively, the width and height of the rectangular region can be determined based on a flag, for example, the sequence level (sequence parameter set) flag of the current sequence image, or the slice level (slice header) flag of the current image. For example, a flag of 0 means that both the width and height of this region are 64 pixels, and a flag of 1 means that both the width and height of the region are 32 pixels. The scan order may be a raster scan order, i.e., the order shown in the figure, or it may be a zig-zag scan order or any other scan order.
[0124] Motion information of the collated image is stored in a minimum memory unit, such as a 4x4 block, an 8x8 block, or a 16x16 block. Only one piece of motion information is stored in the minimum memory unit. Scanning can be performed at the granularity of each minimum memory unit. Reflected in coordinates, if the minimum memory unit is an 8x8 block, then horizontally, 8 is added to the horizontal coordinate of the next scan position, and the vertical coordinate remains unchanged. Vertically, the horizontal coordinate of the next scan position remains unchanged, and 8 is added to the vertical coordinate. Scan granularity can be related to the size of the current block; for example, horizontal granularity is equal to the width of the current block, and vertical granularity is equal to the height of the current block. Reflected in coordinates, horizontally, the width of the current block is added to the horizontal coordinate of the next scan position, and the vertical coordinate remains unchanged. Vertically, the horizontal coordinate of the next scan position remains unchanged, and the height of the current block is added to the vertical coordinate. The granularity can also be determined based on the parameters of the current image sequence. For example, in a 1920×1080 current image sequence, both the horizontal and vertical scan granularity are 16 pixels, and in a 1280×720 sequence, both the horizontal and vertical scan granularity are 32 pixels. The granularity can also be determined based on flags, such as the sequence level flag of the current image sequence, the image level flag of the current image, or the slice level flag of the current image. For example, a flag of 0 means that both the horizontal and vertical scan granularity are 16 pixels, and a flag of 1 means that both the horizontal and vertical scan granularity are 32 pixels.
[0125] During scanning, this method may limit the number of time-motion information candidates that can be added. That is, when predicting time-motion information, it is not necessarily the case that all positions within the rectangular range are scanned.
[0126] In the time motion information prediction method according to the above embodiment of this disclosure, by determining the time motion information of the current block using non-adjacent positions, it is possible to effectively complement motion information for several scenarios that cannot be covered by spatial motion information, such as movement from right to left, movement from the lower right to the upper left, and movement from bottom to top, thereby improving compression efficiency. Furthermore, in the time motion information prediction method disclosed in the above embodiment, it is proposed to extract the time motion information of the current block more efficiently by setting the direction of these non-adjacent positions relative to the current block, the distance between non-adjacent positions, etc.
[0127] As shown in Figure 9, one embodiment of the present disclosure further provides a method for constructing a candidate motion information list. The method for constructing a candidate motion information list includes the following:
[0128] Step 210: Predict spatial motion information and temporal motion information for the current block, and determine the spatial motion information and temporal motion information for the current block.
[0129] Step 220: Add the spatial motion information and temporal motion information to the current block's candidate motion information list in the set order.
[0130] Time motion information prediction utilizes a time motion information prediction method described in any one embodiment of the present disclosure. Time motion information includes first time motion information.
[0131] In one exemplary embodiment of the present disclosure, the method for constructing a list of candidate motion information is used in merge mode.
[0132] In HEVC and VVC, a list of motion information candidates in merge mode is a common scenario for using motion information prediction. In embodiments of this disclosure, spatial motion information and temporal motion information (which may also be collectively referred to as candidate motion information) can be added to the candidate motion information list mergeCandList in a predetermined order (e.g., order of relevance). One possible process for constructing mergeCandList is as follows: In a predetermined order and according to predetermined thresholds, check whether each motion information candidate, such as spatial motion information for adjacent locations, temporal motion information for adjacent locations, spatial motion information for non-adjacent locations, temporal motion information for non-adjacent locations, and history-based motion information, is similar to a motion information candidate that has already been added to mergeCandList or is confirmed to be added. If a motion information candidate awaiting addition to the list is not similar to a motion information candidate that has already been added to mergeCandList or is confirmed to be added, it is confirmed to add the motion information candidate awaiting addition to the list. Otherwise, it is confirmed not to add the motion information candidate awaiting addition to the list.
[0133] The order of correlation can be set as follows: The closer the distance to the current block, the stronger the correlation. When the distances are the same, spatial motion information predictions have a stronger correlation than temporal motion information predictions. However, since NATMVP can provide motion information in several directions that SMVP and NASMVP cannot, how to add these candidates to mergeCandList becomes an issue to consider.
[0134] In one exemplary embodiment of this disclosure, the set order is determined based on one or more of the following rules: If the distance of the spatial motion information is less than or equal to the distance of the temporal motion information, the spatial motion information is preferentially added to the list of candidate motion information. If the distance of spatial motion information is greater than the distance of temporal motion information, the temporal motion information is preferentially added to the list of candidate motion information. If multiple time-motion data points are located at different distances, the time-motion data point with the shortest distance is preferentially added to the list of candidate time-motion data points. If multiple time-motion information points are the same distance apart, the order in which they are added to the candidate motion information list is determined according to the statistical laws governing time-motion information.
[0135] The distance of spatial motion information is the distance from the position from which the spatial motion information is derived to the current block, and the distance of temporal motion information is the distance from the position from which the temporal motion information is derived to the current block. The distance from a position to the current block is determined based on the rectangular box to which that position is located or adjacent to it. The current block is enclosed by the rectangular box and the rectangular box S It is located in the center. The width of the rectangular box is an integer multiple of the width of the current block, and the height of the rectangular box is an integer multiple of the height of the current block. The larger the area of the rectangular box, the greater the distance from that position to the current block. For example, in Figure 12, the distance from positions 11, 14, 15, 13, 12, 17, 16, and 18 to the current block is considered to be the same. All of the above positions can include adjacent and non-adjacent positions. The horizontal coordinate deviation from the above position to the nearest sample in the current block is the absolute value of the difference between the horizontal coordinate of the above position and the horizontal coordinate of the nearest sample in the current block, and the vertical coordinate deviation from the above position to the nearest sample in the current block is the absolute value of the difference between the vertical coordinate of the above position and the vertical coordinate of the nearest sample in the current block.
[0136] The order of the correlations described above can be found in the order of the positions shown in Figure 12. The small black squares in the figure represent positions used for spatial motion information prediction, and the small squares with cutting lines represent positions used for temporal motion information prediction. The following can be seen: If the distance of the spatial motion information is less than or equal to the distance of the temporal motion information, for example, if position 16 is a position used for temporal motion information prediction and position 11 is a position used for spatial motion information prediction, and the distance from position 16 to the current block is the same as the distance from position 11 to the current block, then the spatial motion information derived based on position 11 is preferentially added to the candidate motion information list. As another example, if position 19 is a position used for spatial motion information prediction, and the distance from position 19 to the current block is greater than the distance from position 16 to the current block, then the temporal motion information derived based on position 16 is preferentially added to the candidate motion information list than the spatial motion information derived based on position 19. As another example, both position 10 and position 16 are positions used to derive time motion information, and since the distance from position 10 to the current block is smaller, the time motion information derived based on position 10 is preferentially added to the list of candidate motion information.
[0137] In one example, the first time-motion information includes at least two of the following: the first time-motion information derived based on a non-adjacent position to the lower right of the current block, the first time-motion information derived based on a non-adjacent position to the right of the current block, and the first time-motion information derived based on a non-adjacent position below the current block. When the distances between multiple first time-motion information are the same, the order in which they are added to the candidate motion information list is determined according to the statistical laws of time-motion information, which includes the following: The first time-motion information derived based on a non-adjacent position to the lower right of the current block is added to the candidate motion information list with priority over the first time-motion information derived based on a non-adjacent position to the right of the current block. The first time-motion information derived based on a non-adjacent position to the right of the current block is added to the candidate motion information list with priority over the first time-motion information derived based on a non-adjacent position below the current block.
[0138] Using Figure 12 as an example, positions 18, 16, and 17 are all used to derive time-motion information. According to the statistical laws governing time-motion information, the order in which these are added to the candidate motion information list must be determined. Research has shown that, in this order, the lower right position takes precedence over the right-side position, and the right-side position takes precedence over the lower-side position. Therefore, the time-motion information derived based on position 16 has the highest correlation, the time-motion information derived based on position 17 has the second highest correlation, and the time-motion information derived based on position 18 has the lowest correlation. The order of highly correlated time-motion information is earlier; that is, highly correlated time-motion information is preferentially added to the candidate motion information list.
[0139] The diagram shows two positions 6. If position 6 on the lower right side of the block is unavailable, it can be replaced with position 6 in the center of the block. If position 6 on the lower right side of the block is available, position 6 in the center of the block is not used. In practical applications, considering the trade-off between performance and complexity, it is possible not to use as many candidates as shown in the diagram. Alternatively, more candidates can be used according to the above principle.
[0140] In this embodiment, the position used to derive spatial motion information is indicated by a small black square in Figure 12, with a cutting line drawn in the center. (That is, multiple parallel lines were drawn.)The large square is the current block, and the five positions adjacent to the current block, i.e., positions 1 to 5, are used for spatial motion prediction to derive spatial motion information. Outside of these positions, several non-adjacent positions can be used to perform non-adjacent position-based spatial motion prediction (NASMVP). As shown in Figure 12, spatial motion information can be derived using the non-adjacent positions on the lower left, left, upper left, top, and upper right sides of the current block. The right, lower right, and bottom blocks have not yet been coded and therefore cannot be used for spatial motion prediction. There is a certain distance between these non-adjacent positions and the current block. For example, the horizontal distance between these positions and the current block is related to the length of the current block, and the vertical distance between these positions and the current block is related to the width of the current block. For example, non-adjacent positions are determined by expanding outward by the length and width of the current block. Positions in the inner layers are added to the list preferentially because they are closer in distance.
[0141] An example of spatial motion information prediction is shown in Figure 12. The items are added to the list in the order shown in the figure. This takes into account the characteristics of correlation and the statistical laws of motion.
[0142] In one exemplary embodiment of the present disclosure, the order of the earliest first time-motion information among the first time-motion information is determined by at least one of the following methods: In Geometric Segmentation Mode (GPM), the sequence number set for the earliest first time motion information is less than or equal to the maximum number of candidate motion information items that are allowed to be added to the candidate motion information list. In the configured high-speed motion scenario, the sequence number set for the earliest first time motion information is less than or equal to the maximum number of candidate motion information items that are allowed to be added to the candidate motion information list.
[0143] In this embodiment, in some specific cases, to better utilize time motion information at non-adjacent locations and effectively compensate for the lack of spatial motion information prediction, the sequence number set for the first time motion information in the GPM can be less than or equal to the maximum number of candidate motion information items that are allowed to be added to the candidate motion information list. For example, if the maximum number of candidate motion information items that are allowed to be added to the candidate motion information list is 6, the sequence number of the first time motion information can be set to 6 or a value less than 6. Here, the sequence number means the order in which the motion information is added to the candidate motion information list. A sequence number of 1 means that it is added to the candidate motion information list first, a sequence number of 2 means that it is added to the candidate motion information list second, and so on thereafter. By setting it in this way, it is possible to ensure that even if any of the higher-level motion information items are valid and added to the candidate motion information list, the first time motion information can also be added to the candidate motion information list. In these specific cases, in order to ensure that the first time motion information can be added to the candidate motion information list, the ranking of the first time motion information can be increased, and the maximum number of candidate motion information items that are permitted to be added to the candidate motion information list can also be increased, for example, increasing the maximum number from 6 to 7 in these specific cases. The above high-speed motion scenarios may be set by the user or acquired by the system itself through learning.
[0144] When adding each candidate motion information to mergeCandList, identity and similarity checks can be performed to prevent the same or very similar motion information from being added to mergeCandList. This helps mergeCandList actually provide more candidates. The methods and criteria for checking the above similarity may differ between time motion information prediction and spatial motion information prediction. For example, SMVP and NASMVP based on different locations may provide some similar motion information (such as some gradual changes or subtle changes) within the same object. However, time motion information, especially NATMVP, requires more "different" motion information. So-called "different" refers to motion information that is more significantly different than "similar". The difference between "different" and "similar" can be reflected in the difference in thresholds.
[0145] Each piece of spatial motion information, temporal motion information, and other motion information is added to mergeCandList in a specific order. When constructing the candidate list, the identity check or similarity check described above can be used. One method is to discard identical or similar motion information candidates and not add them to mergeCandList. When adding temporal motion information derived based on non-adjacent locations, different thresholds can be used to determine similarity. For example, the similarity threshold used for temporal motion information derived based on non-adjacent locations is greater than the similarity threshold used for spatial motion information derived based on adjacent locations (e.g., a single threshold set for motion vectors).
[0146] In one exemplary embodiment of the present disclosure, adding spatial motion information and temporal motion information to the candidate motion information list of the current block in a set order includes the following: A similarity check (including an identity check) is performed before adding the spatial motion information or temporal motion information to the candidate motion information list. If the similarity check determines that the spatial motion information or temporal motion information is not similar to any of the candidate motion information already added to or determined to be added to the candidate motion information list, the spatial motion information or temporal motion information is added to the candidate motion information list. A first similarity threshold θ1 is used when performing a similarity check on the first temporal motion information, and a second similarity threshold θ2 is used when performing a similarity check on the spatial motion information, where θ1 > θ2 or θ1 = θ2.
[0147] In one exemplary embodiment of the present disclosure, the time motion information further includes time motion information determined by time motion information prediction based on adjacent positions, and a third similarity threshold θ3 is used when performing a similarity check on the time motion information determined by time motion information prediction based on adjacent positions, such that θ1 > θ3 or θ1 = θ3. That is, in this embodiment, the similarity threshold set for time motion information determined by time motion information prediction based on adjacent positions may be smaller than the similarity threshold set for time motion information determined by NATMVP, so that the time motion information determined by time motion information prediction based on adjacent positions is more likely to pass the similarity check.
[0148] In one exemplary embodiment of this disclosure, a first similarity threshold is determined based on one of the parameters of the current image sequence and the parameters of the current image. In one example, the parameters of the current image sequence include the resolution of the sequence. The parameters of the current image include one or more of the width of the image, the height of the image, and the number of samples in the image. There are multiple first similarity thresholds, and the larger the first similarity threshold, the larger the value of the corresponding parameter. That is, the first similarity threshold in this embodiment is related to several parameters of the sequence or image. As can be understood, the same motion has a larger motion vector in a higher resolution video than in a lower resolution video. Therefore, the similarity threshold of time motion information derived based on non-adjacent positions can be set based on the parameters of the sequence or the parameters of the image, for example, the resolution of the sequence, the width and / or height of the image, or the number of samples in the image. For example, in a 1920×1080 sequence, the threshold is set to 64, and in a 1280×720 sequence, the threshold is set to 16. The threshold can be calculated based on the width and / or height of the image. For example, if the threshold is diffThT, the image width is picWidth, and the image height is picHeight, then diffThT = picWidth * picHeight >> 14. The ">>" indicates a right shift, i.e., division by 2.
[0149] In one exemplary embodiment of the present disclosure, a first similarity threshold is determined based on the reference relationships of the currently existing images, which include unidirectional and bidirectional references. The first similarity threshold determined when the reference relationship is a unidirectional reference is smaller than the first similarity threshold determined when the reference relationship is a bidirectional reference. Unidirectional reference This refers to the ability to use only forward-referenced or backward-referenced images in the current image, while bidirectional referencing refers to the ability to use both forward-referenced and backward-referenced images in the current image.
[0150] In other words, the first similarity threshold can be related to the reference relationships of the images. As mentioned above, in low-latency configurations, it is more difficult to obtain motion information in directions such as right to left or bottom to top compared to random-access configurations. This is because in low-latency configurations, the current image can only reference reference images that precede the current image according to the point of view (POC), or reference images that precede the current image according to the temporal order. In contrast, in random-access configurations, an image can reference both reference images that precede the current image according to the POC and reference images that follow the current image according to the POC. Therefore, the setting of the first similarity threshold can be related to the reference relationships of the current image. One possible method is as follows: If the current image can only reference reference images that precede the current image according to the POC, multiply the threshold of that image by a relatively small coefficient. The current image can only reference reference images that precede the current image according to the POC. And, according to the POC, it is possible to refer to both the current image and the reference image that follows it. In that case, the threshold of the image is relatively big Multiply by a coefficient. For example, a relatively small coefficient is 1, and a relatively large coefficient is 4. For example, the threshold is denoted as diffThT and the base threshold as diffThBase. diffThBase can be obtained by the other methods described above. If the current image can only reference reference images that precede it according to the POC, then diffThT = diffThBase * 1. If the current image can reference both reference images that precede it according to the POC and reference images that follow it according to the POC, then diffThT = diffThBase * 4.
[0151] In one exemplary embodiment of this disclosure, a first similarity threshold is determined based on whether or not template matching is used in the current prediction mode, and the first similarity threshold determined when template matching is used is greater than the first similarity threshold determined when template matching is not used. Because the template matching method enables searching within a certain range, motion information can be optimized to some extent, and the search range can be broadened by templates, so that candidate motion information that is too similar is not required in the candidate motion information list, and the more different the two motion information are, the more effective it becomes. For this reason, when the template matching method is used in the current block, the threshold used to determine similarity can be set higher than when the template matching method is not used.
[0152] In one exemplary embodiment of the present disclosure, the first similarity threshold is determined based on one or more of the following: a preset value, parameters of the current image sequence, parameters of the current image, size of the current block, sequence level flag of the current image sequence, slice level flag of the current image, image level flag of the current image, a flag indicating whether template matching is used in the current prediction mode, and reference relationships of the current image. In one example, the similarity threshold for time-motion information derived based on non-adjacent positions may be a fixed value such as 16, 32, 64, or 128. Since the motion information can support fractional pixel precision, this threshold can represent one pixel unit, two pixel units, four pixel units, eight pixel units, etc., where the pixel unit can be, for example, 1 / 16 of a pixel. In another example, the first similarity threshold may be related to the size of the current block, for example, the width and / or height of the current block, or the number of samples, which is used to determine the threshold. For example, if the number of samples in the current block is greater than 64, the threshold is set to 32; otherwise, the threshold is set to 16. In a further example, the first similarity threshold is determined based on a flag such as a sequence-level flag, an image-level flag, or a slice-level flag. For example, a flag of 0 indicates a threshold of 16, and a flag of 1 indicates a threshold of 32.
[0153] In one example of this embodiment, the first similarity threshold is determined based on any multiple parameters, which means that the first similarity threshold is determined based on multiple parametersThis includes setting the first similarity threshold to the maximum value among those determined based on the non-adjacent locations. Time-motion information derived based on non-adjacent locations can be determined as the maximum value of similarity thresholds set based on multiple factors. If the pre-set similarity threshold for time-motion information derived based on non-adjacent locations is 16 and the threshold determined based on other parameters is 1, then the similarity threshold for time-motion information derived based on non-adjacent locations is set to 16. If the similarity threshold set for time-motion information derived based on non-adjacent locations is 16 and the threshold determined based on other parameters is 32, then the similarity threshold for time-motion information derived based on non-adjacent locations is set to 32. However, in another example, it is also possible to determine different first similarity thresholds based on different combinations of parameters.
[0154] The number of time-motion information candidates that can be added can be limited. For example, you can set it to add a maximum of two time-motion information candidates, and once two time-motion information items that can be added to mergeCandList have been confirmed, the time-motion information prediction based on the subsequent position will end.
[0155] Based on all the methods described above, after determining the movement information to be added to mergeCandList, this movement information can be ordered according to some rule, and then mergeCandList can be finalized.
[0156] In one exemplary embodiment of the present disclosure, the time-motion information prediction method further includes: determining, based on a first flag, whether or not it is permitted to use time-motion information prediction based on non-adjacent locations for the current block; if it is determined that it is permitted to use time-motion information prediction based on non-adjacent locations, Time-based movement information forecasting Execute the method. The first flag includes one or more of the following: A first sequence level flag to indicate whether or not it is currently permitted to use time-motion information prediction based on non-adjacent positions in the image sequence, A first image-level flag to indicate whether or not it is permitted to use time-motion information prediction based on non-adjacent positions for the current image, and A first slice level flag to indicate whether or not it is currently permitted to use time-motion information predictions based on non-adjacent positions for the slice.
[0157] In this embodiment, the first flag includes at least two of the following: a first sequence level flag, a first image level flag, and a first slice level flag. The first sequence level flag has a higher level than the first image level flag, and the first image level flag has a higher level than the first slice level flag. Determining whether or not it is permitted to use non-adjacent position-based time-motion information prediction for the current block based on the first flag includes the following: Analyze the first flags in descending order of level. Analyze the lower level flags if a higher level flag indicates that it is permitted to use non-adjacent position-based time-motion information prediction. If all level flags indicate that it is permitted to use non-adjacent position-based time-motion information prediction, then it is determined that it is permitted to use non-adjacent position-based time-motion information prediction.
[0158] In one exemplary embodiment of the present disclosure, the time-motion information prediction method further includes: determining whether a first flag needs to be analyzed based on a second flag; analyzing the first flag if it is determined that it needs to be analyzed; and determining whether time-motion information prediction based on non-adjacent positions is permitted for the current block based on the first flag. The second flag is the next flag, namely, A flag indicating whether or not time-motion information is currently permitted to be used in the image sequence, and if the flag indicates that time-motion information is not currently permitted to be used in the image sequence, then it is determined that there is no need to analyze the first flag, A flag indicating whether or not it is permitted to use time-motion information in the current image, and if the flag indicates that it is not permitted to use time-motion information in the current image, then it is determined that there is no need to analyze the first flag, A flag indicating whether or not it is permitted to use non-adjacent locations for motion information derivation, where if the flag indicates that it is not permitted to use non-adjacent locations for motion information derivation, it is determined that there is no need to analyze the first flag, and It includes one or more of the following. If, based on the second flag, it is determined that the first flag needs to be analyzed, then the first flag is analyzed.
[0159] In the above embodiment, a flag is used to control whether NATMVP is enabled or disabled.
[0160] A flag can be used to control the enablement or disablement of non-adjacent position-based time-motion information prediction (NATMVP). This flag may be a sequence-level (sequence parameter set), image-level (image header or image parameter set), slice-level (slice header), or block-level flag. The flag can control whether (or whether it is permitted to) the NATMVP technology described in this invention is used for the corresponding sequence, image, slice, or block. For example, if the value of the flag is 1, the NATMVP technology described in this invention is used (or permitted to be used) for the current sequence, image, slice, or block. If the value of the flag is 0, the NATMVP technology described in this invention is not used (or is not permitted to be used) for the current sequence, image, slice, or block. This flag may depend on several other flags, such as sps_temporal_mvp_enabled_flag or ph_temporal_mvp_enabled_flag. The flag sps_temporal_mvp_enabled_flag controls whether or not time-motion information is permitted to be used in the current sequence. The flag ph_temporal_mvp_enabled_flag controls whether or not time-motion information is permitted to be used in the current image. To understand this, if time-motion information is permitted to be used in the current sequence or current image, the decoder needs to parse a flag to control whether or not the use of time-motion information prediction based on non-adjacent positions is permitted. Otherwise, the decoder does not need to parse a flag to control whether or not the use of time-motion information prediction based on non-adjacent positions is permitted. This flag may depend on other flags, such as a flag indicating whether or not the use of motion information derived from non-adjacent positions is permitted.
[0161] Flags for controlling NATMVP can be assigned levels, with lower-level flags depending on higher-level flags. For example, the sequence-level flag sps_natmvp_enabled_flag controls whether NATMVP is currently allowed for the sequence, and the image-level flag ph_natmvp_enabled_flag controls whether NATMVP is currently allowed for the image. If the value of sps_natmvp_enabled_flag is 1, the decoder needs to parse ph_natmvp_enabled_flag. If the value of sps_natmvp_enabled_flag is 0, the decoder does not need to parse ph_natmvp_enabled_flag.
[0162] In the method for constructing a candidate motion information list of the embodiments described herein, temporal motion information derived based on non-adjacent positions can be added to the candidate motion information list, and by considering the order of correlation between temporal motion information derived based on non-adjacent positions and other motion information, motion information for several scenarios that cannot be covered by spatial motion information can be effectively supplemented, thereby improving compression efficiency.
[0163] One embodiment of the present disclosure further provides a video encoding method. As shown in Figure 14, the video encoding method includes the following:
[0164] Step 310: Construct a candidate motion information list for the current block based on the method for constructing a candidate motion information list described in any one embodiment of the present disclosure.
[0165] Step 320: Select one or more candidate motion information from the list of candidate motion information and record the index of the selected candidate motion information.
[0166] Step 330: Determine the predicted block for the current block based on the selected candidate movement information, encode the current block based on the predicted block, and encode the index of the candidate movement information.
[0167] Based on the selected candidate motion information, the position of the current block's reference block in the reference image can be determined. In this case, the reference block can be obtained by adding the motion vector difference to the selected candidate motion information. Based on the reference block (e.g., one or two), a predicted block can be obtained. Then, the residual between the predicted block and the current block can be calculated and encoded.
[0168] One embodiment of this disclosure further provides a video decoding method. 15 As shown, the video decoding method includes the following:
[0169] Step 410: Construct a candidate motion information list for the current block based on the method for constructing a candidate motion information list described in any one embodiment of the present disclosure.
[0170] Step 420: Based on the index of the candidate motion information for the current block obtained by decoding, select one or more candidate motion information from the list of candidate motion information.
[0171] Step 430: Determine the predicted block for the current block based on the selected candidate movement information, and reconstruct the current block based on the predicted block.
[0172] One embodiment of the present disclosure further provides a time motion information prediction device. As shown in Figure 16, the time motion information prediction device comprises a processor 71 and a memory 73 in which a computer program is stored. When the processor 71 executes the computer program, it executes the time motion information prediction method described in any one embodiment of the present disclosure.
[0173] One embodiment of the present disclosure further provides an apparatus for constructing a candidate motion information list. As shown in Figure 16, the apparatus for constructing a candidate motion information list comprises a processor and memory in which a computer program is stored. When the processor executes the computer program, it performs a method for constructing a candidate motion information list as described in any one embodiment of the present disclosure.
[0174] One embodiment of the present disclosure further provides a video encoding device. As shown in Figure 16, the video encoding device comprises a processor and memory in which a computer program is stored. When the processor executes the computer program, it executes the video encoding method described in any one embodiment of the present disclosure.
[0175] One embodiment of the present disclosure further provides a video decoding apparatus. As shown in Figure 16, the video decoding apparatus comprises a processor and memory in which a computer program is stored. When the processor executes the computer program, it executes the video decoding method described in any one embodiment of the present disclosure.
[0176] One embodiment of the present disclosure further provides a video coding system, which includes a video encoding device described in any one embodiment of the present disclosure and a video decoding device described in any one embodiment of the present disclosure.
[0177] One embodiment of the present disclosure further provides a non-temporary computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, a method described in any one embodiment of the present disclosure is performed.
[0178] One embodiment of the present disclosure further provides a bitstream, which is generated by a video encoding method described in any one embodiment of the present disclosure.
[0179] The video encoding and / or video decoding devices of the embodiments described herein can be implemented using any one or any combination of the following circuits: one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic circuits, and hardware. If the disclosure is partially implemented in software, the instructions used in the software may be stored in a suitable non-volatile computer-readable storage medium, and one or more processors may be used to implement the methods of the embodiments of the disclosure by executing the instructions in hardware.
[0180] In one or more exemplary embodiments, the functions described may be implemented by hardware, software, firmware, or any combination thereof. If implemented by software, the functions may be stored as one or more instructions or codes in a computer-readable medium, or transmitted through a computer-readable medium, and executed by a hardware-based processing unit. The computer-readable medium may include tangible media such as data storage media, or any communication media that facilitates the transmission of a computer program from one place to another, for example, according to a communication protocol. Thus, the computer-readable medium may typically be non-transient tangible computer-readable storage media, or communication media such as signals or carriers. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the technology described herein. Computer program products may include computer-readable media.
[0181] As a non-limiting example, such computer-readable storage media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), compact disk ROM (CD-ROM) or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium that stores desired program code in the form of instructions or data structures and is accessible by a computer. Furthermore, any connection may be called a computer-readable storage medium, for example, when transmitting instructions from a website, server or other remote source using coaxial cable, fiber optic cable, twisted-pair cabling, digital subscriber line (DSL), or wireless technologies such as infrared, radio, or microwave, coaxial cable, fiber optic cable, twisted-pair cabling, DSL, or wireless technologies such as infrared, radio, or microwave are included in the definition of a medium. However, computer-readable storage media and data storage media are non-transient tangible storage media that do not include connections, carriers, signals, or other temporary (transient) media. The magnetic disks and optical disks used herein include compact discs (CDs), laserdiscs, optical disks, digital versatile discs (DVDs), floppy disks, or Blu-ray discs, where magnetic disks typically reproduce data magnetically, and optical disks reproduce data optically using a laser. Any combination of the above should also be included within the scope of computer-readable media.
[0182] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated circuits or discrete logic circuits. Thus, the term “processor” as used herein may refer to any one of the above structures or any other structure suitable for carrying out the techniques described herein. Furthermore, in some embodiments, the functions described herein may be provided within dedicated hardware and / or software modules configured for use in encoding and decoding, and may also be incorporated into an integrated encoder-decoder. Additionally, the techniques described herein can be fully implemented in one or more circuits or logic elements.
[0183] The technical proposals of the embodiments of this disclosure can be implemented in various devices or equipment, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., a chipset). Embodiments of this disclosure highlight the functionality of a device configured to perform the described technology using various components, modules, or units, which are not necessarily implemented by different hardware units. As described above, the various units may be combined in a codec hardware unit, or they may be provided in combination with an array of interoperable hardware units (including one or more processors as described above) and appropriate software and / or firmware.
Claims
1. A method for predicting time-movement information, Based on the first flag, determine whether it is permitted to use time-movement information predictions based on at least one non-adjacent position of the current block for the current block, If it is determined that it is permitted to use time-movement information predictions based on the aforementioned non-adjacent locations, Determine at least one non-adjacent location of the current block, Based on the motion information of at least one non-adjacent position in the coded image, the first time motion information of the current block is determined. Includes, The coded image is the reference image of the current block, Determining the first temporal motion information of the current block based on the motion information of at least one non-adjacent position in the coded image is: To determine the collated block corresponding to the non-adjacent position in the aforementioned reference image, To obtain the movement information of the aforementioned collated block, This includes scaling the motion vectors in the acquired motion information to obtain first time motion information for the current block, The first flag is, A first sequence level flag to indicate whether or not it is currently permitted to use time motion information prediction based on the non-adjacent positions in the image sequence, A first image level flag to indicate whether or not it is permitted to use the time motion information prediction based on the non-adjacent positions in the current image, and A first slice level flag to indicate whether or not it is currently permitted to use time-motion information predictions based on the non-adjacent positions for slicing, Includes one or more of the following: A method for predicting time-movement information characterized by the following features.
2. The aforementioned non-adjacent position is a non-adjacent position in the following direction, i.e., The non-adjacent position to the right of the aforementioned current block, The non-adjacent position below the current block, and The non-adjacent position on the lower right side of the aforementioned current block, Includes one or more of the following: If there are at least three positions in the same direction for deriving the first time motion information, the distance between adjacent positions among the at least three positions is the same, or the distance between adjacent positions among the at least three positions is variable, and the distance between two adjacent positions increases as they move away from the current block. The method for predicting time motion information according to feature 1.
3. The horizontal and / or vertical distances between the non-adjacent position and the setting sample within the current block are predetermined fixed values, or, The horizontal and / or vertical distance between the non-adjacent position and the setting sample within the current block is variable and depends on the following parameters: Current block size, Current image sequence parameters, The current sequence level flag for the image sequence, The current image level flag, and Current image slice level flag, Determined based on any one of the following, or any combination thereof. The method for predicting time motion information according to feature 1.
4. All of the non-adjacent positions are located within the range of the largest coding unit (LCU) or coding tree unit (CTU) in which the current block is located, or all of the non-adjacent positions are located within the range of the LCU row or CTU row in which the current block is located. The method for predicting time motion information according to feature 1.
5. The aforementioned time-motion information prediction method further includes time-motion information prediction based on adjacent positions, The adjacent positions include the adjacent position of the lower right corner of the current block, the non-adjacent positions of the current block and the adjacent positions of the lower right corner of the current block are distributed in an array, the upper left corner of the array is the adjacent position of the lower right corner of the current block, and different positions within the array are distributed in different minimum memory units in the coded image. The method for predicting time motion information according to feature 1.
6. The first flag includes at least two of the first sequence level flag, the first image level flag, and the first slice level flag, wherein the first sequence level flag has a higher level than the first image level flag, and the first image level flag has a higher level than the first slice level flag. Based on the first flag, determining whether it is permitted to use time-movement information predictions based on non-adjacent positions for the current block is: The first flag is analyzed sequentially in descending order of level, If a higher-level flag indicates that it is permitted to use time-motion information predictions based on the non-adjacent positions, then the lower-level flag is analyzed, If all levels of the flags indicate that the use of time-motion information predictions based on non-adjacent positions is permitted, then it is confirmed that the use of time-motion information predictions based on non-adjacent positions is permitted. including, The method for predicting time motion information according to feature 1.
7. The aforementioned time movement information prediction method is: Based on the second flag, determine whether or not it is necessary to analyze the first flag, If it is determined that the aforementioned first flag needs to be analyzed, then the aforementioned first flag will be analyzed, The method further includes determining, based on the first flag, whether or not it is permitted to use time-motion information predictions based on the non-adjacent positions for the current block, The aforementioned second flag is the next flag, namely, A flag indicating whether or not it is permitted to use time-motion information in the current image sequence, wherein if it indicates that it is not permitted to use time-motion information in the current image sequence, it is determined that there is no need to analyze the first flag, A flag indicating whether or not it is permitted to use time-motion information in the current image, and if it indicates that it is not permitted to use time-motion information in the current image, then it is determined that there is no need to analyze the first flag, and A flag indicating whether or not it is permitted to use non-adjacent locations for motion information derivation, wherein if it indicates that it is not permitted to use non-adjacent locations for motion information derivation, then it is determined that there is no need to analyze the first flag, Includes one or more of the following: If it is determined that the first flag needs to be analyzed based on the second flag, the first flag is analyzed. The method for predicting time motion information according to feature 1.
8. The at least one non-adjacent location is The coordinates are (xCb+2*k_cbWidth-1, yCb+2*k_cbHeight-1), (xCb+2*k_cbWidth, yCb+2*k_cbHeight-1), (xCb+2*k_cbWidth-1, yCb+2*k_cbHeight), or (xCb+2*k_cbWidth, yCb+2*k_cbHeight), The second position where the coordinates are (xCb + 3 * k_cbWidth - 1, yCb + 3 * k_cbHeight - 1), (xCb + 3 * k_cbWidth, yCb + 3 * k_cbHeight - 1), (xCb + 3 * k_cbWidth - 1, yCb + 3 * k_cbHeight), or (xCb + 3 * k_cbWidth, yCb + 3 * k_cbHeight), The third position where the coordinates are (xCb + 3 * k_cbWidth - 1, yCb + k_cbHeight / 2 - 1), (xCb + 3 * k_cbWidth, yCb + k_cbHeight / 2 - 1), (xCb + 3 * k_cbWidth - 1, yCb + k_cbHeight / 2), or (xCb + 3 * k_cbWidth, yCb + k_cbHeight / 2), The fourth position where the coordinates are (xCb + k_cbWidth / 2, yCb + 3 * k_cbHeight - 1), (xCb + k_cbWidth / 2, yCb + 3 * k_cbHeight), (xCb + k_cbWidth / 2 - 1, yCb + 3 * k_cbHeight - 1), or (xCb + k_cbWidth / 2 - 1, yCb + 3 * k_cbHeight), The fifth position where the coordinates are (xCb + 2 * k_cbWidth - 1, yCb + k_cbHeight / 2 - 1), (xCb + 2 * k_cbWidth, yCb + k_cbHeight / 2 - 1), (xCb + 2 * k_cbWidth - 1, yCb + k_cbHeight / 2), or (xCb + 2 * k_cbWidth, yCb + k_cbHeight / 2), The sixth position where the coordinates are (xCb + k_cbWidth / 2, yCb + 2 * k_cbHeight - 1), (xCb + k_cbWidth / 2, yCb + 2 * k_cbHeight), (xCb + k_cbWidth / 2 - 1, yCb + 2 * k_cbHeight - 1), or (xCb + k_cbWidth / 2 - 1, yCb + 2 * k_cbHeight), The seventh position where the coordinates are (xCb + k_cbWidth, yCb + 2 * k_cbHeight), (xCb + k_cbWidth, yCb + 2 * k_cbHeight - 1), (xCb + k_cbWidth - 1, yCb + 2 * k_cbHeight), or (xCb + k_cbWidth - 1, yCb + 2 * k_cbHeight - 1), and, The coordinates are (xCb+2*k_cbWidth, yCb+k_cbHeight), (xCb+2*k_cbWidth-1, yCb+k_cbHeight), (xCb+2*k_cbWidth, yCb+k_cbHeight-1), or (xCb+2*k_cbWidth-1, 8th position which is yCb+k_cbHeight-1), Includes one or more of the following: xCb is the horizontal coordinate of the upper left corner of the current block, yCb is the vertical coordinate of the upper left corner of the current block, k_cbWidth is the width of the current block, or half the width, or a quarter of the width, or twice the width, k_cbHeight is the height of the current block, or half the height, or a quarter of the height, or twice the height, and "*" indicates a multiplication operation. The method for predicting time motion information according to feature 1.
9. A method for constructing a list of candidate movement information, Currently, spatial motion information prediction and temporal motion information prediction are performed on the block to determine the spatial motion information and temporal motion information of the block. This includes adding the spatial motion information and the temporal motion information to the candidate motion information list of the current block in the set order, The time motion information prediction method described in claim 1 is used for the time motion information prediction, and the time motion information includes the first time motion information. A method for constructing a list of candidate movement information, characterized by the following features.
10. The previously set order follows the following rules, namely: If the distance of the spatial motion information is less than or equal to the distance of the temporal motion information, the spatial motion information is preferentially added to the candidate motion information list. If the distance of the spatial motion information is greater than the distance of the temporal motion information, the temporal motion information is preferentially added to the candidate motion information list. If multiple time-motion information points have different distances, the time-motion information point with the smaller distance is preferentially added to the candidate motion information list, and If the distance between multiple time-motion information items is the same, the order in which they are added to the candidate motion information list is determined according to the statistical laws governing the time-motion information. Determined based on one or more of the following: The distance of the spatial motion information is the distance from the position from which the spatial motion information is derived to the current block, the distance of the temporal motion information is the distance from the position from which the temporal motion information is derived to the current block, the distance from one position to the current block is determined based on the rectangular box on which the position is located or adjacent to the current block, the current block is surrounded by the rectangular box and located at the center of the rectangular box, the width of the rectangular box is an integer multiple of the width of the current block, the height of the rectangular box is an integer multiple of the height of the current block, and the larger the area of the rectangular box, the greater the distance from the position to the current block. The method for constructing a list of candidate motion information according to feature 9.
11. The first time motion information described above is The first time motion information derived based on the non-adjacent position on the lower right side of the current block, The first time motion information derived based on the non-adjacent position to the right of the current block, and First time motion information derived based on the non-adjacent position below the current block, It includes at least two of the following: If multiple first time-motion information points are the same distance apart, the order in which they are added to the candidate motion information list is determined according to the statistical laws of time-motion information. The first time motion information derived based on the non-adjacent position on the lower right side of the current block is added to the candidate motion information list with priority over the first time motion information derived based on the non-adjacent position on the right side of the current block. The first time motion information derived based on the non-adjacent position to the right of the current block is added to the candidate motion information list with priority over the first time motion information derived based on the non-adjacent position below the current block, The method for constructing a list of candidate motion information according to claim 10.
12. The order of the earliest of the first time-motion information is as follows: In geometric segmentation mode (GPM), the sequence number set for the earliest first time motion information is less than or equal to the maximum number of candidate motion information items that are permitted to be added to the candidate motion information list. In the configured high-speed motion scenario, the sequence number set for the earliest first time motion information is less than or equal to the maximum number of candidate motion information items that are permitted to be added to the candidate motion information list. It is determined by at least one of the following: The method for constructing a list of candidate motion information according to feature 9.
13. Adding the spatial motion information and the temporal motion information to the candidate motion information list of the current block in the order set above is: Before adding the spatial motion information or the temporal motion information to the candidate motion information list, a similarity check is performed. If the similarity check determines that the spatial motion information or the temporal motion information is not similar to any of the candidate motion information already added to or confirmed to be added to the candidate motion information list, the spatial motion information or the temporal motion information is added to the candidate motion information list. First similarity threshold θ 1 This is used when performing a similarity check on the first time motion information, and the second similarity threshold θ 2 This is used when performing a similarity check on the aforementioned spatial motion information, θ 1 >θ 2 or θ 1 =θ 2 That is, The method for constructing a list of candidate motion information according to feature 9.
14. The time motion information further includes the time motion information determined by prediction of time motion information based on adjacent positions, and a third similarity threshold θ 3 is used when performing a similarity check on the time motion information determined by prediction of time motion information based on the adjacent positions, and θ 1 > θ 3 or θ 1 = θ 3 is satisfied. The method for constructing a list of candidate motion information according to claim 13.
15. The first similarity threshold is determined based on one of the parameters of the current image sequence and the parameters of the current image. The parameters of the current image sequence include the resolution of the sequence. The parameters of the current image include one or more of the following: image width, image height, and number of image samples. There are multiple first similarity thresholds, and the larger the first similarity threshold, the larger the value of the corresponding parameter. The method for constructing a list of candidate motion information according to claim 13.
16. The first similarity threshold is determined based on the reference relationships of the current images, and these reference relationships include unidirectional and bidirectional references. The first similarity threshold determined when the reference relationship is a unidirectional reference is smaller than the first similarity threshold determined when the reference relationship is a bidirectional reference. Unidirectional referencing means that only the forward-referenced image or the backward-referenced image can be used in the current image, while bidirectional referencing means that both the forward-referenced image and the backward-referenced image can be used in the current image. The method for constructing a list of candidate motion information according to claim 13.
17. The first similarity threshold is determined based on whether or not template matching is used in the current prediction mode, and the first similarity threshold determined when template matching is used is greater than the first similarity threshold determined when template matching is not used. The method for constructing a list of candidate motion information according to claim 13.
18. The first similarity threshold is determined based on one or more of the following parameters: a preset value, the parameters of the current image sequence, the parameters of the current image, the size of the current block, the sequence level flag of the current image sequence, the slice level flag of the current image, the image level flag of the current image, a flag indicating whether or not to use template matching in the current prediction mode, and the reference relationships of the current image. The method for constructing a list of candidate motion information according to claim 13.
19. A video decoding method, Construct a candidate movement information list for the current block based on the method for constructing a candidate movement information list described in any one of claims 9 to 18, Based on the index of the candidate motion information of the current block obtained by decoding, one or more candidate motion information are selected from the candidate motion information list. Based on the selected candidate movement information, the predicted block of the current block is determined, and the current block is reconstructed based on the predicted block. including, A video decoding method characterized by the following features.