Video encoding method and apparatus

By predicting video blocks using initial motion vectors and offsets from correlated reference blocks, the method addresses delays in motion vector updates, enhancing encoding efficiency and reducing processing delays.

JP2025100574AInactive Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025061440
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Conventional video encoding methods experience delays in parallel processing due to the need for motion vector updates, which are completed only after the actual motion vector is determined, affecting encoding efficiency.

Method used

The method involves determining a reference block with a preset temporal or spatial correlation, using its initial motion vector and offsets to predict subsequent blocks, allowing prediction to proceed before the actual motion vector update is completed.

Benefits of technology

This approach improves encoding efficiency by enabling prediction steps to be executed earlier, reducing processing delays and enhancing overall coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100574000001_ABST
    Figure 2025100574000001_ABST
Patent Text Reader

Abstract

To provide a video coding technology.SOLUTION: A method for obtaining a motion vector according to an embodiment of the present application includes steps of: determining a reference block for a block to be processed, a reference block and the block to be processed having a preset temporal or spatial correlation with each other, the reference block having an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block obtained on the basis of the predicted motion vector of the reference block, and the predicted block of the reference block obtained on the basis of the initial motion vector and one or more preset motion vector offsets; and using the initial motion vector of the reference block as the predicted motion vector of the block to be processed.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding technology, and more particularly, to a video coding method and apparatus.

Background Art

[0002] Digital video technology can be widely applied to various devices including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), notebook computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, mobile phones or satellite radiotelephones, video conferencing devices, video streaming transmission devices, etc. Digital video devices implement video decoding technologies, such as MPEG-2, MPEG-4, ITU-T Recommendation H.263, ITU-T Recommendation H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T Recommendation H.265 (also called High Efficiency Video Coding (HEVC)), and video decoding technologies described in the extended parts of these standards. By implementing these video decoding technologies, digital video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently.

[0003] In video compression technology, redundant information inherent in a video sequence can be reduced or removed by performing spatial (intra) prediction and / or temporal (inter) prediction. In the case of block-based video decoding, a video picture can be divided into video blocks. Video blocks may also be referred to as tree blocks, coding units / decoding units (CUs), or coding nodes / decoding nodes. Video blocks in an intra-decoded (I) slice of a picture are encoded by spatial prediction of reference samples in adjacent blocks of the same picture. Video blocks in an inter-decoded (P or B) slice of a picture can be encoded by spatial prediction of reference samples in adjacent blocks of the same picture or by temporal prediction of reference samples in another reference picture. A picture may sometimes be called a frame, and a reference picture may sometimes be called a reference frame.

Summary of the Invention

[0004] Embodiments of the present application provide a video encoding method and related devices, mainly related to the acquisition of motion vectors. In conventional inter-prediction techniques and intra-prediction techniques related to motion estimation, motion vectors are important implementation elements and are used to determine predictors of blocks to be processed in order to reconstruct the blocks to be processed. Generally, a motion vector is composed of a predicted motion vector and a motion vector difference. The motion vector difference is the difference between a motion vector and a predicted motion vector. In some techniques, for example, in the motion vector merge mode (Merge mode), the motion vector difference is not used and the predicted motion vector is directly regarded as the motion vector. The predicted motion vector is usually obtained from a previous encoded or decoded block that has a temporal or spatial correlation with the block to be processed, and the motion vector of the block to be processed is usually used as the predicted motion vector of a subsequent encoded or decoded block.

[0005] However, with the development of technology, techniques related to the update of motion vectors have emerged. The motion vector for determining the predictor of the block to be processed is no longer directly obtained from the predicted motion vector, or the sum of the predicted motion vector and the difference between the motion vectors (here, the predicted motion vector, or the sum of the predicted motion vector and the difference between the motion vectors is called the initial motion vector), but is obtained from the updated value of the initial motion vector. Specifically, after the initial motion vector of the block to be processed is obtained, the initial motion vector is first updated to obtain the actual motion vector, and then the predicted block of the block to be processed is obtained by using the actual motion vector. The actual motion vector is stored for use in the prediction procedure of subsequent encoding or decoding blocks. The motion vector update technique improves prediction accuracy and encoding efficiency. However, for subsequent encoding or decoding blocks, the prediction step can only be executed after the update of the motion vector for one or more previous encoding or decoding blocks is completed, that is, only after the actual motion vector is determined. This causes a delay in the parallel processing or pipeline processing of different blocks compared to the method where the motion vector update is not performed.

Means for Solving the Problems

[0006] According to a first aspect of the present application, there is provided a method for obtaining a motion vector, including determining a reference block of a block to be processed, where the reference block and the block to be processed have a preset temporal or spatial correlation relationship, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on the predicted motion vector of the reference block, and the predicted block of the reference block is obtained based on the initial motion vector and one or more preset motion vector offsets; and using the initial motion vector of the reference block as the predicted motion vector of the block to be processed.

[0007] As described above, the initial motion vector before update is used to replace the actual motion vector and to predict subsequent encoded or decoded blocks. Before the update of the actual motion vector is completed, a prediction step can be executed for subsequent encoded or decoded blocks. This ensures an improvement in encoding efficiency brought about by the update of the motion vector and eliminates processing delays.

[0008] In a first realizable implementation of the first aspect, the initial motion vector of the reference block is specifically obtained in the following manner, that is, using the predicted motion vector of the reference block as the initial motion vector of the reference block, or adding the difference between the predicted motion vector of the reference block and the motion vector of the reference block to obtain the initial motion vector of the reference block.

[0009] In different inter-prediction modes, the initial motion vector can be obtained from the predicted motion vector or the sum of the predicted motion vector and the motion vector difference. This improves the encoding efficiency.

[0010] In a second possible implementation of the first aspect, the prediction block of the reference block is obtained in the following manner, i.e., from the reference frame of the reference block, a picture block indicated by the initial motion vector of the reference block is obtained, and the obtained picture block is used as a temporary prediction block of the reference block; to obtain one or more actual motion vectors, adding the initial motion vector of the reference block and one or more preset motion vector offsets, where each actual motion vector indicates a search position; obtaining one or more candidate prediction blocks at the search positions indicated by the one or more actual motion vectors, where each search position corresponds to one candidate prediction block; and specifically obtaining by selecting, as the prediction block of the reference block, the candidate prediction block with the minimum pixel difference from the temporary prediction block among the one or more candidate prediction blocks.

[0011] In this implementation, the update method of the motion vector is specifically described. Based on the update of the motion vector, the prediction becomes more accurate and the coding efficiency is improved.

[0012] In a third possible implementation of the first aspect, the method is used for bidirectional prediction, the reference frame includes a reference frame in the first direction and a reference frame in the second direction, the initial motion vector includes an initial motion vector in the first direction and an initial motion vector in the second direction, and the step of obtaining, from the reference frame of the reference block, a picture block indicated by the initial motion vector of the reference block and using the obtained picture block as a temporary prediction block of the reference block includes obtaining a first picture block indicated by the initial motion vector in the first direction of the reference block from the reference frame in the first direction of the reference block; obtaining a second picture block indicated by the initial motion vector in the second direction of the reference block from the reference frame in the second direction of the reference block; and weighting the first picture block and the second picture block to obtain a temporary prediction block of the reference block.

[0013] In this implementation form, the update method of the motion vector during bidirectional prediction is particularly described. Based on the update of the motion vector, the prediction becomes more accurate and the coding efficiency is improved.

[0014] In a fourth possible implementation form of the first aspect, the method further includes a step of rounding the motion vector resolution of the actual motion vector so that the motion vector resolution of the processed actual motion vector becomes equal to a preset pixel accuracy when the motion vector resolution of the actual motion vector is higher than the preset pixel accuracy.

[0015] This implementation form ensures that the motion vector resolution of the actual motion vector becomes equal to the preset pixel accuracy, and reduces the computational complexity caused by different motion vector resolutions. It should be understood that even when the initial motion vector before update is used to replace the actual motion vector and the method used to predict subsequent coding blocks or decoding blocks is not used, this implementation form can reduce the delay because the complexity of the motion vector update is reduced when this implementation form is used separately.

[0016] In a fifth possible implementation form of the first aspect, the step of selecting, as the prediction block of the reference block, the candidate prediction block with the smallest pixel difference from the temporary prediction block from one or more candidate prediction blocks includes the step of selecting the actual motion vector corresponding to the candidate prediction block with the smallest pixel difference from the temporary prediction block from one or more candidate prediction blocks, the step of rounding the motion vector resolution of the selected actual motion vector so that the motion vector resolution of the processed selected actual motion vector becomes equal to the preset pixel accuracy when the motion vector resolution of the selected actual motion vector is higher than the preset pixel accuracy, and the step of determining that the prediction block corresponding to the position indicated by the processed selected actual motion vector is the prediction block of the reference block.

[0017] This implementation also ensures that the motion vector resolution of the actual motion vector equals the preset pixel accuracy, and reduces the computational complexity caused by different motion vector resolutions. It should be understood that even when the initial motion vector before update is used to replace the actual motion vector and the method of predicting subsequent encoding blocks or decoding blocks is not used, this implementation can reduce the delay because the complexity of updating the motion vector is reduced when this implementation is used separately.

[0018] In a sixth possible implementation of the first aspect, the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0019] In a seventh possible implementation of the first aspect, the method further includes the step of using the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed.

[0020] In an eighth possible implementation of the first aspect, the method further includes the step of adding the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed to obtain the initial motion vector of the block to be processed.

[0021] In a ninth possible implementation of the first aspect, the method is used for video decoding, and the difference between the motion vectors of the block to be processed is obtained by analyzing the first identification information in the bitstream.

[0022] In different inter prediction modes, the initial motion vector can be obtained from the predicted motion vector or the sum of the predicted motion vector and the difference between the motion vectors. This improves the coding efficiency.

[0023] In a tenth possible implementation of the first aspect, the method is used for video decoding, and the step of determining a reference block for a block to be processed includes the step of analyzing a bitstream to obtain second identification information, and the step of determining a reference block for the block to be processed based on the second identification information.

[0024] In an eleventh possible implementation of the first aspect, the method is used for video encoding, and the step of determining a reference block for a block to be processed includes the step of selecting, from one or more candidate reference blocks of the block to be processed, a candidate reference block with the minimum rate-distortion cost as the reference block for the block to be processed.

[0025] A reference block is a video picture block having a spatial or temporal correlation relationship with the block to be processed, and can be, for example, a spatially adjacent block or a co-located block at the same temporal position. The motion vector of the reference block is used to predict the motion vector of the block to be processed. This improves the coding efficiency of the motion vector.

[0026] According to a second aspect of the present application, there is provided an apparatus for obtaining a motion vector, including a determination module configured such that a reference block for a block to be processed is determined, the reference block and the block to be processed have a preset temporal or spatial correlation relationship, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on the predicted motion vector of the reference block, and the predicted block of the reference block is obtained based on the initial motion vector and one or more preset motion vector offsets, and an acquisition module configured to use the initial motion vector of the reference block as the predicted motion vector of the block to be processed.

[0027] In a first possible implementation of the second aspect, the acquisition module is further configured to use the predicted motion vector of the reference block as the initial motion vector of the reference block, or to add the predicted motion vector of the reference block and the difference between the motion vector of the reference block to obtain the initial motion vector of the reference block.

[0028] In a second possible implementation of the second aspect, the acquisition module obtains a picture block indicated by the initial motion vector of the reference block from the reference frame of the reference block, uses the obtained picture block as a temporary prediction block of the reference block, and adds the initial motion vector of the reference block and one or more preset motion vector offsets to obtain one or more actual motion vectors. Each actual motion vector indicates a search position, and one or more candidate prediction blocks are obtained at the search positions indicated by the one or more actual motion vectors. Each search position corresponds to one candidate prediction block, and a candidate prediction block with the smallest pixel difference from the temporary prediction block is selected as the prediction block of the reference block.

[0029] In a third possible implementation of the second aspect, the apparatus is configured for bidirectional prediction. The reference frame includes a reference frame in a first direction and a reference frame in a second direction. The initial motion vector includes an initial motion vector in the first direction and an initial motion vector in the second direction. The acquisition module obtains a first picture block indicated by the initial motion vector in the first direction of the reference block from the reference frame in the first direction of the reference block, and obtains a second picture block indicated by the initial motion vector in the second direction of the reference block from the reference frame in the second direction of the reference block. The apparatus is particularly configured to weight the first picture block and the second picture block to obtain a temporary prediction block of the reference block.

[0030] In a fourth possible implementation of the second aspect, when the motion vector resolution of the actual motion vector is higher than the preset pixel accuracy, the apparatus further includes a rounding module configured to round the motion vector resolution of the actual motion vector so that the motion vector resolution of the processed actual motion vector is equal to the preset pixel accuracy.

[0031] In a fifth possible implementation of the second aspect, the acquisition module selects an actual motion vector corresponding to a candidate prediction block with the smallest pixel difference from one or more candidate prediction blocks to a temporary prediction block. When the motion vector resolution of the selected actual motion vector is higher than the preset pixel accuracy, the motion vector resolution of the selected actual motion vector is rounded so that the motion vector resolution of the processed selected actual motion vector is equal to the preset pixel accuracy, and a prediction block corresponding to the position indicated by the processed selected actual motion vector is determined to be the prediction block of the reference block.

[0032] In a sixth possible implementation of the second aspect, the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0033] In a seventh possible implementation of the second aspect, the acquisition module is specifically configured to use the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed.

[0034] In an eighth possible implementation of the second aspect, the acquisition module is specifically configured to add the difference between the predicted motion vector of the block to be processed and the motion vector of the block to be processed in order to obtain the initial motion vector of the block to be processed.

[0035] In the ninth possible implementation of the second aspect, the apparatus is used for video decoding, and the difference of the motion vectors of the blocks to be processed is obtained by analyzing the first identification information in the bitstream.

[0036] In the tenth possible implementation of the second aspect, the apparatus is used for video decoding, and the determination module is specifically configured to analyze the bitstream to obtain the second identification information and determine the reference block of the block to be processed based on the second identification information.

[0037] In the eleventh possible implementation of the second aspect, the apparatus is used for video encoding, and the determination module is specifically configured to select, as the reference block of the block to be processed, the candidate reference block with the minimum rate-distortion cost from one or more candidate reference blocks of the block to be processed.

[0038] According to the third aspect of the present application, a method for obtaining a motion vector is provided, including: determining a reference block of a block to be processed, where the reference block and the block to be processed have a preset temporal or spatial correlation relationship; obtaining an initial motion vector of the block to be processed based on the reference block; obtaining a predicted block of the block to be processed based on the initial motion vector of the block to be processed and one or more preset motion vector offsets; and using the initial motion vector of the block to be processed as the predicted motion vector of a subsequent block to be processed that is processed after the block to be processed.

[0039] In a first realizable implementation of the third aspect, the step of obtaining an initial motion vector of a block to be processed based on a reference block includes using the initial motion vector of the reference block as the initial motion vector of the block to be processed, or adding the difference between the initial motion vector of the reference block and the motion vector of the block to be processed to obtain the initial motion vector of the block to be processed.

[0040] In a second realizable implementation of the third aspect, the step of obtaining a predicted block of a block to be processed based on the initial motion vector of the block to be processed and one or more preset motion vector offsets includes obtaining a picture block indicated by the initial motion vector of the block to be processed from the reference frame of the block to be processed, and using the obtained picture block as a temporary predicted block of the block to be processed; adding the initial motion vector of the block to be processed and one or more preset motion vector offsets to obtain one or more actual motion vectors, where each actual motion vector indicates a search position; obtaining one or more candidate predicted blocks at the search positions indicated by the one or more actual motion vectors, where each search position corresponds to one candidate predicted block; and selecting, from the one or more candidate predicted blocks, the candidate predicted block with the smallest pixel difference from the temporary predicted block as the predicted block of the block to be processed.

[0041] In a third possible implementation of the third aspect, the method is used for bidirectional prediction, the reference frames include a reference frame in a first direction and a reference frame in a second direction, the initial motion vector of the block to be processed includes an initial motion vector in the first direction and an initial motion vector in the second direction, obtaining a picture block indicated by the initial motion vector of the block to be processed from the reference frame of the block to be processed, and using the obtained picture block as a temporary prediction block of the block to be processed includes obtaining a first picture block indicated by the initial motion vector in the first direction of the block to be processed from the reference frame in the first direction of the block to be processed, obtaining a second picture block indicated by the initial motion vector in the second direction of the block to be processed from the reference frame in the second direction of the block to be processed, and weighting the first picture block and the second picture block to obtain a temporary prediction block of the block to be processed.

[0042] In a fourth possible implementation of the third aspect, the method further includes rounding the motion vector resolution of the actual motion vector so that the motion vector resolution of the processed actual motion vector is equal to a preset pixel accuracy when the motion vector resolution of the actual motion vector is higher than the preset pixel accuracy.

[0043] In a fifth feasible implementation of the third aspect, the step of selecting, from one or more candidate prediction blocks, a candidate prediction block with the smallest pixel difference from a temporary prediction block as the prediction block of the block to be processed includes the step of selecting, from one or more candidate prediction blocks, an actual motion vector corresponding to the candidate prediction block with the smallest pixel difference from the temporary prediction block, and when the motion vector resolution of the selected actual motion vector is higher than a preset pixel accuracy, rounding the motion vector resolution of the selected actual motion vector so that the motion vector resolution of the processed selected actual motion vector becomes equal to the preset pixel accuracy, and determining that the prediction block corresponding to the position indicated by the processed selected actual motion vector is the prediction block of the block to be processed.

[0044] In a sixth feasible implementation of the third aspect, the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0045] In a seventh feasible implementation of the third aspect, the method further includes the step of using the predicted motion vector of a subsequent processed block to be processed after the block to be processed as the initial motion vector of the subsequent processed block to be processed after the block to be processed.

[0046] In an eighth feasible implementation of the third aspect, the method further includes the step of adding the predicted motion vector and the motion vector of a subsequent processed block to be processed after the block to be processed to obtain the initial motion vector of the subsequent processed block to be processed after the block to be processed.

[0047] In a ninth possible implementation of the third aspect, the method is used for video decoding, and the difference in motion vectors of subsequent processed blocks to be processed after the processed block is obtained by analyzing first identification information in the bitstream.

[0048] In a tenth possible implementation of the third aspect, the method is used for video decoding, and the step of determining the reference block of the processed block includes the step of analyzing the bitstream to obtain second identification information, and the step of determining the reference block of the processed block based on the second identification information.

[0049] In an eleventh possible implementation of the third aspect, the method is used for video encoding, and the step of determining the reference block of the processed block includes the step of selecting, as the reference block of the processed block, the candidate reference block with the minimum rate-distortion cost from one or more candidate reference blocks of the processed block.

[0050] According to a fourth aspect of the present application, a device for obtaining motion vectors is provided. This device can be applied on the encoder side or the decoder side. This device includes a processor and a memory. The processor and the memory are connected to each other (for example, connected to each other via a bus). In a possible implementation, this device may further include a transceiver. The transceiver is connected to the processor and the memory and is configured to receive / transmit data. The memory is configured to store program code and video data. The processor is configured to read the program code stored in the memory in order to execute the method described in the first aspect or the third aspect.

[0051] According to a fifth aspect of the present application, a video encoding system is provided. The video encoding system includes a source device and a destination device. The source device and the destination device can be communicatively connected. The source device generates encoded video data. Thus, the source device may be referred to as a video encoding device or a video encoding apparatus. The destination device can decode the encoded video data generated by the source device. Thus, the destination device may be referred to as a video decoding device or a video decoding apparatus. The source device and the destination device can be examples of a video encoding device or a video encoding apparatus. The method described in the first aspect or the third aspect is applied to a video encoding device or a video encoding apparatus.

[0052] According to a sixth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer is enabled to execute the method described in the first aspect or the third aspect.

[0053] According to a seventh aspect of the present application, a computer program product including instructions is provided. When the computer program product runs on a computer, the computer is enabled to execute the method described in the first aspect or the third aspect.

[0054] It should be understood that the embodiments corresponding to the second aspect to the seventh aspect of the present application and the embodiments corresponding to the first aspect of the present application have the same invention purpose, similar technical features, and the same beneficial technical effects. Details will not be repeatedly described.

Brief Description of the Drawings

[0055]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17A

Figure 17B

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

DETAILED DESCRIPTION OF THE INVENTION

[0056] Hereinafter, the technical solutions of the embodiments of the present application will be clearly described with reference to the accompanying drawings in the embodiments of the present application.

[0057] Figure 1 is a schematic block diagram of a video encoding system 10 according to an embodiment of the present application. As shown in Figure 1, system 10 includes a source device 12. The source device 12 generates encoded video data that is later decoded by a destination device 14. The source device 12 and the destination device 14 can include any one of a variety of devices, such as a desktop computer, a notebook computer, a tablet computer, a set-top box, a telephone handset such as a "smart" phone, a "smart" touch pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming transmission device, etc. In some applications, the source device 12 and the destination device 14 can be equipped for wireless communication.

[0058] The destination device 14 can receive the encoded video data to be decoded via a link 16. The link 16 can include any type of medium or device that can transfer the encoded video data from the source device 12 to the destination device 14. In a realizable implementation, the link 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard (e.g., a wireless communication protocol) and then transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as a radio frequency band or one or more physical transmission lines. The communication medium can form part of a packet-based network (e.g., a local area network, a wide area network, or a global network such as the Internet). The communication medium can include routers, switches, base stations, or any other device that can be used to facilitate communication from the source device 12 to the destination device 14.

[0059] Alternatively, the encoded data may be output to the storage device 24 via the output interface 22. Similarly, the encoded data from the storage device 24 may be accessed via the input interface. The storage device 24 may include any one of a plurality of scattered or local data storage media, such as a hard disk drive, a Blu-ray (registered trademark) disk, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media configured to store encoded video data. In another possible implementation, the storage device 24 may correspond to a file server or another intermediate storage device that can hold the encoded video generated by the source device 12. The destination device 14 may access the stored video data from the storage device 24 via streaming or downloading. The file server may be any type of server that can store the encoded video data and transmit the encoded video data to the destination device 14. In a possible implementation, the file server may include a website server, a file transfer protocol server, a network attached storage device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection including an Internet connection. The data connection may include a wireless channel (e.g., Wi-Fi connection) or a wired connection (e.g., cable modem) suitable for accessing the encoded video data stored in the file server, or a combination thereof. The encoded video data may be transmitted from the storage device 24 in a streaming manner, via downloading, or via a combination thereof.

[0060] The technology of this application is not necessarily limited to wireless applications or settings. This technology can be applied to video decoding to support any one of a plurality of multimedia applications, for example, wireless television broadcasting, cable television transmission, satellite television transmission, video streaming transmission (e.g., via the Internet), encoding of digital video for storage in a data storage medium, decoding of digital video stored in a data storage medium, or another application. In some possible implementations, system 10 can be configured to support unidirectional or bidirectional video transmission to support applications such as video streaming transmission, video playback, video broadcasting, and / or video telephony.

[0061] In the possible implementation of FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. In some applications, the output interface 22 can include a modulator / demodulator (modem) and / or a transmitter. In the source device 12, the video source 18 can include, for example, the following sources, namely, a video capture device (e.g., a video camera), a video archive including previously captured video, a video feed-in interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination thereof. In a possible implementation, when the video source 18 is a video camera, the source device 12 and the destination device 14 can form a camera-equipped mobile phone or a videophone. For example, the technology described in this application may be applied to video decoding or to wireless and / or wired applications.

[0062] Video encoder 20 may encode the captured or pre-captured video, or the video generated by a computer. The encoded video data may be directly transmitted to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may also (or alternatively) be stored in the storage device 24 so that the destination device 14 or another device can later access the encoded video data for decoding and / or playback.

[0063] Destination device 14 includes an input interface 28, a video decoder 30, and a display device 32. In some applications, the input interface 28 may include a receiver and / or a modem. The input interface 28 of the destination device 14 receives the encoded video data via the link 16. The encoded video data transmitted or supplied to the storage device 24 via the link 16 may include a plurality of syntax elements generated by the video encoder 20 and used by the video decoder 30 to decode the video data. These syntax elements may be included in the encoded video data transmitted on a communication medium, stored in a storage medium, or stored in a file server.

[0064] The display device 32 may be integrated with the destination device 14 or disposed outside the destination device 14. In some possible implementations, the destination device 14 can include an integrated display device and can also be configured to connect to the interface of an external display device. In other possible implementations, the destination device 14 can be a display device. Generally, the display device 32 displays the decoded video data to the user and may include any one of a plurality of display devices, such as a liquid crystal display, a plasma display, an organic light emitting diode display, or another type of display device.

[0065] The video encoder 20 and the video decoder 30 can operate, for example, in accordance with the next-generation video encoding compression standard (H.266) currently under development and can conform to the H.266 test model (JEM). Alternatively, the video encoder 20 and the video decoder 30 can operate, for example, in accordance with other proprietary or industrial standards such as ITU-T Recommendation H.265 standard or ITU-T Recommendation H.264 standard, or extensions of these standards. The ITU-T Recommendation H.265 standard is also called the High Efficiency Video Coding standard, and the ITU-T Recommendation H.264 standard is also called MPEG-4 Part 10 or Advanced Video Coding (AVC). However, the technology of this application is not limited to any specific coding standard. In other possible implementations, the video compression standards include MPEG-2 and ITU-T Recommendation H.263.

[0066] Although not shown in FIG. 1, in some embodiments, the video encoder 20 and the video decoder 30 may be integrated with an audio encoder and an audio decoder, respectively, and may include a suitable multiplexer-demultiplexer (MUX-DEMUX) unit, or other hardware and software, to encode both audio and video of the same data stream or separate data streams. If applicable, in some possible implementations, the MUX-DEMUX unit may conform to the ITU Recommendation H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).

[0067] Video encoder 20 and video decoder 30 can each be implemented as any one of a plurality of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If the technology is implemented partially in software, the apparatus can store instructions for the software in a suitable non-transitory computer-readable medium, execute the instructions in hardware by using one or more processors, and perform the technology of this application. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders. Either video encoder 20 or video decoder 30 may be integrated as part of a combined encoder / decoder (codec) of the corresponding apparatus.

[0068] This application may relate to another apparatus where, for example, video encoder 20 signals specific information to video decoder 30. However, it should be understood that video encoder 20 can associate certain syntax elements with the encoded portion of the video data to signal the information. That is, video encoder 20 can store certain syntax elements in the header information of the encoded portion of the video data to signal the data. In some applications, these syntax elements can be encoded and stored before being received and decoded by video decoder 30 (e.g., stored in storage system 34 or file server 36). Thus, the term "signal" can mean, for example, the transmission of syntax or the transmission of other data used to decode compressed video data, regardless of whether the transmission is performed in real time, near real time, or within a certain time span. For example, the transmission may be performed when the syntax elements are stored in the medium during encoding, and then the syntax elements may be re-searched by the decoding apparatus at any time after being stored in the medium.

[0069] JCT-VC is developing the H.265 (HEVC) standard. The HEVC standardization is based on an advanced model of a video decoder, which is called the HEVC Test Model (HM). The latest H.265 standard document is available at http: / / www.itu.int / rec / T-REC-H.265. The latest version of the standard document is H.265(12 / 16), and the entire standard document is incorporated herein by reference. In HM, a video decoder is assumed to have several additional features related to the existing algorithms of ITU-T Recommendation H.264 / AVC. For example, H.264 provides nine intra prediction coding modes, while HM can provide up to 35 intra prediction coding modes.

[0070] JVET is working on the development of the H.266 standard. The H.266 standardization process is based on an advanced model of a video decoder, which is called the H.266 Test Model. The description of the H.266 algorithm is available at http: / / phenix.int-evry.fr / jvet, and the description of the latest algorithm is included in JVET-G1001-v1. The algorithm description document is incorporated herein by reference in its entirety. In addition, the reference software of the JEM Test Model is available at https: / / jvet.hhi.fraunhofer.de / svn / svn_HMJEMSoftware / , and the entire reference is incorporated herein by reference.

[0071] Generally, as described by the HM motion model, a video frame or picture can be divided into a series of tree blocks, or the largest coding units (LCUs) that contain both luminance samples and chrominance samples. The LCU is also called a CTU. The tree block has a function similar to that of the macroblock in the H.264 standard. A slice contains several consecutive tree blocks in decoding order. A video frame or picture can be divided into one or more slices. Each tree block can be divided into coding units based on a quadtree. For example, a tree block functioning as the root node of a quadtree may be divided into four child nodes, and each child node can also function as a parent node and may be divided into another four child nodes. The last non - dividable child node functioning as the leaf node of the quadtree contains a decoding node, for example, a decoded video block. In the syntax data associated with the decoded bitstream, the maximum number of times a tree block can be divided and the minimum size of the decoding node can be defined.

[0072] The coding unit includes a decoding node, a prediction unit (PU), and a transform unit (TU) associated with the decoding node. The size of the CU corresponds to the size of the decoding node, and the shape of the CU needs to be square. The size of the CU can range from 8×8 pixels to the size of a tree block of up to 64×64 pixels or larger. Each CU can include one or more PUs and one or more TUs. For example, the syntax data associated with the CU can describe dividing one CU into one or more PUs. The division pattern can vary depending on whether the CU is encoded in skip mode, direct mode, intra - prediction mode, or inter - prediction mode. The PU obtained by division can have a non - square shape. For example, the syntax data associated with the CU can also describe dividing one CU into one or more TUs based on a quadtree. The TU can have a square or non - square shape.

[0073] The HEVC standard enables TU-based transformation, and the TUs can be different for different CUs. The size of a TU is usually set based on the size of the PUs within a given CU defined for the divided LCU. However, this is not always the case. The size of a TU generally is the same as or smaller than the size of the PU. In some possible implementations, a quadtree structure called "residual quadtree" (RQT) can be used to divide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can be called TUs. The pixel differences associated with the TUs can be transformed to generate transform coefficients, and the transform coefficients can be quantized.

[0074] Generally, a PU contains data related to the prediction process. For example, when a PU is encoded in an intra prediction mode, the PU may contain data describing the intra prediction mode of the PU. In another possible implementation, when a PU is encoded in an inter prediction mode, the PU may contain data defining the motion vector of the PU. For example, the data defining the motion vector of the PU may describe the horizontal component of the motion vector, the vertical component of the motion vector, the resolution of the motion vector (e.g., 1 / 4 pixel accuracy or 1 / 8 pixel accuracy), the reference picture pointed to by the motion vector, and / or the reference picture list of the motion vector (e.g., list 0, list 1, or list C).

[0075] Generally, the transform process and quantization process are used for TUs. A given CU that includes one or more PUs may also include one or more TUs. After prediction, video encoder 20 may calculate residual values corresponding to the PUs. The residual values include pixel differences. The pixel differences may be converted into transform coefficients, which are quantized and passed through a TU scan to generate serialized transform coefficients for entropy decoding. In this application, the term "video block" is commonly used to indicate a decoding node of a CU. In some specific applications of this application, the term "video block" may also be used to indicate a tree block that includes a decoding node, a PU, and a TU. For example, the tree block may be an LCU or a CU.

[0076] A video sequence typically includes a series of video frames or pictures. For example, a group of pictures (GOP) includes a series of video pictures, or one or more video pictures. The GOP may include GOP header information, header information for one or more of the pictures, or syntax data elsewhere, and the syntax data describes the number of pictures included in the GOP. Each slice of a picture may include slice syntax data that describes the coding mode of the corresponding picture. Video encoder 20 typically performs operations on video blocks in several video slices to encode video data. The video block may correspond to a decoding node of a CU. The size of the video block may be fixed or changeable and may vary depending on the specified decoding standard.

[0077] In a feasible implementation form, HM supports the prediction of various PU sizes. Assuming that the size of a given CU is 2N×2N, HM supports intra prediction with a PU size of 2N×2N or N×N, and inter prediction with a symmetric PU size of 2N×2N, 2N×N, N×2N, or N×N. HM also supports an asymmetric split of inter prediction with a PU size of 2N×nU, 2N×nD, nL×2N, and nR×2N. In an asymmetric split, the CU is not split in one direction and is split into two parts in the other direction, where one part occupies 25% of the CU and the other part occupies 75% of the CU. The part that occupies 25% of the CU is indicated by an indicator following "n" with "U (Up)", "D (Down)", "L (Left)", or "R (Right)". Thus, for example, "2N×nU" refers to a 2N×2N CU that is split horizontally and has a 2N×0.5N PU in the upper part and a 2N×1.5N PU in the lower part.

[0078] In this application, "N×M" and "N multiplied by M", for example, 16×16 pixels or 16 multiplied by 16 pixels, can be used interchangeably to indicate the pixel size of a video block in the vertical and horizontal dimensions. Generally, a 16×16 block has 16 pixels (y = 16) in the vertical direction and 16 pixels (x = 16) in the horizontal direction. Similarly, an N×N block has N pixels in the vertical direction and N pixels in the horizontal direction, where N is a non-negative integer. The pixels in a block can be arranged in rows and columns. In addition, in a block, the number of pixels in the horizontal direction and the number of pixels in the vertical direction are not necessarily the same. For example, a block can contain N×M pixels, but M is not necessarily equal to N.

[0079] After performing intra or inter prediction decoding on the PU of the CU, the video encoder 20 may calculate the residual data of the TU of the CU. The PU may include pixel data in the spatial domain (also called the pixel domain). The TU may include coefficients in the transform domain after a transform (e.g., discrete cosine transform (DCT), integer transform, wavelet transform, or a conceptually similar transform) is performed on the residual video data. The residual data may correspond to the pixel difference between the pixels of the non-encoded picture and the predictor corresponding to the PU. The video encoder 20 may generate a TU including the residual data of the CU and then transform the TU to generate the transform coefficients of the CU.

[0080] After performing some transform to generate the transform coefficients, the video encoder 20 may quantize the transform coefficients. Quantization refers to the process of quantizing the coefficients, for example, to reduce the amount of data used to represent the coefficients and to further perform compression. The quantization process can reduce the bit depth associated with some or all of the coefficients. For example, during quantization, an n-bit value may be reduced to an m-bit value by rounding, where n is greater than m.

[0081] The JEM model further improves the encoding structure of video pictures. Specifically, a block encoding structure called the "quad-tree plus binary-tree" (QTBT) structure is introduced. Without using concepts such as CU, PU, and TU in HEVC, the QTBT structure supports a more flexible divided CU shape. One CU can be square or rectangular in shape. The quadtree division is first performed on the CTU, and the binary-tree division is further performed on the leaf nodes of the quadtree. In addition, there are two division patterns for the binary-tree division: symmetric horizontal division and symmetric vertical division. The leaf nodes of the binary tree are called CUs. The CUs of the JEM model cannot be further divided during prediction and transformation. In other words, the CUs, PUs, and TUs of the JEM model have the same block size. In the existing JEM model, the maximum CTU size is 256x256 luma pixels.

[0082] In some possible implementations, the video encoder 20 may scan the quantized transform coefficients in a predefined scan order to generate a serialized vector that can be entropy encoded. In other possible implementations, the video encoder 20 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 20 may perform entropy decoding on the one-dimensional vector by using context-based adaptive variable length coding (CAVLC) or context-based adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) decoding, or another entropy decoding method. The video encoder 20 may further perform entropy encoding on the syntax elements associated with the encoded video data so that the video decoder 30 can decode the video data.

[0083] To perform CABAC, the video encoder 20 can assign a context in the context model to the symbol to be transmitted. The context may be related to whether the adjacent value of the symbol is non-zero. To perform CAVLC, the video encoder 20 can select a variable length code for the symbol to be transmitted. The codewords of variable length coding (VLC) can be constructed such that shorter codes correspond to more probable symbols and longer codes correspond to less probable symbols. In this way, using VLC can reduce the bit rate compared to using codewords of the same length for all symbols to be transmitted. The probability of CABAC can be determined based on the context assigned to the symbol.

[0084] In this embodiment of the present application, the video encoder may perform inter prediction to reduce the temporal redundancy between pictures. As described above, a CU may have one or more prediction units (PUs) according to different video compression encoding standards. In other words, a plurality of PUs may belong to a CU, or a PU and a CU may have the same size. In this specification, when a CU and a PU have the same size, the splitting pattern of the CU is no splitting, or the CU is split into one PU, and the PU is uniformly used for description. When the video encoder performs inter prediction, the video encoder may signal the motion information of the PU to the video decoder. For example, the motion information of the PU may include a reference picture index, a motion vector, and a prediction direction identifier. The motion vector may indicate the displacement between the picture block of the PU (also referred to as a video block, a pixel block, a pixel set, etc.) and the reference block of the PU. The reference block of the PU may be a part of a reference picture similar to the picture block of the PU. The reference block may be arranged in the reference picture indicated by the reference picture index and the prediction direction identifier.

[0085] To reduce the number of encoded bits required to represent the motion information of a PU, the video encoder may generate a list of candidate predicted motion vectors (MVs) for each PU according to the merge prediction mode or the advanced motion vector prediction mode. Each candidate predicted motion vector in the list of candidate predicted motion vectors of the PU may indicate motion information. The motion information indicated by some of the candidate predicted motion vectors in the list of candidate predicted motion vectors may be based on the motion information of other PUs. When a candidate predicted motion vector indicates the motion information of one of the specified spatial candidate predicted motion vector positions or the specified temporal candidate predicted motion vector positions, the candidate predicted motion vector may be referred to herein as an "original" candidate predicted motion vector. For example, in the merge mode, also referred to herein as the merge prediction mode, there may be five original spatial candidate predicted motion vector positions and one original temporal candidate predicted motion vector position. In some examples, the video encoder may generate additional candidate predicted motion vectors by combining some motion vectors from different original candidate predicted motion vectors, modifying the original candidate predicted motion vectors, or inserting only zero motion vectors as candidate predicted motion vectors. The additional candidate predicted motion vectors are not considered original candidate predicted motion vectors and may be referred to herein as artificially generated candidate predicted motion vectors in this application.

[0086] The technology of this application generally includes a technology for generating a candidate-predicted motion vector list in a video encoder and a technology for generating the same candidate-predicted motion vector list in a video decoder. The video encoder and the video decoder can generate the same candidate-predicted motion vector list by implementing the same technology for constructing the candidate-predicted motion vector list. For example, the video encoder and the video decoder can construct a list using the same number of candidate-predicted motion vectors (e.g., five candidate-predicted motion vectors). The video encoder and the video decoder can first consider spatially candidate-predicted motion vectors (e.g., adjacent blocks of the same picture), then consider temporally candidate-predicted motion vectors (e.g., candidate-predicted motion vectors of different pictures), and finally consider artificially generated candidate-predicted motion vectors until the required number of candidate-predicted motion vectors is added to the list. According to the technology of this application, during the construction of the candidate-predicted motion vector list, pruning operations may be performed on some types of candidate-predicted motion vectors to remove overlapping candidate-predicted motion vectors from the candidate-predicted motion vector list, and may not be performed on other types of candidate-predicted motion vectors to reduce the complexity of the decoder. For example, for a set of spatially candidate-predicted motion vectors and temporally candidate-predicted motion vectors, pruning operations can be performed to remove candidate-predicted motion vectors with overlapping motion information from the candidate-predicted motion vector list. However, the artificially generated candidate-predicted motion vectors can be added to the candidate-predicted motion vector list without being pruned.

[0087] After generating a list of candidate predicted motion vectors for a CU's PU, the video encoder may select a candidate predicted motion vector from the list of candidate predicted motion vectors and output a candidate predicted motion vector index to the bitstream. The selected candidate predicted motion vector may be a candidate predicted motion vector for generating a motion vector that most closely matches the predictor of the decoded target PU. The candidate predicted motion vector index may indicate the position of the selected candidate predicted motion vector within the list of candidate predicted motion vectors. The video encoder may further generate a predicted picture block for the PU based on a reference block indicated by the motion information of the PU. The motion information of the PU may be determined based on the motion information indicated by the selected candidate predicted motion vector. For example, in the merge mode, the motion information of the PU may be the same as the motion information indicated by the selected candidate predicted motion vector. In the AMVP mode, the motion information of the PU may be determined based on the difference of the motion vector of the PU and the motion information indicated by the selected candidate predicted motion vector. The video encoder may generate one or more residual picture blocks for the CU based on the predicted picture block of the CU's PU and the original picture block of the CU. Next, the video encoder may encode the one or more residual picture blocks and output the one or more residual picture blocks to the bitstream.

[0088] The bitstream may include data identifying a selected candidate predicted motion vector within a list of candidate predicted motion vectors for a PU. The video decoder may determine motion information for the PU based on the motion information indicated by the selected candidate predicted motion vector within the list of candidate predicted motion vectors for the PU. The video decoder may identify one or more reference blocks for the PU based on the motion information for the PU. After identifying the one or more reference blocks for the PU, the video decoder may generate a predicted picture block for the PU based on the one or more reference blocks for the PU. The video decoder may reconstruct a picture block for a CU based on the predicted picture block for the PU of the CU and one or more residual picture blocks for the CU.

[0089] For ease of explanation, in this application, a position or a picture block may be described as having various spatial relationships with a CU or a PU. This description may be explained as follows, i.e., the position or the picture block has various spatial relationships with a picture block associated with the CU or the PU. Additionally, in this application, a PU being currently decoded by the video decoder may be referred to as the current PU and may also be referred to as the currently processed picture block. In this application, a CU being currently decoded by the video decoder may be referred to as the current CU. In this application, a picture being currently decoded by the video decoder may be referred to as the current picture. It should be understood that this application is also applicable when the PU and the CU have the same size or when the PU is the CU and the PU is uniformly used in the description.

[0090] As briefly described above, video encoder 20 may generate predicted picture blocks and motion information of the PUs of a CU through inter prediction. In many examples, the motion information of a given PU may be the same as or similar to the motion information of one or more adjacent PUs (i.e., PUs whose picture blocks are spatially or temporally adjacent to the picture blocks of the given PU). Since adjacent PUs often have similar motion information, video encoder 20 may encode the motion information of a given PU based on the motion information of the adjacent PUs. Encoding the motion information of a given PU based on the motion information of adjacent PUs can reduce the number of encoded bits required in the bitstream to represent the motion information of the given PU.

[0091] Video encoder 20 may encode the motion information of a given PU based on the motion information of adjacent PUs in various ways. For example, video encoder 20 may indicate that the motion information of a given PU is the same as the motion information of an adjacent PU. In this application, merge mode may be used to indicate that the motion information of a given PU is the same as or can be derived from the motion information of an adjacent PU. In another possible implementation, video encoder 20 may calculate the Motion Vector Difference (MVD) of a given PU. The MVD indicates the difference between the motion vector of a given PU and the motion vector of an adjacent PU. Video encoder 20 may include the MVD in the motion information of a given PU instead of the motion vector of the given PU. In the bitstream, the number of encoded bits required to represent the MVD is less than the number of encoded bits required to represent the motion vector of a given PU. In this application, advanced motion vector prediction mode may be used to indicate that the motion information of a given PU is signaled to the decoder by using the MVD and an index value used to identify a candidate motion vector.

[0092] In merge mode or AMVP mode, to signal motion information of a given PU to a decoder, video encoder 20 may generate a candidate predicted motion vector list for the given PU. The candidate predicted motion vector list may include one or more candidate predicted motion vectors. Each of the candidate predicted motion vectors within the candidate predicted motion vector list for the given PU may specify motion information. The motion information indicated by each candidate predicted motion vector may include a motion vector, a reference picture index, and a prediction direction identifier. The candidate predicted motion vectors within the candidate predicted motion vector list may include "original" candidate predicted motion vectors, and each "original" candidate predicted motion vector indicates the motion information of one of the specified candidate predicted motion vector positions within a PU different from the given PU.

[0093] After generating the candidate predicted motion vector list for the PU, video encoder 20 may select one candidate predicted motion vector from the candidate predicted motion vector list for the PU. For example, the video encoder may compare each candidate predicted motion vector with the PU for which it is being decoded, and may select the candidate predicted motion vector having the desired rate-distortion cost. Video encoder 20 may output a candidate predicted motion vector index. The candidate predicted motion vector index may identify the position of the selected candidate predicted motion vector within the candidate predicted motion vector list.

[0094] In addition, the video encoder 20 may generate a predicted picture block of the PU based on a reference block indicated by the motion information of the PU. The motion information of the PU may be determined based on the motion information indicated by a selected candidate predicted motion vector within the candidate predicted motion vector list of the PU. For example, in the merge mode, the motion information of the PU may be the same as the motion information indicated by the selected candidate predicted motion vector. In the AMVP mode, the motion information of the PU may be determined based on the difference of the motion vector of the PU and the motion information indicated by the selected candidate predicted motion vector. As described above, the video encoder 20 may process the predicted picture block of the PU.

[0095] When the video decoder 30 receives a bitstream, the video decoder 30 may generate a candidate predicted motion vector list for each PU of the CU. The candidate predicted motion vector list generated by the video decoder 30 for the PU may be the same as the candidate predicted motion vector list generated by the video encoder 20 for the PU. The syntax element obtained by analyzing the bitstream may indicate the position of the selected candidate predicted motion vector within the candidate predicted motion vector list of the PU. After generating the candidate predicted motion vector list of the PU, the video decoder 30 may generate a predicted picture block of the PU based on one or more reference blocks indicated by the motion information of the PU. The video decoder 30 may determine the motion information of the PU based on the motion information indicated by the selected candidate predicted motion vector within the candidate predicted motion vector list of the PU. The video decoder 30 may reconstruct the picture block of the CU based on the predicted picture block of the PU and the residual picture block of the CU.

[0096] In a realizable implementation form, on the decoder, constructing the candidate-predicted motion vector list and analyzing the bitstream to obtain the position of the selected candidate-predicted motion vector in the candidate-predicted motion vector list are independent of each other and may be executed in any order or in parallel. It should be understood that this is the case.

[0097] In another realizable implementation form, on the decoder, the position of the selected candidate-predicted motion vector in the candidate-predicted motion vector list is first obtained by analyzing the bitstream, and then the candidate-predicted motion vector list is constructed based on the position obtained by the analysis. In this implementation form, it is not necessary to construct all the candidate-predicted motion vector lists. Specifically, only the candidate-predicted motion vector list at the position obtained by the analysis needs to be constructed under the condition that the candidate-predicted motion vector at that position can be determined. For example, if it is obtained by analyzing the bitstream that the selected candidate-predicted motion vector is the candidate-predicted motion vector with index 3 in the candidate-predicted motion vector list, only the candidate-predicted motion vector list from index 0 to index 3 needs to be constructed, and the candidate-predicted motion vector with index 3 can be determined. This reduces the complexity and improves the decoding efficiency.

[0098] FIG. 2 is a schematic block diagram of a video encoder 20 according to an embodiment of the present application. The video encoder 20 can perform intra coding and inter coding on video blocks in a video slice. Intra coding relies on spatial prediction to reduce or remove the spatial redundancy of video in a given video frame or picture. Inter coding relies on temporal prediction to reduce or remove the temporal redundancy of video in adjacent frames or pictures of a video sequence. The intra mode (I mode) can be any one of several spatial-based compression modes. Inter modes such as the unidirectional prediction mode (P mode) or the bidirectional prediction mode (B mode) can be any one of several time-based compression modes.

[0099] In a realizable implementation form of FIG. 2, the video encoder 20 includes a splitting unit 35, a prediction unit 41, a reference picture memory 64, an adder 50, a transformation processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra prediction unit 46. For the reconstruction of video blocks, the video encoder 20 further includes an inverse quantization unit 58, an inverse transformation unit 60, and an adder 62. The video encoder 20 may further include a deblocking filter (not shown in FIG. 2) for filtering block boundaries to remove blocking artifacts from the reconstructed video. Optionally, the deblocking filter normally filters the output of the adder 62. In addition to the deblocking filter, additional loop filters (inside or after the loop) may be used.

[0100] As shown in FIG. 2, video encoder 20 receives video data, and splitting unit 35 splits the data into video blocks. Such splitting may further include splitting into slices, picture blocks, or other larger units, as well as splitting of video blocks based on, for example, the quadtree structure of LCU and CU. For example, video encoder 20 is a component for encoding video blocks in an encoded video slice. Usually, one slice can be split into a plurality of video blocks (and can be split into a set of video blocks called picture blocks).

[0101] Prediction unit 41 can select, for the current video block, one of a plurality of possible decoding modes, for example, one of a plurality of intra decoding modes or one of a plurality of inter decoding modes, based on the encoding quality and cost calculation results (e.g., rate distortion cost, RD cost, or what is called rate distortion cost). Prediction unit 41 can provide the obtained intra decoded or inter decoded block to adder 50 to generate residual block data, reconstruct the encoded block, and provide the obtained intra decoded or inter decoded block to adder 62 for using the reconstructed encoded block as a reference picture.

[0102] To provide temporal compression, the motion estimation unit 42 and the motion compensation unit 44 of the prediction unit 41 perform inter-prediction decoding on the current video block with respect to one or more prediction blocks of one or more reference pictures. The motion estimation unit 42 may be configured to determine the inter-prediction mode of a video slice based on a pre-set mode of the video sequence. In the pre-set mode, a video slice in the sequence may be designated as a P slice, a B slice, or a GPB slice. Although the motion estimation unit 42 and the motion compensation unit 44 may be highly integrated, they are described separately for the purpose of explanation. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector for estimating a video block. For example, the motion vector may indicate the displacement of the PU of the video block in the current video frame or picture with respect to the prediction block of the reference picture.

[0103] The prediction block is a block of PUs that has been found to closely match the video block to be decoded based on pixel differences. The pixel differences may be determined based on the sum of absolute differences (SAD), the sum of squared differences (SSD), or another difference metric. In some possible implementations, the video encoder 20 may calculate the values of the sub-integer pixel positions of the reference picture stored in the reference picture memory 64. For example, the video encoder 20 may interpolate the values of the 1 / 4 pixel position, the 1 / 8 pixel position, or another fractional pixel position of the reference picture. Therefore, the motion estimation unit 42 can perform motion search for both integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy.

[0104] The motion estimation unit 42 calculates the motion vector of the PU of the video block in the inter-decoded slice by comparing the position of the PU with the position of the predicted block of the reference picture. The reference picture can be selected from the first reference picture list (list 0) or the second reference picture list (list 1). Each item in the list is used to identify one or more reference pictures stored in the reference picture memory 64. The motion estimation unit 42 transmits the calculated motion vector to the entropy encoding unit 56 and the motion compensation unit 44.

[0105] The motion compensation executed by the motion compensation unit 44 may include extracting or generating a predicted block based on the motion vector determined by motion estimation, and interpolation at the sub-pixel level may be performed. After receiving the motion vector of the PU of the current video block, the motion compensation unit 44 may identify the predicted block indicated by the motion vector in one of the reference picture lists. The video encoder 20 obtains the residual video block and subtracts the pixel value of the predicted block from the pixel value of the currently decoded video block to obtain the pixel difference. The pixel difference constitutes the residual data of the block and may include both the luminance difference component and the chrominance difference component. The adder 50 is one or more components that perform a subtraction operation. The motion compensation unit 44 may further generate syntax elements associated with the video block and the video slice for the video decoder 30 to decode the video block in the video slice.

[0106] When the PU is arranged in a B slice, the picture containing the PU may be associated with two reference picture lists called "list 0" and "list 1". In some possible implementations, the picture containing the B slice may be associated with a combination of the lists of list 0 and list 1.

[0107] In addition, when the PU is placed in the B slice, the motion estimation unit 42 may perform unidirectional prediction or bidirectional prediction of the PU. In some possible implementations, the bidirectional prediction is a prediction separately performed based on the pictures in reference picture list 0 and the pictures in reference picture list 1. In some other possible implementations, the bidirectional prediction is a prediction separately performed in display order based on the reconstructed future frame and the reconstructed past frame, which are of the current frame. When the motion estimation unit 42 performs unidirectional prediction of the PU, the motion estimation unit 42 may search for the reference block of the PU in the reference picture in list 0 or list 1. Next, the motion estimation unit 42 may generate a reference index indicating the reference picture including the reference block in list 0 or list 1, and a motion vector indicating the spatial displacement between the reference block and the PU. The motion estimation unit 42 may output the reference index, the prediction direction identifier, and the motion vector as the motion information of the PU. The prediction direction identifier may indicate that the reference index indicates a reference picture in list 0 or list 1. The motion compensation unit 44 may generate a predicted picture block of the PU based on the reference block indicated by the motion information of the PU.

[0108] When the motion estimation unit 42 performs bidirectional prediction of the PU, the motion estimation unit 42 may search for the reference block of the PU in the reference picture in list 0, and further search for another reference block of the PU in the reference picture in list 1. Next, the motion estimation unit 42 may generate a reference index indicating the reference picture including the reference blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference block and the PU. The motion estimation unit 42 may output the reference index and the motion vector of the PU as the motion information of the PU. The motion compensation unit 44 may generate a predicted picture block of the PU based on the reference block indicated by the motion information of the PU.

[0109] In some possible implementations, the motion estimation unit 42 does not output a complete set of motion information of the PU to the entropy encoding unit 56. Instead, the motion estimation unit 42 may signal the motion information of the PU by referring to the motion information of another PU. For example, the motion estimation unit 42 may determine that the motion information of the PU is similar to the motion information of an adjacent PU. In this implementation, the motion estimation unit 42 can indicate an indicator value in the syntax structure associated with the PU, and the indicator value indicates to the video decoder 30 that the motion information of the PU is the same or can be derived from the motion information of an adjacent PU. In another implementation, the motion estimation unit 42 may identify a candidate predicted motion vector and a motion vector difference (MVD) associated with an adjacent PU in the syntax structure associated with the PU. The MVD indicates the difference between the motion vector of the PU and the indicated candidate predicted motion vector associated with the adjacent PU. The video decoder 30 may use the indicated candidate predicted motion vector and the MVD to determine the motion vector of the PU.

[0110] As described above, the prediction unit 41 may generate a list of candidate predicted motion vectors for each PU of the CU. One or more of the candidate predicted motion vector lists may include one or more original candidate predicted motion vectors and one or more additional candidate predicted motion vectors derived from the one or more original candidate predicted motion vectors.

[0111] The intra prediction unit 46 within the prediction unit 41 may perform intra prediction decoding for the current video block with respect to one or more adjacent blocks in the same picture or slice as the current block to be decoded, in order to provide spatial compression. Thus, as an alternative to inter prediction (as described above) performed by the motion estimation unit 42 and the motion compensation unit 44, the intra prediction unit 46 may perform intra prediction for the current block. Specifically, the intra prediction unit 46 may determine an intra prediction mode for encoding the current block. In some possible implementations, the intra prediction unit 46 may use various intra prediction modes (for example) for encoding the current block during a separate encoding traversal, and the intra prediction unit 46 (or in some possible implementations, the mode selection unit 40) may select an appropriate intra prediction mode from the tested modes.

[0112] After the prediction unit 41 generates a predicted block of the current video block by inter prediction or intra prediction, the video encoder 20 subtracts the predicted block from the current video block to obtain a residual video block. The residual video data of the residual block is included in one or more TUs and may be applied to the transform processing unit 52. The transform processing unit 52 performs a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform (for example, discrete sine transform DST), to transform the residual video data into residual transform coefficients. The transform processing unit 52 may transform the residual video data from pixel domain data to transform domain (for example, frequency domain) data.

[0113] The conversion processing unit 52 can send the obtained conversion coefficients to the quantization unit 54. The quantization unit 54 quantizes the conversion coefficients in order to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The quantization degree can be corrected by adjusting the quantization parameter. Next, in some possible implementations, the quantization unit 54 can scan a matrix including the quantized conversion coefficients. Alternatively, the entropy encoding unit 56 can perform the scan.

[0114] After quantization, the entropy encoding unit 56 can perform entropy encoding on the quantized conversion coefficients. For example, the entropy encoding unit 56 can perform context-adaptive variable-length decoding (CAVLC) or context-adaptive binary arithmetic decoding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding, or another entropy encoding method or technique. The entropy encoding unit 56 can further perform entropy encoding on the motion vectors and other syntax elements of the currently decoded video slice. After the entropy encoding unit 56 performs entropy encoding, the encoded bitstream may be sent to the video decoder 30 or archived for subsequent transmission or re-search by the video decoder 30.

[0115] According to the technology of the present application, the entropy encoding unit 56 can encode information indicating the selected intra prediction mode. The video encoder 20 can include in the transmitted bitstream configuration data that can include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also called codeword mapping tables), definitions of the encoding contexts of various blocks, and instructions for the MPM, the intra prediction mode index table, and the modified intra prediction mode index table used for each context.

[0116] The inverse quantization unit 58 and the inverse transform unit 60 perform inverse quantization and inverse transform, respectively, to reconstruct the residual block of the pixel region so that it can be used later as a reference block of the reference picture. The motion compensation unit 44 can calculate the reference block by adding the residual block and the prediction block in one of the reference pictures in the reference picture list. The motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate sub-pixel values for motion estimation. The adder 62 adds the reconstructed residual block and the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block, and the reference block is stored in the reference picture memory 64. The reference block can be used as a reference block for performing inter prediction on the blocks of subsequent video frames or pictures by the motion estimation unit 42 and the motion compensation unit 44.

[0117] Figure 3 is a schematic block diagram of a video decoder 30 according to an embodiment of the present application. In a realizable implementation form of Figure 3, the video decoder 30 includes an entropy decoding unit 80, a prediction unit 81, an inverse quantization unit 86, an inverse transform unit 88, an adder 90, and a reference picture memory 92. The prediction unit 81 includes a motion compensation unit 82 and an intra prediction unit 84. In some realizable implementation forms, the video decoder 30 can execute an exemplary decoding process that is the reverse of the encoding process described with respect to the video encoder 20 of Figure 4.

[0118] During decoding, the video decoder 30 receives from the video encoder 20 an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements. The entropy encoding unit 80 of the video decoder 30 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. The entropy encoding unit 80 transfers the motion vectors and other syntax elements to the prediction unit 81. The video decoder 30 may receive syntax elements at the video slice level and / or the video block level.

[0119] When a video slice is decoded as an intra-decoded (I) slice, the intra prediction unit 84 of the prediction unit 81 may generate prediction data for a video block of the current video slice based on the signaled intra prediction mode and data of previously decoded blocks of the current frame or picture.

[0120] When a video picture is decoded into an inter-decoded slice (e.g., a B slice, a P slice, or a GPB slice), the motion compensation unit 82 of the prediction unit 81 generates a prediction block for a video block of the current video picture based on the motion vectors and other syntax elements received from the entropy encoding unit 80. The prediction block may be generated from one of the reference pictures in one of the reference picture lists. The video decoder 30 may use default construction techniques to construct reference picture lists (list 0 and list 1) based on the reference pictures stored in the reference picture memory 92.

[0121] The motion compensation unit 82 analyzes the motion vector and other syntax elements to determine the prediction information of the video blocks of the current video slice, and uses the prediction information to generate the prediction block of the currently decoded video block. For example, the motion compensation unit 82 uses a part of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for decoding the video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), one or more configuration information of the reference picture list regarding the slice, the motion vector of each inter-coded video block of the slice, the inter prediction state regarding each inter-decoded video block of the slice, and other information for decoding the video blocks of the current video slice.

[0122] The motion compensation unit 82 may further perform interpolation by using an interpolation filter. The motion compensation unit 82 may use, for example, the interpolation filter used by the video encoder 20 during video block encoding to calculate the interpolation value of the sub-integer pixels of the reference block. In the present application, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 based on the received syntax elements and use the interpolation filter to generate the prediction block.

[0123] When the PU is encoded by inter prediction, the motion compensation unit 82 may generate a candidate predicted motion vector list for the PU. The bitstream may include data for identifying the position of a selected candidate predicted motion vector within the candidate predicted motion vector list for the PU. After generating the candidate predicted motion vectors for the PU, the motion compensation unit 82 may generate a predicted picture block for the PU based on one or more reference blocks indicated by the motion information of the PU. The reference block of the PU may be arranged in a temporal picture different from the temporal picture of the PU. The motion compensation unit 82 may determine the motion information of the PU based on the selected motion information within the candidate predicted motion vector list for the PU.

[0124] The inverse quantization unit 86 performs inverse quantization (e.g., dequantization) on the quantized transform coefficients provided in the bitstream and decoded by the entropy encoding unit 80. The inverse quantization process may include steps of determining a quantization degree based on quantization parameters calculated by the video encoder 20 for each video block in the video slice, and similarly determining an applied inverse quantization degree. The inverse transform unit 88 performs an inverse transform (e.g., inverse DCT, inverse integer transform, or a conceptually similar inverse transform process) on the transform coefficients to generate a pixel domain residual block.

[0125] After the motion compensation unit 82 generates a predicted block of the current video block based on the motion vector and other syntax elements, the video decoder 30 adds the residual block from the inverse transform unit 88 and the corresponding predicted block generated by the motion compensation unit 82 to construct the decoded video block. The adder 90 is one or more components that perform the addition operation. Optionally, a deblocking filter may be further used to filter the decoded block to remove blocking artifacts. Another loop filter may be further used (within or after the decoding loop) to smooth the pixels, or the video quality may be improved in another way. Next, the decoded video blocks in a given frame or picture are stored in the reference picture memory 92. The reference picture memory 92 stores the reference pictures used for subsequent motion compensation. The reference picture memory 92 further stores the decoded video for later display on a display device such as the display device 32 of FIG. 1.

[0126] As described above, the technology of the present application relates to, for example, inter-decoding. The technology of the present application can be implemented by any video codec described in the present application, and it should be understood that the video decoder includes, for example, the video encoder 20 and the video decoder 30 shown and described in FIGS. 1 to 3. Specifically, in a feasible implementation, the prediction unit 41 described in FIG. 2 can execute specific techniques described below when inter-prediction is performed during the encoding of a block of video data. In another feasible implementation, the prediction unit 81 described in FIG. 3 can execute specific techniques described below when inter-prediction is performed during the decoding of a block of video data. Therefore, a general reference to a "video encoder" or "video decoder" may include the video encoder 20, the video decoder 30, or another video encoding unit or decoding unit.

[0127] FIG. 4 is a schematic block diagram of an inter prediction module according to an embodiment of the present application. The inter prediction module 121 may include, for example, a motion estimation unit 42 and a motion compensation unit 44. The relationship between the PU and the CU varies depending on the video compression encoding standard. The inter prediction module 121 may divide the current CU into PUs according to a plurality of division patterns. For example, the inter prediction module 121 may divide the current CU into PUs according to division patterns of 2N×2N, 2N×N, N×2N, and N×N. In another embodiment, the current CU is the current PU, which is not limited.

[0128] The inter prediction module 121 may perform integer motion estimation (IME) for each PU and then perform fraction motion estimation (FME). When the inter prediction module 121 performs IME for a PU, the inter prediction module 121 may search for one or more reference pictures of the reference block of the PU. After finding the reference block of the PU, the inter prediction module 121 may generate a motion vector indicating the spatial displacement between the PU and the reference block of the PU with integer accuracy. When the inter prediction module 121 performs FME for a PU, the inter prediction module 121 may improve the motion vector generated by performing IME for the PU. The motion vector generated by performing FME for a PU may have sub-integer accuracy (e.g., 1 / 2 pixel accuracy or 1 / 4 pixel accuracy). After generating the motion vector of the PU, the inter prediction module 121 may generate a predicted picture block of the PU by using the motion vector of the PU.

[0129] In some possible implementation forms where the inter prediction module 121 signals the motion information of the PU to the decoder in the AMVP mode, the inter prediction module 121 may generate a list of candidate predicted motion vectors for the PU. The list of candidate predicted motion vectors may include one or more original candidate predicted motion vectors and one or more additional candidate predicted motion vectors derived from the one or more original candidate predicted motion vectors. After generating the list of candidate predicted motion vectors for the PU, the inter prediction module 121 may select a candidate predicted motion vector from the list of candidate predicted motion vectors and generate a motion vector difference (MVD) for the PU. The MVD of the PU may indicate the difference between the motion vector indicated by the selected candidate predicted motion vector and the motion vector generated for the PU by the IME and FME. In these possible implementation forms, the inter prediction module 121 may output a candidate predicted motion vector index that identifies the position of the selected candidate predicted motion vector within the list of candidate predicted motion vectors. The inter prediction module 121 may further output the MVD of the PU. The following details a possible implementation form of the advanced motion vector prediction (AMVP) mode of FIG. 6 in the present embodiment of the present application.

[0130] In addition to performing IME and FME on the PU to generate motion information of the PU, the inter prediction module 121 may further perform a merge operation on the PU. When the inter prediction module 121 performs a merge operation on the PU, the inter prediction module 121 may generate a list of candidate predicted motion vectors of the PU. The list of candidate predicted motion vectors of the PU may include one or more original candidate predicted motion vectors and one or more additional candidate predicted motion vectors derived from the one or more original candidate predicted motion vectors. The original candidate predicted motion vectors in the list of candidate predicted motion vectors may include one or more spatial candidate predicted motion vectors and temporal candidate predicted motion vectors. The spatial candidate predicted motion vector may indicate the motion information of another PU in the current picture. The temporal candidate predicted motion vector may be based on the motion information of the corresponding PU in a picture different from the current picture. The temporal candidate predicted motion vector may also be called a temporal motion vector prediction (TMVP).

[0131] After generating the list of candidate predicted motion vectors, the inter prediction module 121 may select one candidate predicted motion vector from the list of candidate predicted motion vectors. Next, the inter prediction module 121 may generate a predicted picture block of the PU based on the reference block indicated by the motion information of the PU. In the merge mode, the motion information of the PU may be the same as the motion information indicated by the selected candidate predicted motion vector. FIG. 5 described below is a flowchart of an example of the merge mode.

[0132] The inter prediction module 121 may generate a predicted picture block of a PU by IME and FME, and after generating a predicted picture block of a PU by a merge operation, may select a predicted picture block of a PU generated by performing an FME operation or a predicted picture block of a PU generated by performing a merge operation. In some possible implementations, the inter prediction module 121 may select a predicted picture block of a PU by analyzing the rate distortion cost of a predicted picture block generated by performing an FME operation and a predicted picture block of a PU generated by performing a merge operation.

[0133] After the inter prediction module 121 selects a predicted picture block of a PU generated by splitting the current CU according to each split pattern (in some implementations, after the coding tree unit CTU is split into CUs, the CU is not further split into smaller PUs, and in this case, the PU is equivalent to the CU), the inter prediction module 121 may select the split pattern of the current CU. In some implementations, the inter prediction module 121 may select the split pattern of the current CU by analyzing the rate distortion cost of the selected predicted picture block of a PU generated by splitting the current CU according to each split pattern. The inter prediction module 121 may output the predicted picture block associated with the PU belonging to the selected split pattern to the residual generation module 102. The inter prediction module 121 may output the syntax elements of the motion information of the PU belonging to the selected split pattern to the entropy coding module 116.

[0134] In the schematic diagram shown in FIG. 4, the inter prediction module 121 may include IME modules 180A to 180N (collectively referred to as "IME module 180"), FME modules 182A to 182N (collectively referred to as "FME module 182"), merge modules 184A to 184N (collectively referred to as "merge module 184"), and PU pattern determination modules 186A to 186N (collectively referred to as "PU pattern determination module 186"), and a CU pattern determination module 188 (which may further execute a pattern determination process from a CTU to a CU).

[0135] The IME module 180, the FME module 182, and the merge module 184 may respectively perform an IME operation, an FME operation, and a merge operation on the PU of the current CU. In the schematic diagram shown in FIG. 4, the inter prediction module 121 is described as including a separate IME module 180, a separate FME module 182, and a separate merge module 184 for each PU in each split pattern of the CU. In another possible implementation, the inter prediction module 121 does not include a separate IME module 180, a separate FME module 182, or a separate merge module 184 for each PU in each split pattern of the CU.

[0136] As shown in the schematic diagram shown in FIG. 4, the IME module 180A, the FME module 182A, and the merge module 184A may respectively perform an IME operation, an FME operation, and a merge operation on the PU generated by splitting the CU according to a 2N×2N split pattern. The PU pattern determination module 186A may select one of the predicted picture blocks generated by the IME module 180A, the FME module 182A, and the merge module 184A.

[0137] The IME module 180B, the FME module 182B, and the merge module 184B can respectively perform an IME operation, an FME operation, and a merge operation on the left PU generated by dividing the CU according to an N×2N division pattern. The PU pattern determination module 186B can select one of the predicted picture blocks generated by the IME module 180B, the FME module 182B, and the merge module 184B.

[0138] The IME module 180C, the FME module 182C, and the merge module 184C can respectively perform an IME operation, an FME operation, and a merge operation on the right PU generated by dividing the CU according to an N×2N division pattern. The PU pattern determination module 186C can select one of the predicted picture blocks generated by the IME module 180C, the FME module 182C, and the merge module 184C.

[0139] The IME module 180N, the FME module 182N, and the merge module 184N can respectively perform an IME operation, an FME operation, and a merge operation on the lower-right PU generated by dividing the CU according to an N×N division pattern. The PU pattern determination module 186N can select one of the predicted picture blocks generated by the IME module 180N, the FME module 182N, and the merge module 184N.

[0140] The PU pattern determination module 186 may select a predicted picture block by analyzing the rate-distortion costs of a plurality of possible predicted picture blocks, and may select a predicted picture block that provides an optimal rate-distortion cost in a given decoding scenario. For example, in the case of an application with limited bandwidth, the PU pattern determination module 186 may select a predicted picture block with an increasing compression ratio, and in the case of another application, the PU pattern determination module 186 may select a predicted picture block with improved quality of the reconstructed video. After the PU pattern determination module 186 selects the predicted picture block of the PU of the current CU, the CU pattern determination module 188 selects the partitioning pattern of the current CU and outputs the predicted picture block and motion information of the PU belonging to the selected partitioning pattern.

[0141] FIG. 5 is a flowchart of an example of the merge mode according to an embodiment of the present application. A video encoder (e.g., video encoder 20) may perform a merge operation 200. In another possible implementation, the video encoder may perform a merge operation different from the merge operation 200. For example, in another possible implementation, the video encoder may perform a merge operation, and the video encoder may perform more or fewer steps than the steps of the merge operation 200, or steps different from the steps of the merge operation 200. In another possible implementation, the video encoder may perform the steps of the merge operation 200 in a different order or in parallel. The encoder may further perform the merge operation 200 on a PU encoded in the skip mode.

[0142] After the video encoder starts the merge operation 200, the video encoder may generate a candidate predicted motion vector list for the current PU (202). The video encoder may generate the candidate predicted motion vector list for the current PU in various ways. For example, the video encoder may generate the candidate predicted motion vector list for the current PU according to one of the exemplary techniques described below in connection with FIGS. 8 to 12.

[0143] As described above, the currently predicted motion vector list of the PU may include the temporally predicted motion vectors. The temporally predicted motion vectors may indicate the motion information of the co-located PUs at the same position in the corresponding temporal region. The co-located PUs may be spatially arranged at the same position as the current PU within the picture frame of the reference picture instead of the current picture. In the present application, the reference picture including the corresponding temporal region PU may be referred to as the associated reference picture. In the present application, the reference picture index of the associated reference picture may be referred to as the associated reference picture index. As described above, the current picture may be associated with one or more reference picture lists (e.g., list 0 and list 1). The reference picture index may indicate the reference picture by indicating the position of the reference picture within the reference picture list. In some possible implementations, the current picture may be associated with a combined reference picture list.

[0144] In some video encoders, the associated reference picture index is the reference picture index of the PU covering the reference index source position associated with the current PU. In these video encoders, the reference index source position associated with the current PU is adjacent to the left side of the current PU or adjacent to the upper part of the current PU. In the present application, a PU may "cover" a specific position if the picture block associated with the PU includes the specific position. In these video encoders, if the reference index source position is not available, the video encoder may use the reference picture index 0.

[0145] However, in one example, the reference index source position associated with the current PU is within the current CU. In this example, a PU covering the reference index source position associated with the current PU may be considered available when the PU is above or to the left of the current CU. In this case, the video encoder may need to access the motion information of another PU in the current CU to determine the reference picture including the co-located PU. Therefore, these video encoders may use the motion information (e.g., reference picture index) of the PUs belonging to the current CU to generate the temporally predicted motion vectors of the current PU. In other words, these video encoders may use the motion information of the PUs belonging to the current CU to generate the temporally predicted motion vectors. Therefore, the video encoder may not be able to generate in parallel the candidate predicted motion vector lists of the current PU and the PUs covering the reference index source position associated with the current PU.

[0146] According to the technology of the present application, the video encoder may explicitly set the relevant reference picture index without referring to the reference picture indexes of other PUs. In this way, the video encoder may generate in parallel the candidate predicted motion vector lists of the current PU and other PUs in the current CU. Since the video encoder explicitly sets the relevant reference picture index, the relevant reference picture index is not based on the motion information of other PUs in the current CU. In some possible implementations where the video encoder explicitly sets the relevant reference picture index, the video encoder may always set the relevant reference picture index to a fixed pre-set reference picture index (e.g., 0). In this way, the video encoder can generate the temporally predicted motion vectors based on the motion information of the co-located PUs in the reference frame indicated by the pre-set reference picture index, and the temporally predicted motion vectors may be included in the candidate predicted motion vector list of the current CU.

[0147] In a feasible implementation where the video encoder explicitly sets the associated reference picture index, the video encoder may explicitly signal the associated reference picture index in a syntax structure (e.g., picture header, slice header, APS, or another syntax structure). In this feasible implementation, the video encoder may signal the associated reference picture index of each LCU (i.e., CTU), CU, PU, TU, or another type of sub-block to the decoder. For example, the video encoder may signal that the associated reference picture index for each PU of a CU is equal to "1".

[0148] In some feasible implementations, the associated reference picture index may be set implicitly rather than explicitly. In these feasible implementations, the video encoder may generate each temporal candidate predicted motion vector in the candidate predicted motion vector list of the PU of the current CU by using the motion information of the PU of the reference picture indicated by the reference picture index of the PU covering the position outside the current CU, even if the position outside the current CU is not strictly adjacent to the current PU.

[0149] After generating the current PU's candidate predicted motion vector list, the video encoder may generate predicted picture blocks associated with the candidate predicted motion vectors in the candidate predicted motion vector list (204). The video encoder determines the motion information of the current PU based on the motion information of the shown candidate predicted motion vectors, and then may generate a predicted picture block based on one or more reference blocks indicated by the motion information of the current PU to generate a predicted picture block associated with the candidate predicted motion vector. Next, the video encoder may select one candidate predicted motion vector from the candidate predicted motion vector list (206). The video encoder may select the candidate predicted motion vectors in various ways. For example, the video encoder may select one candidate predicted motion vector by analyzing the rate-distortion cost of each predicted picture block associated with the candidate predicted motion vectors.

[0150] After selecting the candidate predicted motion vector, the video encoder may output a candidate predicted motion vector index (208). The candidate predicted motion vector index may indicate the position of the selected candidate predicted motion vector in the candidate predicted motion vector list. In some possible implementations, the candidate predicted motion vector index may be represented as "merge_idx".

[0151] FIG. 6 is a flowchart of an example of an advanced motion vector prediction (AMVP) mode according to an embodiment of the present application. A video encoder (e.g., video encoder 20) may perform an AMVP operation 210.

[0152] After the video encoder starts the AMVP operation 210, the video encoder may generate one or more motion vectors for the current PU (211). The video encoder may perform integer motion estimation and fractional motion estimation to generate the motion vectors of the current PU. As described above, the current picture may be associated with two reference picture lists (list 0 and list 1). When the current PU is predicted unidirectionally, the video encoder may generate a motion vector of list 0 or a motion vector of list 1 for the current PU. The motion vector of list 0 may indicate the spatial displacement between the picture block corresponding to the current PU and the reference block of the reference picture in list 0. The motion vector of list 1 may indicate the spatial displacement between the picture block corresponding to the current PU and the reference block of the reference picture in list 1. When the current PU is predicted bidirectionally, the video encoder may generate the motion vector of list 0 and the motion vector of list 1 of the current PU.

[0153] After generating one or more motion vectors of the current PU, the video encoder may generate a predicted picture block of the current PU (212). The video encoder may generate the predicted picture block of the current PU based on one or more reference blocks indicated by one or more motion vectors of the current PU.

[0154] In addition, the video coder may generate a candidate predicted motion vector list for the current PU (213). The video coder may generate the candidate predicted motion vector list for the current PU in various ways. For example, the video coder may generate the candidate predicted motion vector list for the current PU according to one or more of the realizable implementations described below in connection with FIGS. 8 to 12. In some realizable implementations, when the video coder generates a candidate predicted motion vector list by an AMVP operation 210, the candidate predicted motion vector list may be limited to two candidate predicted motion vectors. In contrast, when the video coder generates a candidate predicted motion vector list by a merge operation, the candidate predicted motion vector list may include more candidate predicted motion vectors (e.g., five candidate predicted motion vectors).

[0155] After generating the candidate predicted motion vector list for the current PU, the video coder may generate a motion vector difference (MVD) of one or more motion vectors of each candidate predicted motion vector in the candidate predicted motion vector list (214). The video coder may determine the difference between the motion vector indicated by the candidate predicted motion vector and the corresponding motion vector of the current PU to generate the motion vector difference of the candidate predicted motion vector.

[0156] When the current PU is predicted in one direction, the video coder may generate a single MVD for each candidate predicted motion vector. When the current PU is predicted in two directions, the video coder may generate two MVDs for each candidate predicted motion vector. The first MVD may indicate the difference between the motion vector indicated by the candidate predicted motion vector and the motion vector of list 0 of the current PU. The second MVD may indicate the difference between the motion vector indicated by the candidate predicted motion vector and the motion vector of list 1 of the current PU.

[0157] The video encoder may select one or more candidate predicted motion vectors from a candidate predicted motion vector list (215). The video encoder may select one or more candidate predicted motion vectors in various ways. For example, the video encoder may select a candidate predicted motion vector that matches the associated motion vector of the motion vector to be encoded with the least error. This can reduce the number of bits required to represent the difference in motion vectors of the candidate predicted motion vectors.

[0158] After selecting one or more candidate predicted motion vectors, the video encoder may output one or more reference picture indexes of the current PU, one or more candidate predicted motion vector indexes of the current PU, and the difference of one or more motion vectors of one or more selected candidate predicted motion vectors (216).

[0159] In an example where the current picture is associated with two reference picture lists (list 0 and list 1) and the current PU is predicted unidirectionally, the video encoder may output the reference picture index of list 0 ("ref_idx_10") or the reference picture index of list 1 ("ref_idx_11") of the current PU. The video encoder may further output a candidate predicted motion vector index ("mvp_10_flag") indicating the position of the selected candidate predicted motion vector of the motion vector of list 0 of the current PU within the candidate predicted motion vector list. Alternatively, the video encoder may output a candidate predicted motion vector index ("mvp_11_flag") indicating the position of the selected candidate predicted motion vector of the motion vector of list 1 of the current PU within the candidate predicted motion vector list. The video encoder may further output the MVD of the motion vector of list 0 or list 1 of the current PU.

[0160] In an example where the current picture is associated with two reference picture lists (list 0 and list 1) and the current PU is predicted bidirectionally, the video encoder may output a reference picture index of list 0 (“ref_idx_10”) and a reference picture index of list 1 (“ref_idx_11”). The video encoder may further output a candidate predicted motion vector index (“mvp_10_flag”) indicating the position of the selected candidate predicted motion vector of the motion vector of list 0 of the current PU within the candidate predicted motion vector list. Additionally, the video encoder may output a candidate predicted motion vector index (“mvp_11_flag”) indicating the position of the selected candidate predicted motion vector of the motion vector of list 1 of the current PU within the candidate predicted motion vector list. The video encoder may further output the MVD of the motion vector of list 0 of the current PU and the MVD of the motion vector of list 1 of the current PU.

[0161] FIG. 7 is a flowchart of an example of motion compensation performed by a video decoder (e.g., video decoder 30) according to an embodiment of the present application.

[0162] When the video decoder executes the motion compensation operation 220, the video decoder may receive an indication of the selected candidate predicted motion vector of the current PU (222). For example, the video decoder may receive a candidate predicted motion vector index indicating the position of the selected candidate predicted motion vector within the candidate predicted motion vector list of the current PU.

[0163] The current PU motion information is encoded in AMVP mode, and when the current PU is bi - directionally predicted, the video decoder may receive a first candidate predicted motion vector index and a second candidate predicted motion vector index. The first candidate predicted motion vector index indicates the position of the selected candidate predicted motion vector of the motion vector of list 0 of the current PU within the candidate predicted motion vector list. The second candidate predicted motion vector index indicates the position of the selected candidate predicted motion vector of the motion vector of list 1 of the current PU within the candidate predicted motion vector list. In some possible implementations, a single syntax element may be used to identify the two candidate predicted motion vector indexes.

[0164] In addition, the video decoder may generate (224) a candidate predicted motion vector list for the current PU. The video decoder may generate the candidate predicted motion vector list for the current PU in various ways. For example, the video decoder may generate the candidate predicted motion vector list for the current PU by using the techniques described below in connection with FIGS. 8 - 12. When the video decoder generates the temporal candidate predicted motion vectors of the candidate predicted motion vector list, the video decoder may explicitly or implicitly set a reference picture index that identifies a reference picture including co - located PUs, as selected above in connection with FIG. 5.

[0165] After generating the current PU's candidate predicted motion vector list, the video decoder may determine the current PU's motion information based on the motion information indicated by one or more selected candidate predicted motion vectors within the current PU's candidate predicted motion vector list (225). For example, if the current PU's motion information is encoded in merge mode, the current PU's motion information may be the same as the motion information indicated by the selected candidate predicted motion vector. If the current PU's motion information is encoded in AMVP mode, the video decoder may reconstruct one or more motion vectors of the current PU by using one or more motion vectors indicated by one or more selected candidate predicted motion vectors and one or more MVDs indicated in the bitstream. The reference picture index and prediction direction identifier of the current PU may be the same as the reference picture index and prediction direction identifier of one or more selected candidate predicted motion vectors. After determining the current PU's motion information, the video decoder may generate the current PU's predicted picture block based on one or more reference blocks indicated by the current PU's motion information (226).

[0166] FIG. 8 is a schematic diagram of an example of an encoding unit (CU) and adjacent position picture blocks associated with the encoding unit according to an embodiment of the present application. FIG. 8 is a schematic diagram for showing CU250 and schematic candidate predicted motion vector positions 252A to 252E associated with CU250. In the present application, the candidate predicted motion vector positions 252A to 252E can be collectively referred to as the candidate predicted motion vector position 252. The candidate predicted motion vector position 252 represents a spatial candidate predicted motion vector in the same picture as CU250. The candidate predicted motion vector position 252A is arranged on the left side of CU250. The candidate predicted motion vector position 252B is arranged above CU250. The candidate predicted motion vector position 252C is arranged in the upper right of CU250. The candidate predicted motion vector position 252D is arranged in the lower left of CU250. The candidate predicted motion vector position 252E is arranged in the upper left of CU250. FIG. 8 shows a schematic implementation of a manner in which the inter prediction module 121 and the motion compensation module 162 can generate a candidate predicted motion vector list. Hereinafter, this implementation will be described in relation to the inter prediction module 121. However, it should be understood that the motion compensation module 162 can implement the same technology and thus can generate the same candidate predicted motion vector list. In the present embodiment of the present application, the picture block in which the candidate predicted motion vector position is arranged is called a reference block. In addition, the reference block includes a spatial reference block, for example, the picture block in which 252A to 252E are arranged, a temporal reference block, for example, the picture block in which a co-located block is arranged, or a picture block spatially adjacent to the co-located block.

[0167] FIG. 9 is a flowchart of an example of constructing a list of candidate predicted motion vectors according to an embodiment of the present application. The technique of FIG. 9 is described based on a list including five candidate predicted motion vectors, but the techniques described herein may alternatively be used with lists of other sizes. The five candidate predicted motion vectors may each have an index (e.g., from 0 to 4). The technique of FIG. 9 is described based on a general video decoder. A general video decoder may be, for example, a video encoder (e.g., video encoder 20) or a video decoder (e.g., video decoder 30).

[0168] To reconstruct the list of candidate predicted motion vectors according to the implementation of FIG. 9, the video decoder first considers four spatially candidate predicted motion vectors (902). The four spatially candidate predicted motion vectors may include candidate predicted motion vector positions 252A, 252B, 252C, and 252D. The four spatially candidate predicted motion vectors may correspond to the motion information of four PUs arranged in the same picture as the current CU (e.g., CU250). The video decoder may consider the four spatially candidate predicted motion vectors in the list in the specified order. For example, the candidate predicted motion vector position 252A may be considered first. If the candidate predicted motion vector position 252A is available, the candidate predicted motion vector position 252A may be assigned to index 0. If the candidate predicted motion vector position 252A is not available, the video decoder may not add the candidate predicted motion vector position 252A to the candidate predicted motion vector list. The candidate predicted motion vector positions may be unavailable for various reasons. For example, if the candidate predicted motion vector position is not arranged within the current picture, the candidate predicted motion vector position may be unavailable. In another possible implementation, if the candidate predicted motion vector position passes through intra prediction, the candidate predicted motion vector position may be unavailable. In another possible implementation, if the candidate predicted motion vector position is arranged in a slice different from the slice of the current CU, the candidate predicted motion vector position may be unavailable.

[0169] After considering the candidate predicted motion vector position 252A, the video decoder may consider the candidate predicted motion vector position 252B. If the candidate predicted motion vector position 252B is available and different from the candidate predicted motion vector position 252A, the video decoder may add the candidate predicted motion vector position 252B to the candidate predicted motion vector list. In this particular context, the terms "same" or "different" mean that the motion information associated with the candidate predicted motion vector positions is the same or different. Thus, if two candidate predicted motion vector positions have the same motion information, the two candidate predicted motion vector positions are considered the same, or if two candidate predicted motion vector positions have different motion information, the two candidate predicted motion vector positions are considered different. If the candidate predicted motion vector position 252A is not available, the video decoder can assign the candidate predicted motion vector position 252B to index 0. If the candidate predicted motion vector position 252A is available, the video decoder can assign the candidate predicted motion vector position 252 to index 1. If the candidate predicted motion vector position 252B is not available or is the same as the candidate predicted motion vector position 252A, the video decoder skips adding the candidate predicted motion vector position 252B to the candidate predicted motion vector list.

[0170] Similarly, the video decoder examines the candidate predicted motion vector position 252C to determine whether to add the candidate predicted motion vector position 252C to the list. If the candidate predicted motion vector position 252C is available and different from the candidate predicted motion vector positions 252B and 252A, the video decoder can assign the candidate predicted motion vector position 252C to the next available index. If the candidate predicted motion vector position 252C is not available or is the same as at least one of the candidate predicted motion vector positions 252A and 252B, the video decoder does not add the candidate predicted motion vector position 252C to the candidate predicted motion vector list. Next, the video decoder examines the candidate predicted motion vector position 252D. If the candidate predicted motion vector position 252D is available and different from the candidate predicted motion vector positions 252A, 252B, and 252C, the video decoder can assign the candidate predicted motion vector position 252D to the next available index. If the candidate predicted motion vector position 252D is not available or is the same as at least one of the candidate predicted motion vector positions 252A, 252B, and 252C, the video decoder does not add the candidate predicted motion vector position 252D to the candidate predicted motion vector list. In the above implementation, an example of examining whether the candidate predicted motion vector positions 252A to 252D should be included in the candidate predicted motion vector list is outlined. However, in some implementations, all of the candidate predicted motion vector positions 252A to 252D can be first added to the candidate predicted motion vector list, and then the duplicate candidate predicted motion vector positions can be removed from the candidate predicted motion vector list.

[0171] After the video decoder examines the first four spatially predicted motion vectors, the list of candidate predicted motion vectors may contain the four spatially predicted motion vectors, or the list may contain fewer than four spatially predicted motion vectors. If the list contains four spatially predicted motion vectors (904, yes), the video decoder examines the temporally predicted motion vector (906). The temporally predicted motion vector may correspond to the motion information of a co-located PU in a picture different from the current picture. If the temporally predicted motion vector is available and different from the first four spatially predicted motion vectors, the video decoder assigns the temporally predicted motion vector to index 4. If the temporally predicted motion vector is not available or is the same as one of the first four spatially predicted motion vectors, the video decoder does not add the temporally predicted motion vector to the list of candidate predicted motion vectors. Thus, after the video decoder examines the temporally predicted motion vector (906), the list of candidate predicted motion vectors may contain five candidate predicted motion vectors (the first four spatially predicted motion vectors examined in 902 and the temporally predicted motion vector examined in 906), or may contain four candidate predicted motion vectors (the first four spatially predicted motion vectors examined in 902). If the list of candidate predicted motion vectors contains five candidate predicted motion vectors (908, yes), the video decoder completes the construction of the list.

[0172] If the candidate predicted motion vector list contains four candidate predicted motion vectors (908, no), the video decoder may consider a fifth spatially candidate predicted motion vector (910). The fifth spatially candidate predicted motion vector may correspond to, for example, candidate predicted motion vector position 252E. If the candidate predicted motion vector at position 252E is available and different from the candidate predicted motion vectors at positions 252A, 252B, 252C, and 252D, the video decoder may add the fifth spatially candidate predicted motion vector to the candidate predicted motion vector list and assign the fifth spatially candidate predicted motion vector to index 4. If the candidate predicted motion vector at position 252E is not available, or is the same as the candidate predicted motion vectors at candidate predicted motion vector positions 252A, 252B, 252C, and 252D, the video decoder may not add the candidate predicted motion vector at position 252E to the candidate predicted motion vector list. Thus, after the fifth spatially candidate predicted motion vector has been considered (910), the list may contain five candidate predicted motion vectors (the first four spatially candidate predicted motion vectors considered at 902 and the fifth spatially candidate predicted motion vector considered at 910), or may contain four candidate predicted motion vectors (the first four spatially candidate predicted motion vectors considered at 902).

[0173] If the candidate predicted motion vector list contains five candidate predicted motion vectors (912, yes), the video decoder completes generation of the candidate predicted motion vector list. If the candidate predicted motion vector list contains four candidate predicted motion vectors (912, no), the video decoder adds an artificially generated candidate predicted motion vector (914) until the list contains five candidate predicted motion vectors (916, yes).

[0174] After the video decoder has considered the first four spatially predicted motion vectors, if the list contains fewer than four spatially predicted motion vectors (904, no), the video decoder may consider a fifth spatially predicted motion vector (918). The fifth spatially predicted motion vector may correspond to, for example, the predicted motion vector position 252E. If the predicted motion vector at position 252E is available and different from the existing predicted motion vectors in the predicted motion vector list, the video decoder can add the fifth spatially predicted motion vector to the predicted motion vector list and assign the fifth spatially predicted motion vector to the next available index. If the predicted motion vector at position 252E is not available, or is the same as one of the existing predicted motion vectors in the predicted motion vector list, the video decoder may not add the predicted motion vector at position 252E to the predicted motion vector list. Next, the video decoder may consider a temporally predicted motion vector (920). If the temporally predicted motion vector is available and different from the existing predicted motion vectors in the predicted motion vector list, the video decoder can add the temporally predicted motion vector to the predicted motion vector list and assign the temporally predicted motion vector to the next available index. If the temporally predicted motion vector is not available, or is the same as one of the existing predicted motion vectors in the predicted motion vector list, the video decoder may not add the temporally predicted motion vector to the predicted motion vector list.

[0175] After the fifth spatially predicted motion vector (at 918) and the temporally predicted motion vector (at 920) have been considered, if the list of candidate predicted motion vectors contains five candidate predicted motion vectors (922, yes), the video decoder completes the generation of the list of candidate predicted motion vectors. If the list of candidate predicted motion vectors contains fewer than five candidate predicted motion vectors (922, no), the video decoder adds (914) artificially generated candidate predicted motion vectors until the list contains five candidate predicted motion vectors (916, yes).

[0176] According to the technology of the present application, additional merge candidate predicted motion vectors can be artificially generated after the spatially predicted motion vector and the temporally predicted motion vector such that the size of the list of merge candidate predicted motion vectors is fixed and equal to a specified number of merge candidate predicted motion vectors (e.g., five in the realizable implementation of FIG. 9 above). The additional merge candidate predicted motion vectors can include examples of a combined bi-predicted merge candidate predicted motion vector (candidate predicted motion vector 1), a scaled bi-predicted merge candidate predicted motion vector (candidate predicted motion vector 2), and a zero vector merge / AMVP candidate predicted motion vector (candidate predicted motion vector 3).

[0177] FIG. 10 is a schematic diagram of an example of adding a combined candidate motion vector to a motion vector list predicted in a merge mode candidate. The combined dual-prediction merge candidate predicted motion vector can be generated by combining the original merge candidate predicted motion vectors. Specifically, two original candidate predicted motion vectors (having mvL0 and refIdxL0 or mvL1 and refIdxL1) can be used to generate a dual-prediction merge candidate predicted motion vector. In FIG. 10, two candidate predicted motion vectors are included in the original merge candidate predicted motion vector list. The prediction type of one candidate predicted motion vector is unidirectional prediction by using list 0, and the prediction type of the other candidate predicted motion vector is unidirectional prediction by using list 1. In this feasible implementation, mvL0_A and ref0 are obtained from list 0, and mvL1_B and ref0 are obtained from list 1. Next, a dual-prediction merge candidate predicted motion vector (having mvL0_A and ref0 in list 0 and having mvL1_B and ref0 in list 1) can be generated, and it is checked whether the dual-prediction merge candidate predicted motion vector is different from the existing candidate predicted motion vectors in the candidate predicted motion vector list. If the dual-prediction merge candidate predicted motion vector is different from the existing candidate predicted motion vectors, the video decoder can add the dual-prediction merge candidate predicted motion vector to the candidate predicted motion vector list.

[0178] FIG. 11 is a schematic diagram of an example of adding a scaled candidate motion vector to a list of motion vectors predicted as merge mode candidates according to an embodiment of the present application. The scaled dual prediction merge candidate predicted motion vector can be generated by scaling the original merge candidate predicted motion vector. Specifically, one original candidate predicted motion vector (having mvLX and refIdxLX) can be used to generate the dual prediction merge candidate predicted motion vector. In a realizable implementation of FIG. 11, two candidate predicted motion vectors are included in the original merge candidate predicted motion vector list. The prediction type of one candidate predicted motion vector is unidirectional prediction by using list 0, and the prediction type of the other candidate predicted motion vector is unidirectional prediction by using list 1. In this realizable implementation, mvL0_A and ref0 are obtained from list 0, and ref0 is replicated to list 1 and can be shown as a reference index ref0'. Next, mvL0'_A can be calculated by scaling mvL0_A with ref0 and ref0'. The scaling may depend on the POC (Picture Order Count) distance. Next, a dual prediction merge candidate predicted motion vector (having mvL0_A and ref0 in list 0 and having mvL0'_A and ref0' in list 1) can be generated, and it is checked whether the dual prediction merge candidate predicted motion vectors overlap. If the dual prediction merge candidate predicted motion vectors do not overlap, they can be added to the merge candidate predicted motion vector list.

[0179] FIG. 12 is a schematic diagram of an example of adding a zero motion vector to a motion vector list for which a merge mode candidate is predicted according to an embodiment of the present application. The motion vector for which a zero vector merge candidate is predicted can be generated by combining a zero vector and a reference index that can be referred to. If the motion vectors for which a zero vector merge candidate is predicted do not overlap, they can be added to the motion vector list for which a merge candidate is predicted. The motion information of each generated motion vector for which a merge candidate is predicted can be compared with the motion information of the previously candidate-predicted motion vectors in the list.

[0180] In a realizable implementation, if the newly generated candidate-predicted motion vector is different from the existing candidate-predicted motion vectors in the candidate-predicted motion vector list, the generated candidate-predicted motion vector is added to the motion vector list for which a merge candidate is predicted. The process of determining whether the candidate-predicted motion vector is different from the existing candidate-predicted motion vectors in the candidate-predicted motion vector list is sometimes called pruning. By pruning, each newly generated candidate-predicted motion vector can be compared with the existing candidate-predicted motion vectors in the list. In some realizable implementations, the pruning operation may include comparing one or more newly generated candidate-predicted motion vectors with the existing candidate-predicted motion vectors in the candidate-predicted motion vector list and skipping the addition of newly generated candidate-predicted motion vectors that are the same as the existing candidate-predicted motion vectors in the candidate-predicted motion vector list. In some other realizable implementations, the pruning operation may include adding one or more newly generated candidate-predicted motion vectors to the candidate-predicted motion vector list and then removing the overlapping candidate-predicted motion vectors from the list.

[0181] In a feasible implementation of the present application, during inter prediction, a method for predicting motion information of a processed picture block includes: obtaining motion information of at least one picture block for which a motion vector is determined in a picture where the processed picture block is located, where at least one picture block for which the motion vector is determined is not adjacent to the processed picture block and includes the picture block for which the motion vector is determined; obtaining first identification information, where the first identification information is used to determine target motion information in the motion information of at least one picture block for which the motion vector is determined; and predicting the motion information of the processed picture block based on the target motion information.

[0182] FIG. 13 is a flowchart of an example of updating a motion vector in video encoding according to an embodiment of the present application. The block to be processed is the block to be encoded.

[0183] S1301: Obtain an initial motion vector of the block to be processed based on the predicted motion vector of the block to be processed.

[0184] In a feasible implementation, for example, in the merge mode, the predicted motion vector of the block to be processed is used as the initial motion vector of the block to be processed.

[0185] In another feasible implementation, for example, in the AMVP mode, to obtain the initial motion vector of the block to be processed, the predicted motion vector of the block to be processed is added to the difference between the motion vectors of the block to be processed.

[0186] The predicted motion vector of the block to be processed can be obtained according to any one of the methods shown in FIGS. 9 to 12 in the embodiments of the present application, or the existing methods for obtaining the predicted motion vector in the H.265 standard or the JEM reference mode. This is not limited. The difference of the motion vector is obtained by using the block to be processed as a reference, performing motion estimation within the search range determined based on the predicted motion vector of the block to be processed, and calculating the difference between the motion vector of the block to be processed obtained after the motion estimation and the predicted motion vector of the block to be processed.

[0187] During bidirectional prediction, this step specifically includes a step of obtaining a forward initial motion vector of the block to be processed based on the forward predicted motion vector of the block to be processed, and a step of obtaining a backward initial motion vector of the block to be processed based on the backward predicted motion vector of the block to be processed.

[0188] S1302: Obtain a predicted block of the block to be processed based on the initial motion vector and one or more preset motion vector offsets. Specifically,

[0189] S13021: It is of the block to be processed. Obtain a picture block indicated by the initial motion vector of the block to be processed from the reference frame indicated by the reference frame index of the block to be processed, and use the obtained picture block as a temporary predicted block of the block to be processed.

[0190] S13022: Add the initial motion vector of the block to be processed and one or more preset motion vector offsets to obtain one or more actual motion vectors, and each actual motion vector indicates a search position.

[0191] S13023: Obtain one or more candidate prediction blocks at a search position indicated by one or more actual motion vectors, where each search position corresponds to one candidate prediction block.

[0192] S13024: Select, as the prediction block of the block to be processed, the candidate prediction block with the minimum pixel difference from the temporary prediction block among the one or more candidate prediction blocks.

[0193] It should be understood that the pixel difference can be calculated in multiple ways. For example, the sum of the absolute errors between the pixel matrices of the candidate prediction block and the temporary prediction block may be calculated, or the mean squared error between the pixel matrices may be calculated, or the correlation between the pixel matrices may be calculated. This is not limited.

[0194] During bidirectional prediction, this step is for the block being processed. It obtains a first picture block indicated by the forward initial motion vector of the block being processed from the forward reference frame indicated by the forward reference frame index of the block being processed, and it is for the block being processed. It obtains a second picture block indicated by the backward initial motion vector of the block being processed from the backward reference frame indicated by the backward reference frame index of the block being processed. It weights the first picture block and the second picture block to obtain a temporary prediction block for the block being processed. It adds the forward initial motion vector of the block being processed and one or more preset motion vector offsets to obtain one or more forward actual motion vectors, and adds the backward initial motion vector of the block being processed and one or more preset motion vector offsets to obtain one or more backward actual motion vectors. It obtains one or more forward candidate prediction blocks at the search positions indicated by the one or more forward actual motion vectors, and obtains one or more backward candidate prediction blocks at the search positions indicated by the one or more backward actual motion vectors. It selects, from the one or more forward candidate prediction blocks, the candidate prediction block with the smallest pixel difference from the temporary prediction block as the forward prediction block of the block being processed, and selects, from the one or more backward candidate prediction blocks, the candidate prediction block with the smallest pixel difference from the temporary prediction block as the backward prediction block of the block being processed. It weights the forward prediction block and the backward prediction block to obtain the prediction block of the block being processed, specifically including these steps.

[0195] Optionally, in a feasible implementation form, after step S13022, this method further includes the following steps.

[0196] S13025: If the motion vector resolution of the actual motion vector is higher than the pre-set pixel accuracy, round the motion vector resolution of the actual motion vector so that it equals the pre-set pixel accuracy. The pre-set pixel accuracy can be integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy, etc., and this is not limited.

[0197] It should be understood that the motion vector resolution is the pixel accuracy that can be distinguished by the motion vector in the motion estimation or motion compensation process. Rounding can include rounding up, rounding down, etc. based on the type of pixel accuracy. This is not limited.

[0198] For example, rounding can include the following operations.

[0199] The horizontal or vertical component of the motion vector to be processed is decomposed into an integer part a, a fractional part b, and a sign bit. Naturally, a is a non-negative integer, b is a fraction greater than 0 and less than 1, and the sign bit is positive or negative.

[0200] It can be assumed that the pre-set pixel accuracy is N pixel accuracy, where N is greater than 0 and less than or equal to 1, and c is equal to the value obtained by dividing b by N.

[0201] When the rounding rule is used, the fractional part of c is rounded. When the rounding-up rule is used, the integer part of c is incremented by 1 and the fractional part is discarded. When the rounding-down rule is used, the fractional part of c is discarded. It can be assumed that the c obtained after processing is d.

[0202] The absolute value of the processed motion vector component is obtained by multiplying d by N and then adding a, and the positive or negative sign of the motion vector component is not changed.

[0203] For example, when the actual motion vector is (1.25, 1) and the pre-set pixel accuracy is integer pixel accuracy, the actual motion vector is rounded to (1, 1). When the actual motion vector is (-1.7, -1) and the pre-set pixel accuracy is 1 / 4 pixel accuracy, the actual motion vector is rounded to (-1.75, -1).

[0204] In some cases, in another possible implementation, step S13024 includes selecting, from one or more candidate prediction blocks, the actual motion vector corresponding to the candidate prediction block with the minimum pixel difference from the temporary prediction block; when the motion vector resolution of the selected actual motion vector is higher than the pre-set pixel accuracy, rounding the motion vector resolution of the selected actual motion vector so that the processed motion vector resolution of the selected actual motion vector is equal to the pre-set pixel accuracy; and determining that the prediction block corresponding to the position indicated by the processed selected actual motion vector is the prediction block of the block to be processed.

[0205] Similarly, the pre-set pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy. This is not limiting. For rounding, refer to the examples of the possible implementations described above. Details will not be repeated.

[0206] It should be understood that a higher pixel accuracy generally means that more complex pixel interpolation needs to be performed in the search area of the motion estimation or motion compensation process to make the motion vector resolution equal to the pre-set pixel accuracy. This can reduce complexity.

[0207] FIG. 14 is a flowchart of an example of updating a motion vector in video decoding according to an embodiment of the present application. The block to be processed is the block to be decoded.

[0208] S1401: Obtain an initial motion vector of a block to be processed based on a predicted motion vector of the block to be processed.

[0209] In a realizable implementation form, for example, in the merge mode, the predicted motion vector of the block to be processed is used as the initial motion vector of the block to be processed.

[0210] In another realizable implementation form, for example, in the AMVP mode, in order to obtain the initial motion vector of the block to be processed, the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed are added together.

[0211] The predicted motion vector of the block to be processed can be obtained according to any one of the methods shown in FIGS. 9 to 12 in the embodiments of the present application, or the existing methods for obtaining the predicted motion vector in the H.265 standard or the JEM reference mode. This is not limited. The difference between the motion vectors can be obtained by analyzing the bitstream.

[0212] During bidirectional prediction, this step specifically includes the step of obtaining a forward initial motion vector of the block to be processed based on the forward predicted motion vector of the block to be processed, and the step of obtaining a backward initial motion vector of the block to be processed based on the backward predicted motion vector of the block to be processed.

[0213] S1402: Obtain a predicted block of the block to be processed based on the initial motion vector and one or more pre-set motion vector offsets. Specifically,

[0214] S14021: It pertains to the block to be processed. A picture block indicated by the initial motion vector of the block to be processed is obtained from the reference frame indicated by the reference frame index of the block to be processed, and the obtained picture block is used as the temporary prediction block of the block to be processed.

[0215] S14022: To obtain one or more actual motion vectors, the initial motion vector of the block to be processed is added to one or more pre-set motion vector offsets, and each actual motion vector indicates a search position.

[0216] S14023: One or more candidate prediction blocks are obtained at the search positions indicated by one or more actual motion vectors, and each search position corresponds to one candidate prediction block.

[0217] S14024: From one or more candidate prediction blocks, the candidate prediction block with the minimum pixel difference from the temporary prediction block is selected as the prediction block of the block to be processed.

[0218] It should be understood that the pixel difference can be calculated in multiple ways. For example, the sum of the absolute errors between the pixel matrices of the candidate prediction block and the temporary prediction block may be calculated, or the mean squared error between the pixel matrices may be calculated, or the correlation between the pixel matrices may be calculated. This is not limited.

[0219] During bidirectional prediction, this step is for the block being processed. It obtains a first picture block indicated by the forward initial motion vector of the block being processed from the forward reference frame indicated by the forward reference frame index of the block being processed, and a second picture block indicated by the backward initial motion vector of the block being processed from the backward reference frame indicated by the backward reference frame index of the block being processed. It also includes the steps of weighting the first picture block and the second picture block to obtain a temporary prediction block for the block being processed; adding the forward initial motion vector of the block being processed and one or more preset motion vector offsets to obtain one or more forward actual motion vectors, and adding the backward initial motion vector of the block being processed and one or more preset motion vector offsets to obtain one or more backward actual motion vectors; obtaining one or more forward candidate prediction blocks at the search positions indicated by the one or more forward actual motion vectors, and obtaining one or more backward candidate prediction blocks at the search positions indicated by the one or more backward actual motion vectors; selecting, from the one or more forward candidate prediction blocks, the candidate prediction block with the minimum pixel difference from the temporary prediction block as the forward prediction block of the block being processed, and selecting, from the one or more backward candidate prediction blocks, the candidate prediction block with the minimum pixel difference from the temporary prediction block as the backward prediction block of the block being processed, and weighting the forward prediction block and the backward prediction block to obtain the prediction block of the block being processed.

[0220] Optionally, in a feasible implementation form, after step S14022, this method further includes the following steps.

[0221] S14025: If the motion vector resolution of the actual motion vector is higher than the preset pixel accuracy, round the motion vector resolution of the actual motion vector so that it is equal to the preset pixel accuracy. The preset pixel accuracy can be integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy, but is not limited to this.

[0222] Optionally, in another possible implementation, step S14024 includes selecting the actual motion vector corresponding to the candidate prediction block with the smallest pixel difference from one or more candidate prediction blocks to a temporary prediction block; if the motion vector resolution of the selected actual motion vector is higher than the preset pixel accuracy, rounding the motion vector resolution of the selected actual motion vector so that it is equal to the preset pixel accuracy; and determining that the prediction block corresponding to the position indicated by the processed selected actual motion vector is the prediction block of the block to be processed.

[0223] Similarly, the preset pixel accuracy can be integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy, but is not limited to this. For rounding, refer to the examples of the above possible implementations. Details will not be repeated.

[0224] By using some specific embodiments, the following will detail the implementation of updating the motion vector. As described in the encoding method of FIG. 13 and the decoding method of FIG. 14, it should be understood that the update of the motion vector maintains consistency between the encoder and the decoder. Therefore, the following embodiments will be described only from the encoder or the decoder. It should be understood that when the description is provided from the encoder, the implementation in the decoder is consistent with the implementation in the encoder, and when the description is provided from the decoder, the implementation in the encoder is consistent with the implementation in the decoder.

[0225] Embodiment 1 As shown in FIG. 15, the current decoding block is the first decoding block, and the predicted motion information of the current decoding block is obtained. The forward motion vector predictor and the backward motion vector predictor of the current decoding block are (-10, 4) and (5, 6), respectively. Assume that the POC of the picture where the current decoding block is located is 4, which is that of the reference picture, and the POCs indicated by the index values of the reference pictures are 2 and 6, respectively. Therefore, the POC corresponding to the current decoding block is 4, the POC corresponding to the forward prediction reference picture block is 2, and the POC corresponding to the backward prediction reference picture block is 6.

[0226] Forward prediction and backward prediction are separately executed for the current decoding block to obtain the initial forward decoding prediction block (Forward prediction Block, FPB) and the initial backward decoding prediction block (Backward Prediction Block, BPB) of the current decoding block. Assume that the initial forward decoding prediction block and the initial backward decoding prediction block are FPB1 and BPB1, respectively. The first decoding prediction block (Decoding Prediction Block, DPB) of the current decoding block is obtained by performing a weighted sum of FPB1 and BPB1 and is assumed to be DPB1.

[0227] (-10, 4) and (5, 6) are used as reference inputs for the forward motion vector predictor and the backward motion vector predictor, and the motion search with the first accuracy is performed separately for the forward prediction reference picture block and the backward prediction reference picture block. In this case, the first accuracy is 1 / 2 pixel accuracy within a 1-pixel range. The first decoded prediction block DPB1 is used as a reference. The corresponding new forward and backward decoded prediction blocks obtained in each motion search are compared with the first decoded prediction block DPB1 to obtain a new decoded prediction block with the minimum difference from DPB1, and the forward and backward motion vector predictors corresponding to the new decoded prediction block are used as the target motion vector predictors, which are assumed to be (-11, 4) and (6, 6) respectively.

[0228] The target motion vector predictors are updated to (-11, 4) and (6, 6), forward prediction and backward prediction are performed on the first decoded block based on the target motion vector predictors, and the target decoded prediction block is obtained by taking the weighted sum of the new forward and backward decoded prediction blocks obtained, which is assumed to be DPB2, and the decoded prediction block of the current decoded block is updated to DPB2.

[0229] It should be noted that when the motion search with the first accuracy is performed for the forward prediction reference picture block and the backward prediction reference picture block, the first accuracy can be any specified accuracy, for example, integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0230] Embodiment 2 As shown in FIG. 16, the current decoding block is the first decoding block, and the predicted motion information of the current decoding block is obtained. The forward motion vector predictor of the current decoding block is (-21, 18), the POC of the picture in which the current decoding block is located is 4, which is that of the reference picture, and it is assumed that the POC indicated by the index value of the reference picture is 2. Therefore, the POC corresponding to the current decoding block is 4, and the POC corresponding to the forward predicted reference picture block is 2.

[0231] Forward prediction is performed on the current decoding block to obtain the initial forward decoding prediction block of the current decoding block, and it is assumed that the initial forward decoding prediction block is FPB1. In this case, FPB1 is used as the first decoding prediction block of the current decoding block, and the first decoding prediction block is shown as DPB1.

[0232] (-21, 18) is used as the reference input of the forward motion vector predictor, and motion search with the first accuracy is performed on the forward predicted reference picture block. In this case, the first accuracy is 1-pixel accuracy within a 5-pixel range. The first decoding prediction block DPB1 is used as a reference. The corresponding new forward decoding prediction blocks obtained in each motion search are compared with the first decoding prediction block DPB1 to obtain a new decoding prediction block with the smallest difference from DPB1, and the forward motion vector predictor corresponding to the new decoding prediction block is used as the target motion vector predictor, and it is assumed that they are (-19, 19) respectively.

[0233] The target motion vector predictor is updated to (-19, 19), forward prediction is performed on the first decoding block based on the target motion vector predictor, and the obtained new forward decoding prediction block is used as the target decoding prediction block, and it is assumed that it becomes DPB2, and the decoding prediction block of the current decoding block is updated to DPB2.

[0234] When motion search with a first accuracy is performed on the forward prediction reference picture block and the backward prediction reference picture block, it should be noted that the first accuracy can be any specified accuracy, for example, integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0235] Embodiment 3 As shown in FIGS. 17A and 17B, the current coding block is the first coding block, and the predicted motion information of the current coding block is obtained. Assume that the forward motion vector predictor and the backward motion vector predictor of the current coding block are (-6, 12) and (8, 4) respectively, the POC of the picture where the current coding block is located is 8, which is that of the reference picture, and the POCs indicated by the index values of the reference pictures are 4 and 12 respectively. Therefore, the POC corresponding to the current coding block is 4, the POC corresponding to the forward prediction reference picture block is 4, and the POC corresponding to the backward prediction reference picture block is 12.

[0236] The forward prediction and the backward prediction are separately executed for the current coding block to obtain the initial forward coding prediction block and the initial backward coding prediction block of the current coding block. Assume that the initial forward coding prediction block and the initial backward coding prediction block are FPB1 and BPB1 respectively. The first coding prediction block of the current coding block is obtained by performing a weighted sum of FPB1 and BPB1 and is assumed to be DPB1.

[0237] (-6, 12) and (8, 4) are used as reference inputs for the forward motion vector predictor and the backward motion vector predictor, and the motion search with the first accuracy is performed separately for the forward prediction reference picture block and the backward prediction reference picture block. The first encoded prediction block DPB1 is used as a reference. The corresponding new forward and backward encoded prediction blocks obtained in each motion search are compared with the first encoded prediction block DPB1 to obtain a new encoded prediction block with the minimum difference from DPB1, and the forward and backward motion vector predictors corresponding to the new encoded prediction block are used as target motion vector predictors, which are assumed to be (-11, 4) and (6, 6) respectively.

[0238] The target motion vector predictors are updated to (-11, 4) and (6, 6), forward prediction and backward prediction are performed on the first encoded block based on the target motion vector predictors, and the target encoded prediction block is obtained by taking the weighted sum of the obtained new forward and backward encoded prediction blocks, which is assumed to be DPB2, and the encoded prediction block of the current encoded block is updated to DPB2.

[0239] Next, (-11, 4) and (6, 6) are used as reference inputs for the forward motion vector predictor and the backward motion vector predictor, and the motion search with the first accuracy is performed separately for the forward prediction reference picture block and the backward prediction reference picture block. The encoded prediction block DPB2 of the current encoded block is used as a reference. The corresponding new forward and backward encoded prediction blocks obtained in each motion search are compared with the first encoded prediction block DPB2 to obtain a new encoded prediction block with the minimum difference from DPB2, and the forward and backward motion vector predictors corresponding to the new encoded prediction block are used as new target motion vector predictors, which are assumed to be (-7, 11) and (6, 5) respectively.

[0240] Next, the target motion vector predictor is updated to (-7, 11) and (6, 5), forward prediction and backward prediction are performed on the first encoded block based on the latest target motion vector predictor, and the weighted sum of the obtained new forward and backward encoded prediction blocks is taken to obtain the target encoded prediction block, which is assumed to be DPB3, and the encoded prediction block of the current encoded block is updated to DPB3.

[0241] Furthermore, the target motion vector predictor can be continuously updated according to the method described above, and the number of cycles is not limited.

[0242] It should be noted that when motion search with the first accuracy is performed on the forward prediction reference picture block and the backward prediction reference picture block, the first accuracy can be any specified accuracy, for example, integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0243] It should be understood that in some executable embodiments, the cycle ends when the condition is met. For example, the cycle ends when the difference between DPBn and DPBn - 1 is smaller than the threshold, where n is a positive integer greater than 2.

[0244] Embodiment 4 As shown in FIG. 18, the current decoding block is the first decoding block, and the predicted motion information of the current decoding block is obtained. The predicted values of the forward and backward motion vectors of the current decoding block are (-10, 4) and (5, 6) respectively, the differences between the forward and backward motion vectors of the current decoding block are (-2, 1) and (1, 1) respectively, the POC of the picture in which the current decoding block is located is 4, which is that of the reference picture, and it is assumed that the POCs indicated by the index values of the reference pictures are 2 and 6 respectively. Therefore, the POC corresponding to the current decoding block is 4, the POC corresponding to the forward predicted reference picture block is 2, and the POC corresponding to the backward predicted reference picture block is 6.

[0245] The forward prediction and the backward prediction are separately executed for the current decoding block to obtain the initial forward decoding prediction block (FPB) and the initial backward decoding prediction block (BPB) of the current decoding block, and it is assumed that the initial forward decoding prediction block and the initial backward decoding prediction block are FPB1 and BPB1 respectively. The first decoding prediction block (DPB) of the current decoding block is obtained by performing a weighted sum of FPB1 and BPB1, and it is assumed to be DPB1.

[0246] The sum of the forward motion vector predictor and the difference of the forward motion vector, and the sum of the backward motion vector predictor and the difference of the backward motion vector, i.e., (-10,4)+(-2,1)=(-12,5) and (5,6)+(1,1)=(6,7), are used as the forward motion vector and the backward motion vector respectively, and motion search with the first accuracy is performed separately for the forward prediction reference picture block and the backward prediction reference picture block. In this case, the first accuracy is 1 / 4 pixel accuracy within a 1 pixel range. The first decoded prediction block DPB1 is used as a reference. The corresponding new forward and backward decoded prediction blocks obtained in each motion search are compared with the first decoded prediction block DPB1 to obtain a new decoded prediction block with the minimum difference from DPB1, and the forward and backward motion vectors corresponding to the new decoded prediction block are used as target motion vector predictors, assumed to be (-11,4) and (6,6) respectively.

[0247] The target motion vectors are updated to (-11,4) and (6,6), forward prediction and backward prediction are performed separately for the first decoded block based on the target motion vectors, and the target decoded prediction block is obtained by taking the weighted sum of the new forward and backward decoded prediction blocks obtained, assumed to be DPB2, and the decoded prediction block of the current decoded block is updated to DPB2.

[0248] FIG. 19 is a schematic flowchart of a method for obtaining a motion vector by an encoder according to an embodiment of the present application. The method includes the following steps.

[0249] S1901: Determine the reference block of the block to be processed.

[0250] The reference block has been described above in connection with FIG. 8. It should be understood that the reference block includes not only the spatially adjacent blocks of the block to be processed shown in FIG. 8, but also other actual or virtual picture blocks having a pre-set temporal or spatial correlation with the block to be processed.

[0251] It should be understood that the beneficial effect of the present embodiment of the present application is reflected in the scenario where the motion vector of the reference block of the block to be processed is updated. Specifically, the reference block has an initial motion vector and one or more pre-set motion vector offsets, the initial motion vector of the reference block is obtained based on the predicted motion vector of the reference block, and the predicted block of the reference block is obtained based on the initial motion vector and one or more pre-set motion vector offsets.

[0252] Specifically, for the process of updating the motion vector of the reference block and obtaining the initial motion vector, reference may be made to the embodiment related to FIG. 13 of the present application. It should be understood that the reference block in the embodiment related to FIG. 19 is the block to be processed in the embodiment related to FIG. 13.

[0253] In some possible implementations, the step of determining the reference block of the block to be processed specifically includes the step of selecting, as the reference block of the block to be processed, the candidate reference block with the minimum rate distortion cost from one or more candidate reference blocks of the block to be processed.

[0254] In some possible implementations, after determining the reference block of the block to be processed in one or more candidate reference blocks of the block to be processed, the method further includes the step of encoding the identification information of the determined reference block in the one or more candidate reference blocks into a bit stream.

[0255] S1902: Use the initial motion vector of the reference block as the predicted motion vector of the block to be processed.

[0256] In some possible implementations, for example, in the merge mode, after step S1902, the method further includes a step of using the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed. Alternatively, in step S1902, the initial motion vector of the reference block is used as the initial motion vector of the block to be processed.

[0257] In another possible implementation, for example, in the AMVP mode, after step S1902, the method further includes a step of adding the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed to obtain the initial motion vector of the block to be processed.

[0258] FIG. 20 is a schematic flowchart of a method for obtaining a motion vector by a decoder according to an embodiment of the present application. The method includes the following steps.

[0259] S2001: Determine the reference block of the block to be processed.

[0260] It should be understood that the beneficial effects of the present embodiment of the present application are reflected in the scenario where the motion vector of the reference block of the block to be processed is updated. Specifically, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on the predicted motion vector of the reference block, and the predicted block of the reference block is obtained based on the initial motion vector and one or more preset motion vector offsets.

[0261] Specifically, for the process of updating the motion vector of the reference block and obtaining the initial motion vector, refer to the embodiments related to FIG. 14 of the present application. It should be understood that the reference block in the embodiment related to FIG. 20 is the block to be processed in the embodiment related to FIG. 14.

[0262] In some possible implementations, the step of determining the reference block of the block to be processed specifically includes the step of analyzing the bitstream to obtain the second identification information and the step of determining the reference block of the block to be processed based on the second identification information.

[0263] S2002: Use the initial motion vector of the reference block as the predicted motion vector of the block to be processed.

[0264] In a possible implementation, for example, in the merge mode, after step S2002, the method further includes the step of using the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed. Alternatively, in step S2002, the initial motion vector of the reference block is used as the initial motion vector of the block to be processed.

[0265] In another possible implementation, for example, in the AMVP mode, after step S2002, the method further includes the step of adding the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed to obtain the initial motion vector of the block to be processed, where the difference between the motion vectors of the block to be processed is obtained by analyzing the first identification information in the bitstream.

[0266] In the above-described implementation form, the initial motion vector before update is used to replace the actual motion vector and to predict subsequent encoding blocks or decoding blocks. Before the update of the actual motion vector is completed, a prediction step can be executed for subsequent encoding blocks or decoding blocks. This ensures the improvement in encoding efficiency brought about by the update of the motion vector and eliminates the delay in processing.

[0267] FIG. 21 is a schematic block diagram of an apparatus 2100 for obtaining a motion vector according to an embodiment of the present application. The apparatus 2100 is configured to determine a reference block of a block to be processed, where the reference block and the block to be processed have a preset temporal or spatial correlation relationship, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on the predicted motion vector of the reference block, and the predicted block of the reference block is obtained based on the initial motion vector and one or more preset motion vector offsets, and a determination module 2101 and an acquisition module 2102 configured to use the initial motion vector of the reference block as the predicted motion vector of the block to be processed. The apparatus 2100 includes the determination module 2101 and the acquisition module 2102.

[0268] In a realizable implementation form, the acquisition module 2102 is further configured to use the predicted motion vector of the reference block as the initial motion vector of the reference block, or to add the difference between the predicted motion vector of the reference block and the motion vector of the reference block to obtain the initial motion vector of the reference block.

[0269] In a possible implementation, the acquisition module 2102 acquires a picture block indicated by the initial motion vector of the reference block from the reference frame of the reference block, uses the acquired picture block as a temporary prediction block of the reference block, adds the initial motion vector of the reference block and one or more preset motion vector offsets to obtain one or more actual motion vectors, where each actual motion vector indicates a search position, acquires one or more candidate prediction blocks at the search positions indicated by the one or more actual motion vectors, each search position corresponds to one candidate prediction block, and further selects, as the prediction block of the reference block, the candidate prediction block with the minimum pixel difference from the temporary prediction block among the one or more candidate prediction blocks.

[0270] In a possible implementation, the apparatus 2100 is configured for bidirectional prediction, the reference frame includes a reference frame in a first direction and a reference frame in a second direction, the initial motion vector includes an initial motion vector in the first direction and an initial motion vector in the second direction, the acquisition module 2102 acquires a first picture block indicated by the initial motion vector in the first direction of the reference block from the reference frame in the first direction of the reference block, acquires a second picture block indicated by the initial motion vector in the second direction of the reference block from the reference frame in the second direction of the reference block, and is specifically configured to weight the first picture block and the second picture block to obtain a temporary prediction block of the reference block.

[0271] In a possible implementation, when the motion vector resolution of the actual motion vector is higher than the preset pixel accuracy, the apparatus 2100 further includes a rounding module 2103 configured to round the motion vector resolution of the actual motion vector so that the motion vector resolution of the processed actual motion vector is equal to the preset pixel accuracy.

[0272] In a possible implementation form, the acquisition module 2102 selects, from one or more candidate prediction blocks, an actual motion vector corresponding to the candidate prediction block with the smallest pixel difference from the temporary prediction block. When the motion vector resolution of the selected actual motion vector is higher than a preset pixel accuracy, the motion vector resolution of the selected actual motion vector is rounded so that the motion vector resolution of the processed selected actual motion vector is equal to the preset pixel accuracy, and it is specifically configured such that the prediction block corresponding to the position indicated by the processed selected actual motion vector is determined to be the prediction block of the reference block.

[0273] In a possible implementation form, the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

[0274] In a possible implementation form, the acquisition module 2102 is specifically configured to use the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed.

[0275] In a possible implementation form, the acquisition module 2102 is specifically configured to add the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed to obtain the initial motion vector of the block to be processed.

[0276] In a possible implementation form, the apparatus 2100 is configured for video decoding, and the difference between the motion vectors of the block to be processed is obtained by analyzing first identification information in the bitstream.

[0277] In a possible implementation form, the apparatus 2100 is configured for video decoding, and the determination module 2101 is specifically configured to analyze the bitstream to obtain second identification information and determine the reference block of the block to be processed based on the second identification information.

[0278] In a possible implementation form, the apparatus 2100 is configured for video encoding, and the determination module 2101 is specifically configured to select, from one or more candidate reference blocks of the block to be processed, the candidate reference block with the minimum rate distortion cost as the reference block of the block to be processed.

[0279] FIG. 22 is a schematic block diagram of a video encoding device according to an embodiment of the present application. The device 2200 may be applied to an encoder or a decoder. The device 2200 includes a processor 2201 and a memory 2202. The processor 2201 and the memory 2202 are connected to each other (for example, connected to each other via a bus 2204). In a possible implementation form, the device 2200 may further include a transceiver 2203. The transceiver 2203 is connected to the processor 2201 and the memory 2202 and is configured to receive / transmit data.

[0280] The memory 2202 includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory 2202 is configured to store related program codes and video data.

[0281] The processor 2201 may be one or more central processing units (CPUs). When the processor 2201 is one CPU, the CPU may be a single-core CPU or a multi-core CPU.

[0282] The processor 2201 is configured to read the program code stored in the memory 2202 and execute operations in any implementation solution corresponding to FIGS. 13 to 20 and various possible implementation forms of the implementation solution.

[0283] For example, an embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores instructions. When the instructions are executed on a computer, the computer can execute operations in any implementation solution corresponding to FIGS. 13 to 20 and various possible implementation forms of the implementation solution.

[0284] For example, an embodiment of the present application further provides a computer program product including instructions. When the computer program product is executed on a computer, the computer can execute operations in any implementation solution corresponding to FIGS. 13 to 20 and various possible implementation forms of the implementation solution.

[0285] Those skilled in the art should be aware that, in combination with the examples described in each embodiment disclosed herein, the units and algorithm steps can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether the function is performed by hardware or software depends on the specific application of the technical solution and the design constraints. Those skilled in the art can use different methods to implement the functions described for each specific application form, but such implementations should not be considered as exceeding the scope of the present application.

[0286] Those skilled in the art should clearly understand that, for the sake of simplicity of description, for the detailed operation processes of the above systems, devices and units, reference should be made to the corresponding processes in the above method embodiments, and the details will not be repeated here.

[0287] All or part of the above-described embodiments may be implemented by using software, hardware, firmware, or any combination thereof. When software is used to implement an embodiment, the embodiment may be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the procedures or functions according to the embodiments of the present invention are generated. The computer may be a general-purpose computer, a dedicated computer, a computer network, or another programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted in a wired (e.g., coaxial cable, optical fiber, or digital subscriber line) or wireless (e.g., infrared or microwave) manner from a website, computer, server, or data center to another website, computer, server, or data center. The computer-readable storage medium may be any usable medium accessible by a computer, or a data storage device such as a server or data center that integrates one or more usable media. The usable medium may be a magnetic medium (e.g., floppy disk, hard disk, or magnetic tape), an optical medium (e.g., DVD), a semiconductor medium (e.g., solid state disk), etc.

[0288] In the above-described embodiments, each embodiment description has its own focus. For parts not described in detail in one embodiment, please refer to the relevant descriptions in other embodiments.

[0289] The above description is merely a specific implementation of the present invention and is not intended to limit the protection scope of the present invention. Any modifications and replacements that are within the technical scope disclosed in the present invention and can be easily conceived by those skilled in the art shall be included within the protection scope of the present invention. Therefore, the protection scope of the present invention belongs to the protection scope of the claims.

Description of Reference Signs

[0290] 10 Video Encoding System 12 Source Device 14 Destination Device 18 Video Source 20 Video Encoder 22 Output Interface 24 Storage Device 28 Input Interface 30 Video Decoder 32 Display Device 34 Storage System 35 Splitting Unit 36 File Server 40 Mode Selection Unit 41 Prediction Unit 42 Motion Estimation Unit 44 Motion Compensation Unit 46 Intra Prediction Unit 50 Adder 52 Transformation Processing Unit 54 Quantization Unit 56 Entropy Encoding Unit 58 Inverse Quantization Unit 60 Inverse Transformation Unit 62 Adder 64 Reference Picture Memory 80 Entropy Encoding Unit 81 Prediction Unit 82 Motion Compensation Unit 84 Intra Prediction Unit 86 Inverse Quantization Unit 88 Inverse Transformation Unit 90 Adder 92 Reference Picture Memory 102 Residual Generation Module 121 Inter Prediction Module 162 Motion Compensation Module 180 IME Module 182 FME Module 184 Merge Module 186 PU Pattern Determination Module 188 CU Pattern Determination Module 200 Merge Operation 210 AMVP Operation 220 Motion Compensation Operation 250 Encoding Unit 2100 Device 2101 Decision Module 2102 Acquisition Module 2103 Rounding Module 2200 Device 2201 Processor 2202 Memory 2203 Transceiver 2204 Bus 180A IME Module 180B IME Module 180C IME Module 180N IME Module 182A FME Module 182B FME Module 182C FME Module 182N FME Module 184A Merge Module 184B Merge Module 184C Merge Module 184N Merge Module 186A PU Pattern Determination Module 186B PU Pattern Determination Module 186C PU Pattern Determination Module 186N PU Pattern Determination Module 252 Candidate Predicted Motion Vector Position 252A Candidate Predicted Motion Vector Position 252B Candidate Predicted Motion Vector Position 252C Candidate Predicted Motion Vector Position 252D Candidate Predicted Motion Vector Position 252E Candidate Predicted Motion Vector Position

Claims

1. Determining a reference block for a block to be processed, wherein the reference block and the block to be processed have a preset temporal or spatial correlation relationship, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on a predicted motion vector of the reference block, and a predicted block of the reference block is obtained based on the initial motion vector and the one or more preset motion vector offsets; Using the initial motion vector of the reference block as the predicted motion vector of the block to be processed; A method for obtaining a motion vector, comprising the above steps.

2. The initial motion vector of the reference block is obtained in the following manner, that is, Using the predicted motion vector of the reference block as the initial motion vector of the reference block, or Adding the difference between the predicted motion vector of the reference block and the motion vector of the reference block to obtain the initial motion vector of the reference block The method according to claim 1, wherein the initial motion vector is specifically obtained in the above manner.

3. The predicted block of the reference block is obtained in the following manner, that is, Obtaining a picture block indicated by the initial motion vector of the reference block from a reference frame of the reference block, and using the obtained picture block as a temporary predicted block of the reference block; Adding the initial motion vector of the reference block and the one or more preset motion vector offsets to obtain one or more actual motion vectors, wherein each actual motion vector indicates a search position; Obtaining one or more candidate predicted blocks at search positions indicated by the one or more actual motion vectors, wherein each search position corresponds to one candidate predicted block; Selecting, as the predicted block of the reference block, a candidate predicted block with the smallest pixel difference from the temporary predicted block from the one or more candidate predicted blocks The method according to claim 1 or 2, wherein the predicted block is specifically obtained in the above manner.

4. wherein the method is used for bidirectional prediction, the reference frame includes a reference frame in a first direction and a reference frame in a second direction, the initial motion vector includes an initial motion vector in the first direction and an initial motion vector in the second direction, a picture block indicated by the initial motion vector of the reference block is obtained from the reference frame of the reference block, and the obtained picture block is used as a temporary prediction block of the reference block, the step comprising: obtaining a first picture block indicated by the initial motion vector in the first direction of the reference block from the reference frame in the first direction of the reference block; obtaining a second picture block indicated by the initial motion vector in the second direction of the reference block from the reference frame in the second direction of the reference block; weighting the first picture block and the second picture block to obtain the temporary prediction block of the reference block The method according to claim 3, comprising:

5. when the motion vector resolution of the actual motion vector is higher than a preset pixel accuracy, rounding the motion vector resolution of the actual motion vector so that the motion vector resolution of the processed actual motion vector is equal to the preset pixel accuracy The method according to claim 3 or 4, further comprising:

6. the step of selecting, as the prediction block of the reference block, a candidate prediction block having the smallest pixel difference from the temporary prediction block from the one or more candidate prediction blocks, comprising: selecting an actual motion vector corresponding to the candidate prediction block having the smallest pixel difference from the temporary prediction block from the one or more candidate prediction blocks; when the motion vector resolution of the selected actual motion vector is higher than a preset pixel accuracy, rounding the motion vector resolution of the selected actual motion vector so that the motion vector resolution of the processed selected actual motion vector is equal to the preset pixel accuracy; Determining that a predicted block corresponding to a position indicated by the processed selected actual motion vector is the predicted block of the reference block The method according to claim 3 or 4, comprising: **Claim 7** The method according to claim 5 or 6, wherein the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy. **Claim 8** Using the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed The method according to any one of claims 1 to 7, further comprising: **Claim 9** Adding the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed to obtain the initial motion vector of the block to be processed The method according to any one of claims 1 to 7, further comprising: **Claim 10** The method according to claim 9, wherein the method is used for video decoding, and the difference between the motion vectors of the block to be processed is obtained by analyzing first identification information in a bitstream. **Claim 11** The method is used for video decoding, and the step of determining a reference block of a block to be processed comprises analyzing a bitstream to obtain second identification information; and determining the reference block of the block to be processed based on the second identification information. The method according to any one of claims 1 to 9, comprising: **Claim 12** The method is used for video encoding, and the step of determining a reference block of a block to be processed comprises selecting, as the reference block of the block to be processed, a candidate reference block having the minimum rate-distortion cost from one or more candidate reference blocks of the block to be processed. The method according to any one of claims 1 to 9, comprising: **Claim 13** Determine a reference block of a block to be processed, wherein the reference block and the block to be processed have a preset temporal or spatial correlation relationship, the reference block has an initial motion vector and one or more preset motion vector offsets, the initial motion vector of the reference block is obtained based on a predicted motion vector of the reference block, and a predicted block of the reference block is obtained based on the initial motion vector and the one or more preset motion vector offsets, and a determination module configured as such; An acquisition module configured to use the initial motion vector of the reference block as a predicted motion vector of the block to be processed; An apparatus for obtaining a motion vector, comprising the above.

14. The acquisition module is using the predicted motion vector of the reference block as the initial motion vector of the reference block, or adding the predicted motion vector of the reference block and a difference between motion vectors of the reference block to obtain the initial motion vector of the reference block. The apparatus according to claim 13, further configured as such.

15. The acquisition module is obtaining a picture block indicated by the initial motion vector of the reference block from a reference frame of the reference block, using the obtained picture block as a temporary predicted block of the reference block, adding the initial motion vector of the reference block and the one or more preset motion vector offsets to obtain one or more actual motion vectors, each actual motion vector indicating a search position, obtaining one or more candidate predicted blocks at search positions indicated by the one or more actual motion vectors, each search position corresponding to one candidate predicted block, selecting, from the one or more candidate predicted blocks, a candidate predicted block having the smallest pixel difference from the temporary predicted block as the predicted block of the reference block. The apparatus according to claim 13 or 14, further configured as such.

16. The apparatus is configured for bidirectional prediction, the reference frame includes a reference frame in a first direction and a reference frame in a second direction, the initial motion vector includes an initial motion vector in the first direction and an initial motion vector in the second direction, and the acquisition module acquires a first picture block indicated by the initial motion vector in the first direction of the reference block from the reference frame in the first direction of the reference block, acquires a second picture block indicated by the initial motion vector in the second direction of the reference block from the reference frame in the second direction of the reference block, weights the first picture block and the second picture block to obtain the temporary prediction block of the reference block, The apparatus according to claim 15, which is particularly configured as described above.

17. A rounding module configured to round the motion vector resolution of the actual motion vector so that, when the motion vector resolution of the actual motion vector is higher than a preset pixel accuracy, the motion vector resolution of the processed actual motion vector is equal to the preset pixel accuracy The apparatus according to claim 15 or 16, further comprising the same.

18. The acquisition module selects an actual motion vector corresponding to the candidate prediction block with the smallest pixel difference from the one or more candidate prediction blocks to the temporary prediction block, when the motion vector resolution of the selected actual motion vector is higher than a preset pixel accuracy, rounds the motion vector resolution of the selected actual motion vector so that the motion vector resolution of the processed selected actual motion vector is equal to the preset pixel accuracy, determines that the prediction block corresponding to the position indicated by the processed selected actual motion vector is the prediction block of the reference block, The apparatus according to claim 15 or 16, which is particularly configured as described above.

19. The apparatus according to claim 17 or 18, wherein the preset pixel accuracy is integer pixel accuracy, 1 / 2 pixel accuracy, 1 / 4 pixel accuracy, or 1 / 8 pixel accuracy.

20. The acquisition module Use the predicted motion vector of the block to be processed as the initial motion vector of the block to be processed The apparatus according to any one of claims 13 to 19, which is particularly configured to **Claim 21** The acquisition module In order to obtain the initial motion vector of the block to be processed, add the predicted motion vector of the block to be processed and the difference between the motion vectors of the block to be processed The apparatus according to any one of claims 13 to 19, which is particularly configured to **Claim 22** The apparatus according to claim 21, wherein the apparatus is used for video decoding, and the difference between the motion vectors of the block to be processed is obtained by analyzing first identification information in a bitstream **Claim 23** The apparatus is used for video decoding, and the determination module Analyze the bitstream to obtain second identification information Based on the second identification information, determine the reference block of the block to be processed The apparatus according to any one of claims 13 to 21, which is particularly configured to **Claim 24** The apparatus is used for video encoding, and the determination module Select, as the reference block of the block to be processed, a candidate reference block with the minimum rate-distortion cost from one or more candidate reference blocks of the block to be processed The apparatus according to any one of claims 13 to 22, which is particularly configured to **Claim 25** A device for obtaining a motion vector, wherein the device is used for video encoding or video decoding A processor and a memory, wherein the processor and the memory are connected to each other The memory is configured to store program code and video data The processor is configured to read the program code stored in the memory in order to execute the method according to any one of claims 1 to 12, the processor and the memory A device comprising **Claim 26** A computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a computer, the instructions enable the computer to execute the method according to any one of claims 1 to 12, the computer-readable storage medium