Prediction of overlap in video code processing
Through adaptive overlapping block prediction technology, weighted prediction is used to use the motion information of neighboring blocks to perform weighted prediction, which solves the problem of low inter-prediction and intra-prediction efficiency in video encoding, and achieves more efficient video compression and coding complexity reduction, which is suitable for real-time video systems.
Patent Information
- Application Number
- CN202380085447.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-17
- Filing Date
- 2023-12-11
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing video encoding technology processes video stream data, it is difficult to effectively reduce the amount of data and improve the compression efficiency. Especially in inter-frame prediction and intra-frame prediction, especially for complex video content such as screen content and multi-reference frames, there are problems of high encoding complexity and low efficiency.
Adaptive overlapping block prediction technology is adopted to optimize the prediction process of the current block by using the motion information of neighboring blocks, including rounding motion vectors, overlapping prediction using the same reference frame and motion vector difference is less than the threshold, reducing encoding complexity and improving compression efficiency.
Improves the compression efficiency and encoding complexity of video encoding, especially in the case of screen content and multi-reference frames, reduces encoding complexity and improves video quality, and is suitable for real-time video systems.
Smart Images

Figure CN120303939A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This application claims the benefit and priority of U.S. Provisional Patent Application Serial No. 63 / 434,991, filed on December 23, 2022, and U.S. Provisional Patent Application Serial No. 63 / 480,262, filed on January 17, 2023. The entire disclosures of these provisional patent applications are incorporated herein by reference. Background Art
[0002] Digital video can be used for, e.g., remote business meetings via video conferencing, high - definition video entertainment, video advertising, or sharing of user - generated videos. Due to the large amount of data involved in video data, high - performance compression is required for transmission and storage. Various methods have been proposed to reduce the amount of data in a video stream, including compression and other encoding and decoding techniques. Summary of the Invention
[0003] This application relates to encoding and decoding video stream data for transmission or storage. Aspects of systems, methods, and apparatus related to adaptive overlapping block prediction in variable - block - size video coding are disclosed herein.
[0004] A system of one or more computers can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which in operation cause the system to perform those actions. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by data - processing apparatus, cause the apparatus to perform the actions.
[0005] In one general aspect, a method for coding a current block of a current frame includes: obtaining a first prediction block based on motion information associated with a first prediction block for the current block; obtaining a second prediction block for at least a portion of the current block based on motion information associated with neighboring blocks; obtaining a prediction difference metric between the first prediction block and the second prediction block; and determining whether to combine the first prediction block and the second prediction block of the portion of the current block based on the prediction difference metric. Implementations can further include one or more of the following features.
[0006] In this method, determining whether to combine the first prediction block and the second prediction block for this portion of the current block based on a prediction error metric may include determining not to combine the first prediction block and the second prediction block in response to the prediction error metric exceeding a threshold. In this method, the threshold may be a power of 2. In this method, the prediction error metric may be the sum of absolute differences (SAD) between the first prediction block and the second prediction block. In this method, the prediction error metric may be the sum of squared errors (SSE) between the first prediction block and the second prediction block. In this method, the prediction error metric may be the absolute maximum value of the pairwise differences between the first prediction block and the second prediction block.
[0007] In this method, the prediction error metric may be calculated based on at least one of the maximum absolute difference or the average absolute difference. In this method, the maximum absolute difference may be the absolute maximum value of the pairwise differences between the first prediction block and the second prediction block. In this method, the average absolute difference is the average of the pairwise differences. In this method, the prediction error metric may be calculated as the absolute difference between the maximum absolute difference and the average absolute difference. In this method, the prediction error metric may be calculated based on the ratio of the maximum absolute difference to the average absolute difference.
[0008] This method may include determining to obtain a second prediction block in response to determining that the motion vector difference between the first motion vector of the current block and the second motion vector of a neighboring block is less than a motion vector threshold. This method may include determining to obtain a second prediction block in response to determining that at least a portion of the multiple first reference frames used to predict the current block is the same as the multiple second reference frames used to predict the neighboring block. This method may include determining to obtain a second prediction block in response to determining that the current block is a block of a P frame or P slice, the current block and the neighboring block use the same reference frame, and the corresponding absolute motion vector difference between the motion vector of the current block and the motion vector of the neighboring block is below the motion vector threshold.
[0009] In another general aspect, a method for coding a current block of a current frame includes: obtaining a first prediction block for the current block based on a first reference frame and a first motion vector; determining, at least in part based on information related to a neighboring block of the current block, to obtain at least a portion of a second prediction block for the current block using an overlapped prediction mode that uses a second reference frame and a second motion vector of the neighboring block, where the second motion vector of the neighboring block is obtained by rounding the motion vector used to predict the neighboring block to an integer pixel position; obtaining the second prediction block using the overlapped prediction mode; and combining the first prediction block and the second prediction block. Implementations may further include one or more of the following features.
[0010] In this method, determining that a second predicted block for a portion of a current block is to be obtained using an overlapping prediction mode based at least in part on information related to neighboring blocks of the current block may include determining that a motion vector difference between a first motion vector and a second motion vector is less than a motion vector threshold. The method may include decoding a one-frame-distance-motion-vector threshold from a compressed bitstream; and calculating the motion vector threshold based on a frame difference between a first reference frame and a second reference frame and the one-frame-distance-motion-vector threshold. In this method, the motion vector threshold may be proportional to a temporal distance between the first reference frame and the second reference frame.
[0011] The method may include, in response to determining that the first reference frame is different from the second reference frame, calculating a difference between the first motion vector and the second motion vector by the following steps. First, in response to determining that a plurality of first motion vectors including the first motion vector are used to predict the current block, obtaining a scaled first motion vector by scaling the first motion vector based on a temporal distance to point to a target reference frame; and averaging the scaled first motion vectors to obtain a first normalized motion vector. Second, in response to determining that a plurality of second motion vectors including the second motion vector are used to predict a neighboring block, obtaining a scaled second motion vector by scaling the second motion vector to point to the target reference frame; and averaging the scaled second motion vectors to obtain a second normalized motion vector. Third, calculating the motion vector difference based on: 1) one of the first normalized motion vector or the first motion vector and 2) one of the second normalized motion vector or the second motion vector.
[0012] In this method, determining that a second predicted block for a portion of a current block is to be obtained using an overlapping prediction mode based at least in part on information related to neighboring blocks of the current block may include determining that a plurality of first reference frames for predicting the current block and a plurality of second reference frames for predicting neighboring blocks are at least partially the same, where the plurality of first reference frames includes the first reference frame and the plurality of second reference frames includes the second reference frame, and where the first reference frame is the same as the second reference frame.
[0013] In this method, determining that a second predicted block for a portion of a current block is to be obtained using an overlapping prediction mode based at least in part on information related to neighboring blocks of the current block may include determining that reference samples for obtaining the second predicted block using a subpixel interpolation filter are available.
[0014] In this method, determining a second prediction block for a portion of a current block to be obtained using an overlapping prediction mode based at least in part on information related to neighboring blocks of the current block may include determining that the current block is a block of a P-frame or P-slice, that the current block and the neighboring blocks use the same reference frame, and that a corresponding absolute motion vector difference between a motion vector of the current block and motion vectors of the neighboring blocks is below a motion vector threshold.
[0015] In this method, at least one of rounding toward positive infinity, rounding toward negative infinity, or rounding toward zero is used to round a motion vector that may be used to predict a neighboring block. In this method, in a case where a component of the motion vector is positive, the component may be rounded toward positive infinity, and in a case where the component is negative, the component may be rounded toward negative infinity.
[0016] In a case where a current block may be along a first boundary of a parent block, the method may include determining whether to perform an overlapping prediction mode with respect to a second boundary of the parent block, different from the first boundary, in parallel with determining whether to obtain a second prediction block for at least a portion of the current block using an overlapping prediction mode.
[0017] Variations of these and other aspects will be described in more detail below. It should be understood that the aspects can be implemented in any convenient form. For example, the aspects can be implemented by a suitable computer program that can be carried on a suitable carrier medium, which can be a tangible carrier medium (e.g., a disk) or an intangible carrier medium (e.g., a communication signal). The aspects can also be implemented using suitable apparatus, which can take the form of a programmable computer running a computer program arranged to implement the methods and / or techniques disclosed herein. The aspects can be combined such that features described in the context of one aspect can be implemented in another aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The description herein refers to the accompanying drawings, where like reference numerals refer to like parts throughout several views, and where:
[0019] Figure 1 is a diagram of a computing device according to an implementation of the present disclosure.
[0020] Figure 2 is a diagram of a computing and communication system according to an implementation of the present disclosure.
[0021] Figure 3 is a diagram of a video stream for use in encoding and decoding according to an implementation of the present disclosure.
[0022] Figure 4 is a block diagram of an encoder according to an implementation of the present disclosure.
[0023] Figure 5 is a block diagram of a decoder according to an implementation of the present disclosure.
[0024] Figure 6 is a flowchart of an example process for adaptive overlapping block prediction according to an implementation of the present disclosure.
[0025] Figure 7 is a block diagram of an example block-based prediction at variable block sizes according to an implementation of the present disclosure.
[0026] Figure 8 is a block diagram of an example size change of an overlapping region according to an implementation of the present disclosure.
[0027] Figure 9 is a block diagram of an example weighting function for overlapping prediction according to an implementation of the present disclosure.
[0028] Figure 10 is a block diagram showing prediction of sub-block overlap.
[0029] Figure 11 is a flowchart of a technique for coding a current block of a video frame using overlapping prediction.
[0030] Figure 12 is a flowchart of another technique for coding a current block of a video frame using overlapping prediction. DETAILED DESCRIPTION
[0031] Video compression schemes can include decomposing each image or frame into smaller parts, such as blocks, and using techniques that limit the information included in the output for each block to generate an output bitstream. The encoded bitstream can be decoded to recreate the source image based on the limited information. In some implementations, the information included in the output for each block can be limited by reducing spatial redundancy, reducing temporal redundancy, or a combination thereof. For example, temporal redundancy or spatial redundancy can be reduced by predicting a frame based on information available to both the encoder and the decoder and including in the encoded video stream information representing the difference or residual between the predicted frame and the original frame.
[0032] In some implementations, a frame can be segmented into variable-sized blocks, the pixel values of each block can be predicted using previously coded information, and the prediction parameters and residual data for each block can be encoded as output. The decoder can receive the prediction parameters and residual data in the compressed bitstream and can reconstruct the frame, which can include predicting the block based on previously decoded image data.
[0033] Overlapped prediction is a weighted prediction that can improve the prediction of a block by using prediction information from adjacent blocks. The prediction from the current block and the prediction based on the motion information from neighboring blocks can be weighted to form a final prediction. When applying overlapped prediction, the motion information of one or more neighboring blocks can be used to refine at least some pixels of the current block (e.g., the top pixels and / or the left pixels near the boundary).
[0034] In some implementations, the prediction block size of adjacent blocks can vary between adjacent blocks and can be different from the prediction block size of the current block. The corresponding overlapping region within the current block corresponding to the respective adjacent block can be identified, and the overlapped prediction can be determined for the corresponding overlapping region based on the prediction parameters from the corresponding adjacent block. In some implementations, the overlapped prediction can be optimized by adjusting the size of each overlapping region in the current block—such as by comparing the prediction parameters of the adjacent block and the prediction parameters of the current block.
[0035] Figure 1 FIG. is a diagram of a computing device 100 according to an implementation of the present disclosure. The computing device 100 may include a communication interface 110, a communication unit 120, a user interface (UI) 130, a processor 140, a memory 150, instructions 160, a power supply 170, or any combination thereof. As used herein, the term "computing device" includes any unit or combination of units capable of performing any technology or any one or more parts thereof disclosed herein.
[0036] The computing device 100 can be a fixed computing device, such as a personal computer (PC), a server, a workstation, a minicomputer, or a mainframe computer; or a mobile computing device, such as a mobile phone, a personal digital assistant (PDA), a laptop computer, or a tablet PC. Although shown as a single unit, any one or more elements of the computing device 100 can be integrated into any number of separate physical units. For example, the UI 130 and the processor 140 can be integrated in a first physical unit, and the memory 150 can be integrated in a second physical unit.
[0037] The communication interface 110 can be a wireless antenna as shown, a wired communication port such as an Ethernet port, an infrared port, a serial port, or any other wired or wireless unit capable of docking with a wired or wireless electronic communication medium 180.
[0038] The communication unit 120 can be configured to send or receive signals via a wired or wireless electronic communication medium 180. For example, as shown, the communication unit 120 is operatively connected to an antenna configured to communicate via wireless signals. Although in Figure 1is not explicitly shown, but communication unit 120 may be configured to transmit, receive, or both via any wired or wireless communication medium, such as radio frequency (RF), ultraviolet (UV), visible light, fiber optic, wired line, or a combination thereof. Although Figure 1 a single communication unit 120 and a single communication interface 110 are shown, any number of communication units and any number of communication interfaces may be used.
[0039] UI 130 may include any unit capable of interfacing with a user, such as a virtual or physical keyboard, touchpad, display, touch display, speaker, microphone, camera, sensor, or any combination thereof. UI 130 may be operatively coupled to the processor, as shown, or to any other element of computing device 100, such as power supply 170. Although shown as a single unit, UI 130 may include one or more physical units. For example, UI 130 may include an audio interface for performing audio communication with the user, and a touch display for performing visual and touch-based communication with the user. Although shown as separate units, communication interface 110, communication unit 120, and UI 130 or portions thereof may be configured as a combined unit. For example, communication interface 110, communication unit 120, and UI 130 may be implemented as a communication port capable of interfacing with an external touchscreen device.
[0040] Processor 140 may include any device or system capable of manipulating or processing existing or later-developed signals or other information, including optical processors, quantum processors, molecular processors, or a combination thereof. For example, processor 140 may include a dedicated processor, digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic array, a programmable logic controller, microcode, firmware, any type of integrated circuit (IC), a state machine, or any combination thereof. As used herein, the term "processor" includes a single processor or multiple processors. The processor may be operatively coupled to communication interface 110, communication unit 120, UI 130, memory 150, instructions 160, power supply 170, or any combination thereof.
[0041] The memory 150 may include any non-transitory computer-usable or computer-readable medium, such as any tangible device that can, for example, contain, store, communicate, or transport instructions 160 or any information associated therewith for use by or in conjunction with the processor 140. The non-transitory computer-usable or computer-readable medium may be, for example, a solid-state drive, a memory card, a removable medium, a read-only memory (ROM), a random access memory (RAM), any type of disk (including a hard disk, a floppy disk, an optical disk, a magnetic card, or an optical card), an application specific integrated circuit (ASIC), or any type of non-transitory medium suitable for storing electronic information, or any combination thereof. The memory 150 may be connected to the processor 140, for example, via a memory bus (not explicitly shown).
[0042] The instructions 160 may include guidance for performing any of the techniques disclosed herein, or any one or more portions thereof. The instructions 160 may be implemented in hardware, software, or any combination thereof. For example, the instructions 160 may be implemented as information stored in the memory 150, such as a computer program, that may be executed by the processor 140 to perform any one of the corresponding techniques, algorithms, aspects, or combinations thereof described herein. One or more portions of the instructions 160 may be implemented as a dedicated processor or circuitry that may include dedicated hardware for practicing any one of the techniques, algorithms, aspects, or combinations thereof described herein. Portions of the instructions 160 may be distributed across multiple processors on the same machine or different machines, or across a network such as a local area network, a wide area network, the Internet, or any combination thereof.
[0043] The power supply 170 may be any suitable device for powering the communication interface 110. For example, the power supply 170 may include a wired power supply; one or more dry batteries, such as nickel-cadmium (NiCd) batteries, nickel-zinc (NiZn) batteries, nickel-metal hydride (NiMH) batteries, lithium-ion (Li-ion) batteries; solar cells; fuel cells; or any other device capable of powering the communication interface 110. The communication interface 110, the communication unit 120, the UI 130, the processor 140, the instructions 160, the memory 150, or any combination thereof may be operatively coupled to the power supply 170.
[0044] Although shown as separate elements, the communication interface 110, the communication unit 120, the UI 130, the processor 140, the instructions 160, the power supply 170, the memory 150, or any combination thereof may be integrated in one or more electronic units, circuits, or chips.
[0045] Figure 2FIG. is a diagram of a computing and communication system 200 according to an implementation of the present disclosure. The computing and communication system 200 may include one or more computing and communication devices 100A, 100B, 100C, one or more access points 210A, 210B, one or more networks 220, or a combination thereof. For example, the computing and communication system 200 may be a multi-access system that provides communications (such as voice, data, video, messaging, broadcasting, or a combination thereof) to one or more wired or wireless communication devices (such as computing and communication devices 100A, 100B, 100C). Although for simplicity, Figure 2 Three computing and communication devices 100A, 100B, 100C, two access points 210A / 210B, and one network 220 are shown, but any number of computing and communication devices, access points, and networks may be used.
[0046] The computing and communication devices 100A, 100B, 100C may be, for example, computing devices such as Figure 1 the computing device 100 shown in. For example, as shown, the computing and communication devices 100A, 100B may be user devices such as mobile computing devices, laptop computers, thin clients, or smartphones, and the computing and communication device 100C may be a server such as a mainframe or a cluster. Although the computing and communication devices 100A, 100B are described as user devices and the computing and communication device 100C is described as a server, any computing and communication device may perform some or all of the functions of a server, some or all of the functions of a user device, or some or all of the functions of a server and a user device.
[0047] Each computing and communication device 100A, 100B, 100C may be configured to perform wired or wireless communication. For example, the computing and communication devices 100A, 100B, 100C may be configured to send or receive wired or wireless communication signals and may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a cellular phone, a personal computer, a tablet computer, a server, a consumer electronic product, or any similar device. Although each computing and communication device 100A, 100B, 100C is shown as a single unit, the computing and communication device may include any number of interconnected elements.
[0048] Each access point 210A, 210B can be any type of device configured to communicate with computing and communication devices 100A, 100B, 100C, network 220, or both via wired or wireless communication links 180A, 180B, 180C. For example, access points 210A, 210B can include base stations, base transceiver stations (BTSs), Node Bs, enhanced Node Bs (eNode-Bs), home Node Bs (HNode-Bs), wireless routers, wired routers, hubs, repeaters, switches, or any similar wired or wireless device. Although each access point 210A, 210B is shown as a single unit, an access point can include any number of interconnected elements.
[0049] Network 220 can be any type of network configured to provide services such as voice, data, applications, Voice over Internet Protocol (VoIP), or any other communication protocol or combination of communication protocols via wired or wireless communication links. For example, network 220 can be a local area network (LAN), wide area network (WAN), virtual private network (VPN), mobile or cellular phone network, the Internet, or any other means of electronic communication. The network can use communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Internet Protocol (IP), Real-Time Transport Protocol (RTP), Hypertext Transfer Protocol (HTTP), or any combination thereof.
[0050] Computing and communication devices 100A, 100B, 100C can communicate with each other via network 220 using one or more wired or wireless communication links, or via a combination of wired and wireless communication links. For example, as shown, computing and communication devices 100A, 100B can communicate via wireless communication links 180A, 180B, and computing and communication device 100C can communicate via wired communication link 180C. Any one of computing and communication devices 100A, 100B, 100C can communicate using any one or more wired or wireless communication links. For example, first computing and communication device 100A can communicate via a first type of communication link via first access point 210A, second computing and communication device 100B can communicate via a second type of communication link via second access point 210B, and third computing and communication device 100C can communicate via a third type of communication link via a third access point (not shown). Similarly, access points 210A, 210B can communicate with network 220 via one or more types of wired or wireless communication links 230A, 230B. Although Figure 2It is shown that the computing and communication devices 100A / 100B / 100C communicate via the network 220. However, the computing and communication devices 100A, 100B, 100C can communicate with each other via any number of communication links (such as direct wired or wireless communication links).
[0051] Other implementations of the computing and communication system 200 are also possible. For example, the network 220 can be an ad-hoc network, and one or more of the access points 210A, 210B can be omitted. The computing and communication system 200 can include Figure 2 devices, units, or elements not shown in []. For example, the computing and communication system 200 can include more communication devices, networks, and access points.
[0052] Figure 3 is a diagram of a video stream 300 for use in encoding and decoding according to an implementation of the present disclosure. The video stream 300 (such as a video stream captured by a camera or a video stream generated by a computing device) can include a video sequence 310. The video sequence 310 can include a sequence of adjacent frames 320. Although three adjacent frames 320 are shown, the video sequence 310 can include any number of adjacent frames 320. Each frame 330 from the adjacent frames 320 can represent a single image from the video stream. The frame 330 can include blocks 340. Although Figure 3 not shown in [], the blocks can include pixels. For example, a block can include a group of 16x16 pixels, a group of 8x8 pixels, a group of 8x16 pixels, or any other group of pixels. Unless otherwise indicated herein, the term 'block' can include superblocks, macroblocks, segments, slices, or any other part of a frame. The frame, block, pixel, or a combination thereof can include display information, such as luminance information, chrominance information, or any other information that can be used to store, modify, communicate, or display the video stream or a part thereof.
[0053] Figure 4 is a block diagram of an encoder 400 according to an implementation of the present disclosure. The encoder 400 can be implemented in a device such as the computing device 100 shown in Figure 1 or the computing and communication device 100A / 100B / 100C shown in Figure 2 as, for example, a computer software program stored in a data storage unit such as the memory 150 shown in Figure 1 . The computer software program can include machine instructions executable by a processor such as the processor 140 shown in Figure 1 , and can cause the device to encode video data as described herein. The encoder 400 can be implemented as, for example, dedicated hardware included in the computing device 100.
[0054] Encoder 400 can encode an input video stream (i.e., video stream 402) to generate an encoded (compressed) bitstream 404, and the input video stream can be Figure 3 the video stream 300 shown in. In some implementations, encoder 400 may include a forward path for generating the compressed bitstream 404. The forward path may include an intra / inter prediction unit 410, a transform unit 420, a quantization unit 430, an entropy coding unit 440, or any combination thereof. In some implementations, encoder 400 may include a reconstruction path (indicated by the dashed connection line) for reconstructing frames for encoding of further blocks. The reconstruction path may include an inverse quantization unit 450, an inverse transform unit 460, a reconstruction unit 470, a loop filter unit 480, or any combination thereof. Other structural variations of encoder 400 may be used to encode video stream 402.
[0055] To encode video stream 402, each frame in video stream 402 can be processed on a block-by-block basis. Thus, a current block can be identified from the blocks in a frame, and the current block can be encoded.
[0056] In the intra / inter prediction unit 410, the current block can be encoded using intra prediction (which can be within a single frame) or inter prediction (which can be from frame to frame). Intra prediction may include generating a predicted block based on samples that have been previously encoded and reconstructed in the current frame. Inter prediction may include generating a predicted block based on samples in one or more previously constructed reference frames. Generating a predicted block for the current block in the current frame may include performing motion estimation to generate a motion vector indicating an appropriate reference block in the reference frame.
[0057] The intra / inter prediction unit 410 can subtract the predicted block from the current block (the original block) to produce a residual block. The transform unit 420 can perform a block-based transform, which may include transforming the residual block into transform coefficients in, for example, the frequency domain. Examples of block-based transforms include the Karhunen-Loève transform (KLT), the discrete cosine transform (DCT), and the singular value decomposition transform (SVD). In an example, the DCT may include transforming a block into the frequency domain. The DCT may include using transform coefficient values based on spatial frequencies, with the lowest frequency (DC) coefficient at the upper left of the matrix and the highest frequency coefficient at the lower right of the matrix.
[0058] The quantization unit 430 can convert the transform coefficients into discrete quantum values, which can be referred to as quantized transform coefficients or quantization levels. The quantized transform coefficients can be entropy encoded by the entropy encoding unit 440 to produce entropy encoded coefficients. Entropy encoding can include using probability distribution metrics. The entropy encoded coefficients and the information for decoding the block can be output to the compressed bitstream 404, and this information can include the prediction type used, motion vectors, and quantizer values. The compressed bitstream 404 can be formatted using various techniques such as run-length encoding (RLE) and zero-run coding.
[0059] The reconstruction path can be used to maintain reference frame synchronization between the encoder 400 and the corresponding decoder such as Figure 5 the decoder 500 shown in. The reconstruction path can be similar to the decoding process discussed below and can include dequantizing the quantized transform coefficients at the dequantization unit 450 and performing an inverse transform on the dequantized transform coefficients at the inverse transform unit 460 to produce a derived residual block. The reconstruction unit 470 can add the prediction block generated by the intra / inter prediction unit 410 to the derived residual block to create a reconstructed block. The loop filter unit 480 can be applied to the reconstructed block to reduce distortions such as block artifacts.
[0060] Other variants of the encoder 400 can be used to encode the compressed bitstream 404. For example, a transform-free encoder 400 can directly quantize the residual block without the transform unit 420. In some implementations, the quantization unit 430 and the dequantization unit 450 can be combined into a single unit.
[0061] Figure 5 is a block diagram of a decoder 500 according to an implementation of the present disclosure. The decoder 500 can be implemented in a device such as Figure 1 the computing device 100 shown in or Figure 2 the computing and communication devices 100A / 100B / 100C shown in as, for example, a computer software program stored in a data storage unit such as Figure 1 the memory 150 shown in. The computer software program can include machine instructions executable by a processor such as Figure 1 the processor 140 shown in and can cause the device to decode video data as described herein. The decoder 500 can be implemented as, for example, dedicated hardware included in the computing device 100.
[0062] The decoder 500 can receive the compressed bitstream 502, such as Figure 4The compressed bitstream 404 shown in [reference], and the compressed bitstream 502 can be decoded to generate an output video stream 504. The decoder 500 may include an entropy decoding unit 510, a dequantization unit 520, an inverse transformation unit 530, an intra / inter prediction unit 540, a reconstruction unit 550, a loop filter unit 560, a deblocking filter unit 570, or any combination thereof. Other structural variations of the decoder 500 may be used to decode the compressed bitstream 502.
[0063] The entropy decoding unit 510 may decode data elements within the compressed bitstream 502 using, for example, context adaptive binary arithmetic decoding to produce a set of quantized transform coefficients. The dequantization unit 520 may dequantize the quantized transform coefficients, and the inverse transformation unit 530 may inverse-transform the dequantized transform coefficients to produce a derived residual block, which may correspond to the derived residual block generated by the inverse transformation unit 460 shown in [reference]. Using the header information decoded from the compressed bitstream 502, the intra / inter prediction unit 540 may generate a prediction block corresponding to the prediction block created in the encoder 400. At the reconstruction unit 550, the prediction block may be added to the derived residual block to create a reconstructed block. The loop filter unit 560 may be applied to the reconstructed block to reduce block artifacts. The deblocking filter unit 570 may be applied to the reconstructed block to reduce block distortion, and the result may be output as the output video stream 504. Figure 4 Other variations of the decoder 500 may be used to decode the compressed bitstream 502. For example, the decoder 500 may produce the output video stream 504 without the deblocking filter unit 570.
[0064] In some implementations, reducing temporal redundancy may include encoding frames using a relatively small amount of data based on the similarity between frames using one or more reference frames, which may be previously encoded, decoded, and reconstructed frames of the video stream. For example, blocks or pixels of the current frame may be similar to the spatially corresponding blocks or pixels of the reference frame. In some implementations, blocks or pixels of the current frame may be similar to blocks or pixels of the reference frame in different parts, and reducing temporal redundancy may include generating motion information indicating the spatial difference or translation between the position of a block or pixel in the current frame and the corresponding position of a block or pixel in the reference frame.
[0065] In some implementations, reducing temporal redundancy may include encoding frames using a relatively small amount of data based on the similarity between frames using one or more reference frames, which may be previously encoded, decoded, and reconstructed frames of the video stream. For example, blocks or pixels of the current frame may be similar to the spatially corresponding blocks or pixels of the reference frame. In some implementations, blocks or pixels of the current frame may be similar to blocks or pixels of the reference frame in different parts, and reducing temporal redundancy may include generating motion information indicating the spatial difference or translation between the position of a block or pixel in the current frame and the corresponding position of a block or pixel in the reference frame.
[0066] In some implementations, reducing temporal redundancy may include identifying a block or pixel in a reference frame or a portion of the reference frame that corresponds to a current block or pixel of a current frame. For example, a reference frame or a portion of the reference frame that may be stored in a memory may be searched to obtain an optimal block or pixel to be used for encoding a current block or pixel of the current frame. For example, the search may identify a block in the reference frame for which the difference in pixel values between its reference block and the current block is minimized, and may be referred to as motion search. In some implementations, the portion of the reference frame that is searched may be restricted. For example, the portion of the reference frame that is searched may include a restricted number of rows of the reference frame, and this portion may be referred to as a search area. In an example, identifying the reference block may include calculating a cost function between the pixels of the blocks in the search area and the pixels of the current block, such as the sum of absolute differences (SAD). In some implementations, more than one reference frame may be provided. For example, three reference frames may be selected from eight candidate reference frames.
[0067] In some implementations, the spatial difference between the position of a reference block in a reference frame and the position of a current block in a current frame may be represented as a motion vector. The difference in pixel values between the reference block and the current block may be referred to as differential data, residual data, or a residual block. In some implementations, generating the motion vector may be referred to as motion estimation, and the pixels of the current block may be indicated as f based on their positions using Cartesian coordinates x,y . Similarly, the pixels of the search area of the reference frame may be indicated as r based on their positions using Cartesian coordinates x,y . The motion vector (MV) for the current block may be determined based on, for example, the SAD between the pixels of the current frame and the corresponding pixels of the reference frame.
[0068] In some implementations, for inter-frame prediction, the encoder 400 may transmit, at the block endpoints, the encoded information for the predicted block, including but not limited to the prediction mode, the prediction reference frame, the motion vector (if required), and the sub-pixel interpolation filter type.
[0069] Figure 6 is a flowchart of an example of a technique 600 for adaptive overlapping block prediction according to an implementation of the present disclosure. Adaptive overlapping block prediction may be implemented in an encoder, such as the prediction performed by the intra / inter prediction unit 410 of the encoder 400 shown in Figure 4 ; or may be implemented in a decoder, such as by Figure 5The prediction based on the compressed bitstream 502 performed by the intra / inter prediction unit 540 in the decoder 500 shown. In some implementations, the adaptive overlapping block prediction may include: determining a base prediction for a current block based on prediction parameters for the current block at 610; identifying adjacent prediction parameters for adjacent blocks at 620; determining an overlapping region in the current block that is adjacent to an adjacent block at 630; determining an overlapping prediction for the overlapping region that is a weighted function of the base prediction and a prediction based on the adjacent prediction parameters at 640; generating an overlapping prediction block based on combining the overlapping predictions at 650; or a combination thereof.
[0070] At 610, the base prediction for the current block may be performed using the current prediction parameters for the current block. For example, the prediction parameters for inter prediction may include a reference frame and a motion vector for the current block. The base prediction block may be determined using the base prediction for the current block.
[0071] At 620, the adjacent prediction parameters may be identified. In some implementations, identifying the adjacent prediction parameters may include identifying a previously encoded or decoded adjacent block, and for each previously encoded or decoded adjacent block among the previously encoded or decoded adjacent blocks, identifying the prediction parameters used to encode or decode the adjacent block.
[0072] At 630, the overlapping region may be determined. In some implementations, the overlapping region in the current block may be determined for one or more of the encoded or decoded adjacent blocks identified at 620. The overlapping region may include a region within the current block that is adjacent to a corresponding adjacent block, such as a pixel grouping. The overlapping region determination may be conditional on the presence of at least one previously encoded or decoded adjacent block that is smaller in size than the current block.
[0073] At 640, the overlapping prediction may be determined. In some implementations, the overlapping prediction for the overlapping region identified at 630 may be determined based on a weighted function of the base prediction determined at 610 and a prediction generated using the adjacent prediction parameters from the corresponding adjacent block to predict the pixel values in the current block within the overlapping region. For example, for the overlapping region, a prediction block having a size equal to the size of the overlapping region may be determined using the prediction parameters of the corresponding adjacent block. The overlapping prediction may be performed for the overlapping region based on a weighted combination of the base prediction block pixel values and the prediction block pixel values generated for the overlapping region based on the prediction parameters of the corresponding adjacent block. For example, the pixel value of a pixel in the overlapping region may be a weighted average of the pixel value from the base prediction block and the corresponding pixel value from the prediction block generated for the overlapping region based on the prediction parameters of the corresponding adjacent block. In some implementations, generating the prediction block for the corresponding overlapping region may be omitted, and the overlapping prediction block may be generated on a pixel-by-pixel basis.
[0074] At 650, overlapping predictions from one or more adjacent blocks can be used to generate an overlapping prediction block. For example, the overlapping prediction at 640 can be repeated for one or more overlapping regions within the current block to form an overlapping prediction block.
[0075] In some implementations, portions of the current block that do not spatially correspond to the overlapping regions used for the current block can be predicted based on a base prediction.
[0076] In some implementations, the overlapping prediction block used for the current block can be compared with a base prediction block, and either the base prediction block or the overlapping prediction block can be used as the prediction block for the current block. For example, the comparison can be based on a residual-based error metric, and the encoder 400 can select the prediction block that yields a lower error value.
[0077] In some implementations, information indicating that overlapping prediction has been performed on the current block can be included in the encoded bitstream. For example, an indication of the type of weighting function used for overlapping prediction can be indicated in the encoded bitstream. In some implementations, the indication of the weighting function can be omitted from the encoded bitstream, and decoding the encoded bitstream can include using context information from a previously decoded adjacent frame to determine the weighting function. For example, decoding can include identifying the weighting function based on which the adjacent block prediction parameters yield the minimum residual-based error.
[0078] Figure 7 is a block diagram of an example block-based prediction at variable block sizes according to an implementation of the present disclosure. In some implementations, at least one side of the current block can be adjacent to two or more previously encoded or decoded blocks. As shown, the current block 720 for prediction is surrounded by previously encoded or decoded top adjacent blocks 721, 722, 723 and left adjacent block 724. Although the previously encoded or decoded adjacent blocks are Figure 7 shown above and to the left of the current block 720, in some implementations, the previously encoded or decoded adjacent blocks can be below or to the right of the current block, or some combination of top, left, bottom, or right.
[0079] As Figure 7 shown, the current block 720 is a block, the adjacent block 721 is a block, the adjacent blocks 722, 723 are blocks, and the adjacent block 724 is a block. Although Figure 7 shows , and blocks, any other block sizes can be used according to the present disclosure.
[0080] In some implementations, for overlapping prediction of the current block 720, an overlapping region can be determined with respect to one or more previously encoded or decoded neighboring blocks. For example, pixels in the current block 720 can be grouped within the defined overlapping regions, where the overlapping regions can be determined for one or more top neighboring blocks, such as overlapping regions 731, 732, and 733 corresponding to neighboring blocks 721, 722, and 723 respectively, and an overlapping region 734 shown at the left half of the current block 720 corresponding to the left neighboring block 724. As Figure 7 shown, the overlapping regions can overlap, such as overlapping regions 731 and 734, where overlapping region 731 includes the intersection of the overlapping regions corresponding to the top neighboring block 721 and the left neighboring block 724. As shown, the overlapping regions 731–734 are within the current block 720 and are adjacent to the corresponding neighboring blocks 721–724.
[0081] In some implementations, a weighting function for overlapping prediction can determine the overlapping region size. The size of the overlapping region can correspond to the size of the corresponding neighboring block, such as the corresponding column size v, row size w, or both. In some implementations, v w the overlapping region size can correspond to the neighboring block size, where and . For example, an overlapping region as shown can be determined with respect to Figure 7 neighboring block 721, such as overlapping region 731 within the current block 720.
[0082] The size of the overlapping region can correspond to the size of the current block, such as the corresponding column size v, row size w, or both. In some implementations, v w the overlapping region size can correspond to the current block size, where and ’. As an example, the current block can be smaller than the neighboring block, and the overlapping region size for one dimension can be limited to the size of the current block at the boundary of the neighboring block. As another example, as Figure 7 shown the overlapping regions 732 and 733 can be determined with respect to neighboring blocks 722 and 733 respectively, and the number of rows corresponds to half of the 16 16 current block size dimension, where w = 1 / 2 y’ = 8. In some implementations, v w the overlapping region size can correspond to the current block size, where and 。For example, it can be determined with respect to the left adjacent block 724 the 16 overlapping regions 734. In some implementations, the size of the overlapping regions can correspond to both the adjacent block size and the current block size. Other variations of the overlapping region size can be used.
[0083] In some implementations, a weighted function index indicating which of the various discrete overlapping region sizes is used as a common size for all overlapping regions can be included in the encoded bitstream, and decoding the block can include decoding the index to determine which of the discrete overlapping region sizes is to be used for the prediction weighted function of the overlap. As an example, a first index can indicate that all overlapping regions have a size with a first dimension equal to the length of the adjacent block edge and a second dimension extending ½ the length of the current block, such as Figure 7 the overlapping region 732 shown in. A second index can indicate that all overlapping regions have a first dimension equal to the adjacent block edge and a second dimension extending ¼ the length of the current block, such as Figure 8 the overlapping region 904 shown in. In some implementations, encoding can include determining a weighted function that maps different relative sizes for each of the overlapping regions in the overlap depending on the prediction parameters used for the adjacent block. For example, encoding can include generating multiple predicted block candidates according to various weighted functions, determining rate-distortion cost estimates for each candidate, and selecting the weighted function that provides the best rate-distortion optimization.
[0084] Figure 8 is a block diagram of an example size variation of an overlapping region according to an implementation of the present disclosure. In some implementations, the size of the overlapping region can exceed the corresponding size of the corresponding adjacent block. For example, as Figure 8 shown, the overlapping region 902 corresponding to the adjacent block 722 can be determined to have a horizontal size greater than the number of horizontal pixels in the corresponding adjacent block 722 and equal in vertical size. As another example, the overlapping region 904 corresponding to the adjacent block 722 can be determined according to the horizontal size of the adjacent block and ¼ the vertical size of the current block. In some implementations, both the horizontal dimension and the vertical dimension of the overlapping region can exceed the corresponding size of the corresponding adjacent block. In some implementations, both the horizontal dimension and the vertical dimension of the overlapping region can be exceeded by the corresponding size of the corresponding adjacent block.
[0085] Multiple overlapping region sizes can be determined by using a discrete set of size determination functions, and an overlapping region size can be adaptively selected from the multiple overlapping region sizes as a function of the difference between a prediction parameter of a current block and an adjacent prediction parameter of a corresponding adjacent block. In some implementations, a comparison between a motion vector of a current block for an overlapping region and a motion vector of a corresponding adjacent block may indicate a motion vector difference exceeding a threshold, and one or more dimensions of a default overlapping region size may be adjusted. In some implementations, determination of a difference between prediction parameters between an adjacent block and a current block may be based on a comparison of temporal distances between a reference frame used for the adjacent block and a reference frame used for the current block. For example, a reference frame for an adjacent block may be a previously encoded frame, and a reference frame for the current block may be a frame encoded prior to the previously encoded frame, and the difference may be measured by the number of frames or the temporal distance between the reference frames.
[0086] In some implementations, both an adjacent block and a current block may be predicted based on inter-frame prediction, in which case the overlapping region size determination of the weighting function may be according to the above description. In some implementations, one of the adjacent block or the current block may be predicted based on intra-frame prediction while the other of the adjacent block or the current block is predicted based on inter-frame prediction, and thus a usable comparison of prediction parameters may not be made. When a comparison of prediction parameters cannot be made, the weighting function may define an overlapping region size according to a predetermined function of the current block size. For example, the overlapping region size may be defined as a small overlapping region, such as ¼ of the current block length. As another example, the size for the overlapping region may be set to zero, or there may be no overlapping region, because the adjacent prediction may be considered too different from the current block prediction and the overlapping prediction may be omitted.
[0087] In some implementations, the defined range of overlapping region sizes may be between (0,0) which may indicate no overlapping region and which may indicate the current block size. The weighting function for the overlapping prediction may adjust the defined overlapping region size based on the difference between prediction parameters. For example, for an overlapping region, such as Figure 7 the overlapping region 732 shown in, the motion vector value of the adjacent block 722 may be very similar to the motion vector value of the current block of the current block, such as Figure 7 shown in, and the size adjustment of the defined overlapping size may be omitted. As another example, the motion vector value of an adjacent block of an adjacent block, such as Figure 7 shown in, may be different from that of a current block of a current block, such as Figure 7The motion vector value of the current block of the current block 720 shown in the figure may have a difference that exceeds the established threshold, and the size of the overlapping region may be adjusted. For example, the overlapping region may be expanded, as shown for the overlapping region 902, or may be shrunk, as shown for the overlapping region 904, as Figure 8 shown. In some implementations, adapting the size of the overlapping region based on the difference between prediction parameters may include adapting the weighted function of the overlapping prediction such that the weighting is processed to bias towards the contribution of the current block prediction parameters or the adjacent block prediction parameters depending on which prediction parameters optimize the overlapping prediction of the current block. For example, by setting at least one dimension of the overlapping region to be smaller than the corresponding dimension of the current block, for some pixels in the current block, the weighted function may weight the contribution from the adjacent block prediction parameters to zero.
[0088] In some implementations, under the condition that the difference between the prediction parameters of the current block and the adjacent block exceeds the threshold, the overlapping region may be omitted (i.e., the size of the overlapping region is 0 0). In some implementations, under the condition that there is a small difference or no difference between the prediction parameters of the current block and the adjacent block, the overlapping region may be omitted. For example, the current block prediction may be generally similar to the adjacent block prediction, the difference between the prediction parameters may be less than the minimum threshold, and the size of the overlapping region may be 0 0.
[0089] In some implementations, the prediction parameters for the current block 720 may be used to determine a base prediction for the current block 720. The base prediction may then be the base prediction for each of the overlapping regions 731 to 734. For example, a base prediction block may be determined for the entire current block such that the pixel values for the base prediction can be stored for later use in determining the overlapping prediction for each pixel in the overlapping regions of the current block 720.
[0090] In some implementations, for an overlapping region such as Figure 7 shown in the figure, the prediction may be determined for each of the overlapping regions 731–734 based on the prediction parameters of the adjacent block associated with the overlapping region. For example, prediction parameters including the corresponding reference frame and motion vector of the adjacent block such as Figure 7 shown in the figure for the adjacent block 722 may be used to determine the prediction for the pixels in the overlapping region such as Figure 7 shown in the figure for the overlapping region 732.
[0091] In some implementations, for one or more of the overlapping regions such as Figure 7 shown in the figure for the overlapping regions 731–734, the overlapping prediction may be determined as a weighted function of the base prediction and the prediction based on the corresponding adjacent prediction parameters. For example, for such asFigure 7 The prediction of the overlap for each pixel in the overlap region of the overlap region shown in may be an average of a base prediction value and a predicted pixel value generated based on corresponding neighboring prediction parameters. In some implementations, there may be more than one overlap region for the pixels in the current block. For example, two or more neighboring overlap regions may overlap, such as Figure 7 the overlap regions 731, 732, and 733 shown in, and the prediction of the overlap may be determined as an average of a base prediction based on the prediction parameters for the current block and n predictions based on the corresponding prediction parameters for each of the n neighboring blocks associated with the overlap region. For example, referring to Figure 7 , the pixels in two overlap regions 731 and 734 correspond to two predictions based on corresponding neighboring prediction parameters (i.e., n = 2), and these two predictions may be averaged with the base prediction to determine the prediction of the overlap. In some implementations, each pixel in the overlap region 731 may be determined as an average of a base prediction using the prediction parameters of the current block 720, a prediction based on the prediction parameters of the neighboring block 721, and a prediction based on the prediction parameters of the neighboring block 724.
[0092] In some implementations, the weighting function for the prediction of the overlap may be a function of the distance between the center of the current block and the centers of the neighboring blocks associated with the overlap region. For example, the weighting function may determine a prediction of the overlap that is biased towards smaller-sized neighboring blocks, which may include pixels that are on average located closer to the current block compared to larger neighboring blocks, may be more reliable, and provide a better prediction of the current block. For example, the weighting function may weight the overlap region 732 to contribute more heavily to the prediction of the overlap of the current block 720 than the larger overlap region 734, because the center of the neighboring block 722 is closer to the center of the current block 720 compared to the center of the neighboring block 724.
[0093] Figure 9 is a block diagram of an example weighting function for the prediction of the overlap according to an implementation of the present disclosure. The prediction of the overlap may be optimized by a weighted average of a first prediction and n neighboring-block-based predictions. For example, P0 may indicate a prediction using the current block prediction parameters, ω0 may indicate the weight for predicting P0, P n may indicate a prediction using the neighboring block prediction parameters, ω n may indicate the weight for predicting P n , and the weighting of the prediction OP of the overlap of the pixel 952 may be expressed as follows:
[0094] In some implementations, one or more predicted pixel values at each pixel in the overlap region may be weighted according to a weighting function based on the relative pixel positions with respect to adjacent blocks associated with the overlap region. For example, the overlapping predictions may be weighted such that when a pixel is located relatively closer to an adjacent block, the contribution made by the prediction based on the prediction parameters of the adjacent block is greater. For example, Figure 9 the pixel 952 in the overlap region 734 shown in Figure 9 has a relative distance 954 to the center of the corresponding adjacent block 724 and a relative distance 955 to the center of the current block 720. In some implementations, the overlapping prediction weights 、 may be a function of the relative distances 954, 955. For example, may indicate the relative distance from the pixel to the center of the current block, may indicate the relative distance from the pixel to the center of the adjacent block n, and the weighting function may be a ratio of the relative distance values, which may be expressed as follows:
[0095] In some implementations, the overlapping prediction weights 、 may be a function of the directional relative distance between the pixel and the boundaries between the adjacent block and the current block, such as a function of the horizontal relative distance 964 for the left adjacent block 724. For example, the weighting function may be based on a raised cosine window function, where for pixels located at the adjacent edge of the overlap region n, the weights 、 are equal, and for pixels located at the edge of the overlap region farthest from the adjacent block n, the weights are = 1, = 0. As another example, the overlapping prediction weights 、 may be a function of the vertical relative distance between the pixel and the nearest edge of the adjacent block, such as a function of the vertical relative distance 963 of the pixel 953 with respect to the top adjacent block 723.
[0096] In some implementations, the type of the weighting function for overlapping predictions (such as that shown by the encoder 400 in Figure 4 ) is encoded using an index and is included in a compressed video bitstream such as the compressed bitstream 404 shown in Figure 4 as being for (such as by the Figure 4 encoder 400 shown in Figure 4 ) and is included in a compressed video bitstream such as the compressed bitstream 404 shown in Figure 4 as being for (such as by the Figure 4 compressed bitstream 404 shown in Figure 4 ) in the compressed video bitstream as being for (such as by the Figure 5The decoder 500 shown in decodes an indication of which weighting function is used for the overlapping prediction. For example, various raised cosine weightings can be mapped to a first index set, and various weighting functions based on the relative distance to the block center point can be mapped to a second index set.
[0097] The weighting function for the overlapping prediction can be a combination of any or all of the weighting functions described in the present disclosure. For example, the weighting function can be implemented to weight the overlapping prediction by adaptively adjusting the overlapping region size, by weighting each of the base prediction and the overlapping prediction for the current block, or a combination thereof.
[0098] The above describes an implementation of the overlapping prediction that only uses the motion information of the neighboring (peripheral and adjacent) blocks of the current block, where the neighboring blocks are above or to the left of the current block. However, other implementations are possible. In some cases, the overlapping prediction can be applied to sub-blocks of a block, and the block can be the largest coding unit (which can be referred to as a macroblock or superblock), or a block smaller than the largest coding unit. In an example, the coding mode can indicate that the block is to be predicted at a certain sub-block level. Thus, for example, a block of size N×N (e.g., 16×16) can be divided into b 2 (e.g., b = 4) blocks of M×M (e.g., 4×4), where N = b*M. The overlapping prediction can be performed on at least some of the b 2 blocks. For ease of reference, the overlapping prediction at the sub-block level is also referred to as sub-block overlapping prediction in this document.
[0099] Figure 10 FIG. 1000 is a block diagram showing sub-block overlapping prediction. The sub-block overlapping prediction can be used to smooth (e.g., correct) the boundaries of the sub-blocks of a block, thereby reducing the blocking artifacts of the sub-blocks. Similar to the overlapping prediction described above, in the sub-block overlapping prediction, the prediction obtained using the motion information (e.g., motion vector and reference frame) of the current sub-block is combined (e.g., weighted) with the prediction obtained using the corresponding motion information of one or more neighboring blocks. However, in the sub-block overlapping prediction, the neighboring blocks can be peripheral neighboring blocks, sub-blocks of the same block as the current sub-block, blocks after the current sub-block in raster scan order, or a combination thereof.
[0100] Block diagram 1000 includes a block 1002 divided into sub-blocks. The sub-blocks of block 1002 are numbered from 0 to 15. Although Figure 10 FIG. shows that block 1002 is divided into 16 sub-blocks, the present disclosure is not limited thereto. Block 1002 can be divided into more or fewer sub-blocks. The number of sub-blocks can depend on the size of block 1002.
[0101] In the prediction of sub-block overlap, motion information of at least some of the available blocks to the left, above, to the right, and below the current sub-block can be used. For illustration, when obtaining the prediction for sub-block 1004, the current prediction P0 of sub-block 1004 is obtained using the motion information (e.g., motion vector and reference frame) determined for sub-block 1004 (the block numbered 9), and P L the predicted block is obtained using the motion information of the upper left sub-block 1006, and P T the predicted block is obtained using the motion information of the upper sub-block 1008, and P R the predicted block is obtained using the motion information of the right sub-block 1010, and P B the predicted block.
[0102] P0, P L , P T , P R , and P B can be used to obtain the final predicted block as a weighted sum. In an example, the predictions can be combined in a certain order. In an example, the order can be a loop starting from the left neighboring block. For example, P1 can be obtained as (P0 + P L ) / 2, then P2 can be obtained as (P1 + P T ) / 2, then P3 can be obtained as (P2 + P R ) / 2, and then the final prediction is obtained as (P3 + P B ) / 2. In this way, only shift operations need to be performed. As another illustration, the motion information of block 1016, block 1018, sub-block 1020, and sub-block 1022 can be used to obtain the prediction of sub-block 1014.
[0103] Statements such as "applying (or performing) overlap prediction for the current block using neighboring blocks" or "obtaining overlap prediction based on neighboring blocks" in this article should be understood to mean that the motion information of neighboring blocks will be used to obtain the prediction for the current block, and the obtained prediction will be included in the calculation of the final predicted block for the current block.
[0104] In some cases, if one or more conditions are applicable, the overlap prediction may not be applied to or may not be used for the current block (e.g., sub-block). Examples of such conditions are now provided.
[0105] In an example, overlapping prediction can be disabled for all frames of a video sequence. For example, a syntax element in a Sequence Parameter Set (SPS) can indicate that overlapping prediction will not be applied to any block of any frame of the video sequence. As is known, the SPS can contain parameters common to the entire video sequence (i.e., each frame in the frames of the video sequence). In an example, overlapping prediction can be disabled for a group of frames. For example, a syntax element in a Picture Parameter Set (PPS) can indicate that overlapping prediction will not be applied to the group of frames corresponding to the PPS. As is known, the PPS can contain parameters common to all frames in the group of frames. In an example, overlapping prediction may not be performed on a current block that is intra-predicted.
[0106] In an example, if the size of a block is less than or equal to a threshold size, overlapping prediction may not be applied to the block (e.g., not performed on the block). For example, overlapping prediction may not be performed on a block smaller than or equal to 32×32 pixels. In an example, the block header of the current block can include one or more syntax elements indicating whether overlapping prediction is to be performed on the block. For example, the syntax element can be a prediction mode indicating that overlapping prediction is to be performed on the block. In an example, one or more syntax elements can be a flag indicating whether overlapping prediction is to be performed. Thus, if one or more syntax elements indicate that overlapping prediction will not be performed on the current block, overlapping prediction is not performed on the current block.
[0107] In some examples, if the current block is not coded using the SKIP model or the MERGE mode, a flag indicating whether overlapping prediction is to be performed can be included in the header of the current block. That is, if the current block is encoded using one of the SKIP or MERGE modes, overlapping prediction is to be performed on the current block. The SKIP and MERGE modes are now briefly described. If a block is encoded using the SKIP mode, for the current block, no residual information is sent from the encoder to the decoder. The decoder can estimate the motion of the current block encoded using the SKIP mode from a candidate motion vector list and can use (e.g., select) the motion vector to calculate the motion-compensated prediction for the current block. In the MERGE mode, a motion vector from the candidate motion vector list is inherited for coding the current block. The candidate motion vector list can also be referred to as a merge list, where the merge list can reference blocks whose motion vectors (or more generally, motion information) are used to select the motion vector (or more generally, motion information) for the current block.
[0108] In some cases, it may not be desirable to use overlapping prediction.
[0109] For example, when the current block includes screen content, overlapping prediction may not be efficient, even if the current block is encoded using, for example, one of the MERGE or SKIP modes. In such cases, overlapping prediction may cause blurring of sharp edges in the screen content during decoding. As mentioned above, flags can be signaled for non-MERGE and non-SKIP predicted blocks. However, it may be useful to further indicate whether to perform or not perform overlapping prediction for such blocks.
[0110] As another example, when multiple reference frames are available for the current frame, overlapping prediction may not be efficient. When multiple reference frames are available, overlapping prediction may require obtaining different samples (pixel values) from different reference pictures, which can significantly increase the memory bandwidth requirements. For illustration, reference frames R1 and R3 can be used to predict (e.g., bidirectionally) sub-block 1014, reference frames R1 and R2 can be used to bidirectionally predict block 1016, reference frames R1 and R3 can be used to bidirectionally predict sub-block 1020, reference frames R3 and R4 can be used to bidirectionally predict sub-block 1020, and reference frame R3 can be used to unidirectionally predict sub-block 1022. Thus, performing overlapping prediction on sub-block 1014 will require obtaining samples from four different reference frames, namely reference frames R1, R2, R3, and R4. This process may have to be performed for each block in the picture.
[0111] Additionally, performing overlapping prediction on blocks of a P slice or frame greatly increases the coding processing complexity of the P slice or frame. The coding processing of blocks of a P slice or frame is desirably of the lowest possible complexity, especially in the case of real-time use of video coding.
[0112] The following techniques can be used to solve (or at least mitigate) problems such as the foregoing regarding overlapping prediction.
[0113] Figure 11 is a flowchart of a technique 1100 for coding processing a current block of a video frame using overlapping prediction. Technique 1100 can be implemented as, for example, a software program executable by a computing device of one or more of computing and communication devices 100A / 100B / 100C such as those for Figure 2 The software program can include machine-readable instructions that can be stored in a memory 150 such as Figure 1 and, when executed by a processor 140 such as Figure 1 can cause the computing device to perform technique 1100. Technique 1100 can be implemented in whole or in part by an intra / inter prediction unit 540 of a decoder 500 such as Figure 5 Technique 1100 can be implemented in whole or in part by Figure 4Implemented by the intra / inter prediction unit 410 of the encoder 400. Technique 1100 can be implemented using dedicated hardware or firmware. Multiple processors, memories, or both can be used.
[0114] Technique 1100 can conditionally apply overlapping prediction (e.g., prediction of sub-block overlap) to the current block (e.g., current sub-block) based on information available at the decoder such as motion information or predicted sample values of blocks adjacent to the current block and that can be used to perform overlapping prediction, such as as Figure 10 described.
[0115] In an example, technique 1100 can be applied to all blocks predicted using inter prediction in the current frame. In an example, it can be inferred at the decoder based on information available at the decoder whether to apply overlapping prediction, and no block-level syntax element is required to indicate whether overlapping prediction will be performed on a block. In another example, the condition can be applied only when block-level overlapping prediction is signaled or derived to be executed.
[0116] In an example, technique 1100 may not be performed on P slices (i.e., all blocks of P slices). Instead, to reduce complexity, when the current block and adjacent blocks share the same reference frame and the absolute motion vector difference between the motion vector of the current block and the motion vector of the adjacent block is below a predefined or signaled motion vector threshold, the motion information of the adjacent block can be used to perform overlapping prediction on the current block in the P slice, which can be as described elsewhere in this document.
[0117] In an example, overlapping prediction is applied to the current block only when the current block is coded using a specific prediction mode. In an example, the specific prediction mode can include modes known as affine mode and sub-block based temporal motion vector prediction (SbTMVP) mode in the MPEG versatile video coding (VVC), ITU-T H.266 video standard.
[0118] In short, the affine mode uses more degrees of freedom (parameters) than classical translation using motion vectors that use two parameters. For example, the affine mode can use four parameters (for implementing translation, rotation, and scaling) or six parameters (for implementing translation, rotation, scaling, shearing, and aspect ratio change). In short, the SbTMVP mode uses the motion field within the collocated frame of the current frame to improve motion vector prediction (MVP) and the merge mode of the coding processing units within the current frame. In the SbTMVP mode, motion prediction can be performed at the sub-block level or at the sub-coding processing unit (sub-CU) level. Additionally, SbTMVP applies a motion shift from the collocated frame and thereafter derives temporal motion information. The motion shift can include the process of obtaining a motion vector from one of the spatial neighboring blocks of the current block and shifting that motion vector.
[0119] As already mentioned, applying overlapping prediction from neighboring blocks of the current block can be defined as or include obtaining a prediction based on the motion information of the neighboring blocks (e.g., one or more motion vectors) and including this prediction in a weighted prediction that includes a prediction obtained for the current block using the motion information associated with the current block.
[0120] At 1110, a first prediction block is obtained for the current block based on the motion information associated with the first current block. The motion information can be or include a first reference frame and a first motion vector. More generally, the motion information can include more than one motion vector (and correspondingly, more than one reference frame). In an example, when technique 1100 is implemented by a decoder, the motion information can be decoded from a compressed bitstream 502 such as Figure 5 of the compressed bitstream.
[0121] At 1120, technique 1100 determines to use an overlapping prediction mode to obtain at least a portion of a second prediction block for the current block, the overlapping prediction mode using a second reference frame and a second motion vector of a neighboring block. The second reference frame and the second motion vector are associated with a neighboring block of the current block. The motion information of the neighboring block can include more motion vectors than the second motion vector (and correspondingly, more reference frames than the second reference frame). The determination is at least partially based on information related to the neighboring block of the current block.
[0122] As used herein, the second motion vector can be the actual motion vector of the neighboring block (i.e., the motion vector used to obtain the prediction of the neighboring block) or a motion vector obtained therefrom. Thus, and for ease of reference, in one example, the "second motion vector" refers to the actual motion vector of the neighboring block; and in another example, the "second motion vector" refers to a motion vector obtained from the actual motion vector.
[0123] For example, a second prediction block may be obtained using a second motion vector obtained by rounding an actual motion vector to an integer position. By rounding the actual motion vector to an integer position, an interpolation process that would otherwise be performed to obtain a sub-pixel value may be bypassed (e.g., avoided). The horizontal and vertical components of the actual motion vector (MV x , MV y ) may be rounded towards positive infinity, towards negative infinity, or towards zero. In another example, the horizontal and vertical components may be rounded independently depending on their respective values. That is, if the component of the actual motion vector is positive, the corresponding component of the second motion vector may be obtained by rounding towards positive infinity; and if the component of the actual motion vector is negative, the corresponding component of the second motion vector may be obtained by rounding towards negative infinity.
[0124] In an example, it is determined to obtain a second prediction block in response to determining that the current block and a neighboring block have at least some same reference frames in the same reference frame (i.e., are predicted using at least some same reference frames in the same reference frame). Thus, if the current block and the neighboring block are not predicted using at least some same reference frames in the same reference frame, overlapping prediction is not performed (applied) from the neighboring block. To illustrate, if the current block is bi-directionally predicted using reference frames R1 and R2, and the neighboring block is also bi-directionally predicted using reference frames R2 and R3, then at 1130, only reference frame R2 (and the motion vectors used with it to predict the neighboring block) is used to obtain the second prediction block. In an example, it is determined to obtain a second prediction block in response to determining that the current block and the neighboring block have the same reference frame (i.e., are predicted using the same reference frame). Thus, in this case, overlapping prediction is performed from the neighboring block using the common reference frame between the current block and the neighboring block.
[0125] Not performing (applying) the overlapping prediction from the neighboring block means that the prediction block for the current block is not obtained based on the motion information of the neighboring block and is not used (included) in obtaining the final prediction block for the current block.
[0126] In an example, it is determined to obtain a second prediction block in response to determining that the motion vector difference between a first motion vector and a second motion vector is less than a motion vector threshold. In an example, the motion vector threshold may depend on the distance between a first reference frame and a second reference frame. For example, the motion vector threshold may be proportional to the frame distance. That is, the motion vector threshold may be proportional to the temporal distance between the first reference frame and the second reference frame.
[0127] In an example, the motion vector threshold may be less than or equal to n pixels per single frame distance (i.e., "single frame distance motion vector threshold"). In an example, n is equal to 16 pixels per single frame distance. For example, assuming that the first reference frame has a frame number of R and the second reference frame has a frame number of S, the motion vector threshold is calculated as to be calculated. In an example, the motion vector difference between the first motion vector given by (MV x,1 , MV y,1 ) and the second motion vector given by (MV x,2 , MV y,2 ) is calculated as to be calculated. If the motion vector difference is not less than the motion vector threshold, prediction of overlap from neighboring blocks may not be performed.
[0128] In an example, the single frame distance motion vector threshold may be decoded from the compressed bitstream at the decoder. Thus, the motion vector threshold can be obtained by decoding the single frame distance motion vector threshold from the compressed bitstream and calculating the motion vector threshold based on the frame difference between the first reference frame and the second reference frame and the single frame distance motion vector threshold (e.g., as their product). The single frame distance motion vector threshold may be signaled (e.g., encoded) in the SPS, PPS, slice header, block header, or some other header. In another example, the single frame distance motion vector threshold may be predefined (e.g., preconfigured). When implemented at the encoder, the single frame distance motion vector threshold may be calculated by the encoder and encoded in the Figure 4 compressed bitstream 404.
[0129] In an example, if the first reference frame is different from the second reference frame, the first motion vector and the second motion vector may be scaled to a target reference frame. Then the scaled first motion vector and the scaled second motion vector are used to calculate the motion vector difference. As already mentioned, more than one motion vector may be used to predict one or both of the current block and neighboring blocks.
[0130] The scaled motion vectors can be summarized as follows. In response to determining that multiple first motion vectors (e.g., more than one motion vector) including a first motion vector are used to predict a current block, a scaled first motion vector is then obtained by scaling the first motion vector based on a temporal distance to point to a target reference frame. Then, the scaled first motion vectors are averaged to obtain a first normalized motion vector. Additionally, in response to determining that multiple second motion vectors (e.g., more than one motion vector) including a second motion vector are used to predict a neighboring block, a scaled second motion vector is then obtained by scaling the second motion vector based on a temporal distance to point to a target reference frame. Then, the scaled second motion vectors are averaged to obtain a first normalized motion vector. Then, a motion vector difference is obtained based on: 1) one of the first normalized motion vector (if computed) or the first motion vector (if the first normalized motion vector is not computed), and 2) one of the second normalized motion vector (if computed) or the second motion vector (if the second normalized motion vector is not computed).
[0131] Thus, the motion vector threshold can be directly based on the normalized motion vectors (e.g., compared with these normalized motion vectors). By way of illustration, the normalized motion vectors for the bottom neighboring block and the current block can be MV n_B and MV n_C . respectively. If |MV n_B - MV n_B | > the threshold (i.e., the motion vector threshold), then the prediction according to the overlap of the bottom neighboring block is not applied.
[0132] Now, the scaling of the motion vectors is described. Assume that the current block is a block of the current frame with a temporal index C0, the motion vector (MV x , MV y ) of this block points to a reference frame with a temporal index R1, and the motion vector is to be scaled to a target reference frame with a temporal index R2. Further assume that the distance between C0 and R1 is b (i.e., b = |C0 - R1|); and the distance between C0 and R2 is d (d = |C0 - R2|). Thus, MV x,scaled = b / d MV x and MV y,scaled = b / d MV y can be used to obtain the scaled motion (MV x,scaled , MV y,scaled ).
[0133] In another example, a second prediction block for obtaining at least a portion of a current block using an overlapping prediction mode is determined based on a determination of the availability of reference samples required for interpolation, the overlapping prediction mode using a second reference frame and a second motion vector of neighboring blocks. Thus, if the reference samples required for interpolation are not available, the overlapping prediction mode is not performed on the current block.
[0134] To illustrate, if the second motion vector includes a fractional part (i.e., references a sub-pixel position), sub-pixel values are obtained via interpolation using an interpolation filter. Assume the interpolation filter has T taps and the current block has size M×N, then a reference block larger than M×N is required to generate the second prediction block. The reference block size will have to be of size (M+T−1)×(N+T−1). The upper left block corner of the reference block can be given by an integer pixel at position (MV x −(T / 2−1), MV y −(T / 2−1)), where (MVx, MV y ) is the second motion vector for the reference block. Thus, to reiterate, a determination to obtain a second prediction block is made in response to a determination that reference samples of a reference sample block of size (M+T−1)×(N+T−1) are available.
[0135] In an example, determinations to obtain respective second prediction blocks for different boundaries of a current block can be made in parallel. In other words, determinations as to whether overlapping prediction is to be applied to different boundaries can be generated in parallel. In this case, the first prediction block obtained at 1110 can be a first prediction block that includes a parent block of the current block. A parent block is a block of which the current block is a child block. More specifically, the first prediction block can be a portion of the prediction block of the parent block, where the portion corresponds to the current block (e.g., is co-extensive with the current block). The prediction block of the parent block is referred to herein as the “original prediction block”.
[0136] Reference Figure 7, and for convenience, the original prediction for the parent block 735 (or equivalently, the original reference sample used to obtain the original prediction) may be referred to as P735. For the boundary between the parent block 735 and the neighboring block 724 (i.e., for the sub-blocks along the boundary), the predicted sample (or the reference sample for it) and the motion vectors of the neighboring block 724 and P735 may be examined as described herein; and for the boundary between the neighboring block 721 and the parent block 735, the predicted sample (or reference sample) together with the motion vectors of the neighboring block 721 and P735 are also examined as described herein. Since P735 is used for each of the boundaries, overlapping prediction decisions regarding different boundaries (i.e., where to perform overlapping prediction for the sub-blocks along different boundaries) can be made in parallel. In the case where the parent block 735 is bi-directionally predicted, the result of the bi-directional prediction or the average of the two reference sample blocks used for the bi-directional prediction can be used.
[0137] Thus, the technique 1100 can determine whether to perform an overlapping prediction mode regarding a second boundary of the parent block in parallel with determining whether to obtain a second predicted block of at least that portion of the current block using an overlapping prediction mode, where the second boundary is different from the first boundary. In an example, the determination of whether to perform overlapping prediction can be performed in parallel for each boundary of the parent block as a whole.
[0138] At 1130, an overlapping prediction mode is used to obtain a second predicted block. At 1140, the first predicted block and the second predicted block are combined. Combining the first predicted block and the second predicted block may mean including the first predicted block and the second prediction in one or more computations that result in obtaining a final predicted block for the current block.
[0139] Figure 12 is a flowchart of another technique 1200 for coding a current block of a video frame using overlapping prediction. The technique 1100 can be implemented as, for example, a software program executable by a computing device of one or more of the computing and communication devices 100A / 100B / 100C such as those for Figure 2 The software program may include machine-readable instructions that can be stored in a memory 150 such as Figure 1 and when executed by a processor 140 such as Figure 1 can cause the computing device to execute the technique 1200. The technique 1200 can be implemented in whole or in part by Figure 5 the intra / inter prediction unit 540 of the decoder 500. The technique 1100 can be implemented in whole or in part by Figure 4 the intra / inter prediction unit 410 of the encoder 400. The technique 1200 can be implemented using dedicated hardware or firmware. Multiple processors, memories, or both can be used.
[0140] Technique 1200 determines whether to apply prediction based on the overlap of a particular neighboring block based on the prediction difference between the predicted sample of the current block and the prediction of the current block based on the motion information (e.g., motion vector) of the particular neighboring block. In an example, the prediction difference can be the sum of absolute differences (SAD). In another example, the prediction difference can be the sum of squared errors (SSE). However, as further described herein, other metrics are possible for the prediction difference.
[0141] At 1210, a first prediction block for the current block is obtained based on the motion information associated with the current block. The motion information associated with the first block can include one or more motion vectors and the corresponding one or more reference frames. At 1220, a second prediction block for at least a portion of the current block is obtained based on the motion information associated with the neighboring block. The motion information associated with the neighboring block can include one or more motion vectors and the corresponding one or more reference frames.
[0142] At 1230, a prediction difference metric between the first prediction block and the second prediction block is determined. The prediction difference metric is calculated as a function of the pairwise differences between the values of the first prediction block and the values of the second prediction block. As already mentioned, the prediction difference metric can be SAD, SSE, or some other metric.
[0143] In an example, the prediction difference metric can be the absolute maximum of the pairwise differences between the first prediction block and the second prediction block. In another example, the prediction difference metric can be obtained as the absolute difference between the maximum absolute difference and the average absolute difference. The maximum absolute difference is the absolute maximum of the pairwise differences between the first prediction block and the second prediction block; and the average absolute difference is the average of the pairwise differences between the first prediction block and the second prediction block. In another example, the prediction difference metric is calculated based on the ratio of the maximum absolute difference to the average absolute difference.
[0144] At 1240, based on the prediction difference metric, it is determined whether to combine the first prediction block and the second prediction block of this portion of the current block. In an example, if the prediction difference metric exceeds a threshold, the second prediction block is not included in the prediction based on the overlap for the current block. In an example, the threshold can be signaled (e.g., encoded) in the SPS, PPS, slice header, block header, or some other header. In an example, and to simplify the hardware implementation of Technique 1200, the threshold can be a power of 2 (e.g., 2 7 = 128). In an example, the threshold can be adjusted or scaled based on the coding / decoding bit depth. For illustration, assume a base threshold T (e.g., T = 128) and a bit depth of d bits per pixel (e.g., d = 10 bits), then the threshold can be calculated as while being signaled. In an example, the base threshold can be signaled.
[0145] As described with respect to Technique 1100, in some examples, if certain conditions apply (i.e., are met), then it is determined to perform Technique 1200 on an adjacent block. In an example, a second predicted block is obtained in response to determining that the motion vector difference between a first motion vector of a current block and a second motion vector of an adjacent block is less than a motion vector threshold. In an example, a second predicted block is obtained in response to determining that a first reference frame used to predict the current block and a second reference frame used to predict the adjacent block are at least partially the same. In an example, a second predicted block is obtained in response to determining that the current block is a block of a P frame or a P slice, the current block and the adjacent block use the same reference frame, and the corresponding absolute motion vector difference between the motion vector of the current block and the motion vector of the adjacent block is below the motion vector threshold. In another example, compliant bitstream constraints may be applied to prohibit overlapping predictions in a P slice.
[0146] For simplicity of explanation, Figure 6 、 Figure 11 and Figure 12 Techniques 600, 1100, and 1200 of [、] and [、] are depicted and described as respective series of steps. However, the steps according to the present disclosure may occur in various orders, concurrently, and / or iteratively. Additionally, the steps according to the present disclosure may occur in conjunction with other steps not presented and described herein. Furthermore, not all of the illustrated steps may be required to implement the techniques according to the disclosed subject matter.
[0147] The word "example" or "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "example" or "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects or designs. Instead, the use of the word "example" or "exemplary" is intended to present concepts in a concrete manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X includes A or B" is intended to mean any of the natural inclusive permutations. That is, if X includes A; X includes B; or X includes both A and B, then "X includes A or B" is satisfied under any of the foregoing instances. Additionally, as used in this application and the appended claims, the articles "a" and "an" shall generally be construed to mean "one or more" unless otherwise specified or clear from the context that it refers to the singular form. Furthermore, the use of the terms "embodiment" or "an embodiment" or "implementation" or "an implementation" throughout the text is not intended to mean the same embodiment or implementation unless so described. As used herein, the terms "determine" and "identify", or any variant thereof, include using Figure 1 one or more of the means shown in [] to select, confirm, calculate, find, receive, determine, establish, obtain, or otherwise identify or determine in any manner.
[0148] Implementations of computing and communication devices, such as one or more of computing and communication devices 100A / 100B / 100C (and algorithms, techniques, methods, instructions, etc. stored on and / or executed by them), can be implemented in hardware, software, or any combination thereof. The hardware can include, for example, a computer, an intellectual property (IP) core, an application specific integrated circuit (ASIC), a programmable logic array, an optical processor, a programmable logic controller, microcode, a microcontroller, a server, a microprocessor, a digital signal processor, or any other suitable circuitry. In the claims, the term "processor" should be understood to cover any one of the foregoing hardware, either individually or in combination. The terms "signal" and "data" may be used interchangeably. Further, the various parts of computing and communication devices 100A / 100B / 100C do not necessarily have to be implemented in the same manner.
[0149] Further, in one implementation, for example, a computing and communication device can be implemented using a computer program that, when executed, implements any one of the corresponding methods, algorithms, and / or instructions described herein. Additionally or alternatively, for example, a dedicated computer / processor can be utilized that can include dedicated hardware for executing any one of the methods, algorithms, or instructions described herein.
[0150] A computing and communication device can be implemented, for example, on a computer in a real-time video system. Alternatively, one computing and communication device (e.g., computing and communication device 100A) can be implemented on a server, and another computing and communication device (e.g., computing and communication device 100B) can be implemented on a device separate from the server, such as a handheld communication device.
[0151] Furthermore, all or part of an implementation can take the form of a computer program product accessible from, for example, a tangible computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can tangibly contain, store, transmit, or transport a program for use by or in connection with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media are also available.
[0152] The foregoing implementations have been described to facilitate understanding of the present application, but are not limiting. On the contrary, the present application covers various modifications and equivalent arrangements within the scope of the appended claims, which scope should be given the broadest interpretation so as to cover all such modifications and equivalent structures as are permitted by law.
Claims
1. A method for coding a current block of a current frame, comprising: Obtaining a first prediction block based on motion information associated with a first prediction block for the current block; Obtaining a second prediction block for at least a portion of the current block based on motion information associated with neighboring blocks; Obtaining a prediction difference metric between the first prediction block and the second prediction block; And Determining whether to combine the first prediction block and the second prediction block of the portion of the current block based on the prediction difference metric.
2. The method according to claim 1, wherein determining whether to combine the first prediction block and the second prediction block of the portion of the current block based on the prediction difference metric comprises: Determining not to combine the first prediction block and the second prediction block in response to the prediction difference metric exceeding a threshold.
3. The method according to claim 2, wherein the threshold is a power of 2.
4. The method according to any one of claims 1 to 3, wherein the prediction difference metric is the sum of absolute differences SAD between the first prediction block and the second prediction block.
5. The method according to any one of claims 1 to 3, wherein the prediction difference metric is the sum of squared errors SSE between the first prediction block and the second prediction block.
6. The method according to any one of claims 1 to 3, wherein the prediction difference metric is the absolute maximum value of pairwise differences between the first prediction block and the second prediction block.
7. The method according to any one of claims 1 to 6, wherein the prediction difference metric is calculated based on at least one of the maximum absolute difference or the average absolute difference, wherein the maximum absolute difference is the absolute maximum value of pairwise differences between the first prediction block and the second prediction block, and wherein the average absolute difference is the average value of the pairwise differences.
8. The method according to claim 7, wherein the prediction difference metric is calculated as the absolute difference between the maximum absolute difference and the average absolute difference.
9. The method according to claim 7, wherein the prediction difference metric is calculated based on the ratio of the maximum absolute difference to the average absolute difference.
10. The method according to any one of claims 1 to 9, further comprising: Determining to obtain the second prediction block in response to determining that the motion vector difference between a first motion vector of the current block and a second motion vector of the neighboring block is less than a motion vector threshold.
11. The method according to any one of claims 1 to 9, further comprising: Determining to obtain the second prediction block in response to determining that a plurality of first reference frames for predicting the current block and a plurality of second reference frames for predicting the neighboring block are at least partially the same.
12. The method according to any one of claims 1 to 9, further comprising: Determining to obtain the second prediction block in response to determining that the current block is a block of a P frame or a P slice, the current block and the neighboring block use the same reference frame, and the corresponding absolute motion vector difference between the motion vector of the current block and the motion vector of the neighboring block is below a motion vector threshold.
13. A method for coding a current block of a current frame, comprising: Obtain a first prediction block for the current block based on a first reference frame and a first motion vector; Determine, at least in part based on information related to neighboring blocks of the current block, to obtain a second prediction block for at least a portion of the current block using an overlapping prediction mode that uses a second reference frame and a second motion vector of the neighboring blocks, wherein the second motion vector of the neighboring blocks is obtained by rounding the motion vector used to predict the neighboring block to an integer pixel position; Obtain the second prediction block using the overlapping prediction mode; And Combine the first prediction block and the second prediction block.
14. The method according to claim 13, wherein determining to obtain the second prediction block for the portion of the current block using the overlapping prediction mode based at least in part on the information related to the neighboring blocks of the current block includes: Determine that a motion vector difference between the first motion vector and the second motion vector is less than a motion vector threshold.
15. The method according to claim 14, further comprising: Decode a single-frame distance motion vector threshold from a compressed bitstream; And Calculate the motion vector threshold based on a frame difference between the first reference frame and the second reference frame and the single-frame distance motion vector threshold.
16. The method according to claim 14, wherein the motion vector threshold is proportional to a temporal distance between the first reference frame and the second reference frame.
17. The method according to claim 14, further comprising: In response to determining that the first reference frame is different from the second reference frame, calculate a difference between the first motion vector and the second motion vector by: In response to determining to predict the current block using a plurality of first motion vectors including the first motion vector: Obtain a scaled first motion vector by scaling the first motion vector based on a temporal distance to point to a target reference frame; And Average the scaled first motion vector to obtain a first normalized motion vector; In response to determining to predict the neighboring block using a plurality of second motion vectors including the second motion vector: Obtain a scaled second motion vector by scaling the second motion vector to point to the target reference frame; and Average the scaled second motion vector to obtain a second normalized motion vector; And Calculate the motion vector difference based on: 1) one of the first normalized motion vector or the first motion vector and 2) one of the second normalized motion vector or the second motion vector.
18. The method according to any one of claims 13 to 17, wherein determining to obtain the second prediction block for the portion of the current block using the overlapping prediction mode based at least in part on the information related to the neighboring blocks of the current block includes: Determine that a plurality of first reference frames for predicting the current block and a plurality of second reference frames for predicting the neighboring block are at least partially the same, where the plurality of first reference frames includes the first reference frame and the plurality of second reference frames includes the second reference frame, and where the first reference frame is the same as the second reference frame.
19. The method according to any one of claims 13 to 17, wherein determining the second prediction block for the portion of the current block to be obtained using the overlapping prediction mode is at least partially based on the information related to the neighboring block of the current block, and includes: Determine that reference samples for obtaining the second prediction block using a sub-pixel interpolation filter are available.
20. The method according to any one of claims 13 to 17, wherein determining the second prediction block for the portion of the current block to be obtained using the overlapping prediction mode is at least partially based on the information related to the neighboring block of the current block, and includes: Determine that the current block is a block of a P frame or a P slice, that the current block and the neighboring block use the same reference frame, and that the corresponding absolute motion vector difference between the motion vector of the current block and the motion vector of the neighboring block is lower than a motion vector threshold.
21. The method according to any one of claims 13 to 20, wherein at least one of rounding towards positive infinity, rounding towards negative infinity, or rounding towards zero is used to round the motion vector for predicting the neighboring block.
22. The method according to claim 21, wherein in the case where the component of the motion vector is positive, the component is rounded towards positive infinity, and in the case where the component is negative, the component is rounded towards negative infinity.
23. The method according to any one of claims 13 to 22, wherein the current block is along a first boundary of a parent block, and the method further includes: Determine whether to perform the overlapping prediction mode with respect to a second boundary of the parent block, which is different from the first boundary, in parallel with determining whether to obtain the second prediction block for at least the portion of the current block using the overlapping prediction mode.
24. An apparatus, comprising: A processor configured to execute the method according to any one of claims 1 to 12.
25. An apparatus, comprising: A memory; And A processor configured to execute instructions stored in the memory to perform the method according to any one of claims 1 to 12.
26. A non-transitory computer-readable storage medium, comprising executable instructions that, when executed by a processor, facilitate the execution of operations, the operations including performing the operations of the method according to any one of claims 1 to 12.
27. A non-transitory computer-readable storage medium, on which an encoded bitstream is stored, wherein the encoded bitstream is configured to be decoded by the method according to any one of claims 1 to 12.
28. A non - transitory computer - readable storage medium having an encoded bitstream stored thereon, wherein the encoded bitstream is generated by an encoder performing the method according to any one of claims 1 to 12.
29. An apparatus, comprising: a processor configured to perform the method according to any one of claims 13 to 23.
30. An apparatus, comprising: a memory; and a processor configured to execute instructions stored in the memory to perform the method according to any one of claims 13 to 23.
31. A non - transitory computer - readable storage medium comprising executable instructions that, when executed by a processor, facilitate the execution of operations including performing the operations of the method according to any one of claims 13 to 23.
32. A non - transitory computer - readable storage medium having an encoded bitstream stored thereon, wherein the encoded bitstream is configured to be decoded by the method according to any one of claims 13 to 23.
33. A non - transitory computer - readable storage medium having an encoded bitstream stored thereon, wherein the encoded bitstream is generated by an encoder performing the method according to any one of claims 13 to 14 or 16 to 23.