Frame level non-linear motion offset in video code processing
By employing frame-level nonlinear motion offset technology in video coding, the problems of inaccurate nonlinear motion representation and wasted computational resources in existing methods are solved, achieving more efficient video coding and decoding.
Patent Information
- Application Number
- CN202480047253.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-18
- Filing Date
- 2024-07-09
- Publication Date
- 2026-02-13
AI Technical Summary
When dealing with nonlinear motion, existing video coding methods, such as conventional block-level motion offset methods, increase computational and signal transmission costs while lacking accuracy and failing to effectively represent nonlinear motion.
The frame-level nonlinear motion offset technique is used to determine a single nonlinear motion offset during encoding or decoding and apply it to the motion vectors of all or part of the current frame to more accurately represent nonlinear motion and reduce additional computation and signal transmission overhead.
It improves the accuracy and compression efficiency of video encoding, reduces the size of the bitstream and the computational resource requirements, while maintaining high compression performance.
Smart Images

Figure CN121533014A_ABST
Abstract
Description
BACKGROUND
[0001] Digital video streams can represent video using a sequence of frames or still images. Digital video can be used for a variety of applications, including, for example, video conferencing, high-definition video entertainment, video advertising, or sharing of user-generated videos. Digital video streams can contain a large amount of data and consume a substantial amount of computing or communication resources of a computing device for processing, transmission, or storage of the video data. Various ways have been proposed to reduce the amount of data in a video stream, including encoding or decoding techniques. SUMMARY
[0002] Herein Among other things Systems and techniques for frame-level non-linear motion offset in video coding are disclosed.
[0003] According to one implementation of the disclosure, a method for decoding an encoded video frame using frame-level non-linear motion offset includes decoding a frame-level non-linear motion offset for the encoded video frame from a bitstream in which the encoded video frame is encoded, performing motion compensation for a block of the encoded video frame including applying the frame-level non-linear motion offset to one or more motion vectors determined for the block, reconstructing the encoded video frame into a reconstructed frame based on an output of the motion compensation, and outputting the reconstructed frame within an output video stream.
[0004] In some implementations of the method, the frame-level non-linear motion offset is determined during encoding of a current frame into the encoded video frame.
[0005] In some implementations of the method, the frame-level non-linear motion offset is determined based on linear motion estimated between the current frame and each of a backward reference frame for the current frame and a forward reference frame for the current frame.
[0006] In some implementations of the method, the linear motion is represented by a single motion vector predicted for the current frame, and the frame-level non-linear motion offset is determined based on a difference between a non-linear motion of the encoded video frame and the single motion vector.
[0007] In some implementations of the method, the non-linear motion of the encoded video frame is determined based on a location of an object associated with the non-linear motion within the current frame.
[0008] In some implementations of the method, the frame-level non-linear motion offset is determined based on block-level offsets determined for a plurality of blocks of the current frame.
[0009] In some implementations of the method, the frame-level non-linear motion offset is one of a simple average, a weighted average, or a mode of the block-level offsets.
[0010] In some implementations of this method, decoding the frame-level nonlinear motion offset of the encoded video frame includes decoding the frame-level nonlinear motion offset from the frame header associated with the encoded video frame within the bitstream.
[0011] In some implementations of this method, motion compensation is performed on blocks of encoded video frames, including applying frame-level nonlinear motion offsets to one or more motion vectors determined for the block, including adding frame-level nonlinear motion offsets to a single motion vector predicted for all encoded video frames in the encoded video frame.
[0012] In some implementations of this method, motion compensation is performed on blocks of encoded video frames, including applying frame-level nonlinear motion offsets to one or more motion vectors determined for the block, including: for each block in the block, adding frame-level nonlinear motion offsets to block-level motion vectors determined for that block.
[0013] In some implementations of this method, motion compensation is performed on blocks of encoded video frames, including applying frame-level nonlinear motion offsets to one or more motion vectors determined for the blocks. This includes: determining an updated motion field of the encoded video frame by applying frame-level nonlinear motion offsets to an initial motion field; determining a temporally interpolated picture reference frame based on the updated motion field; and using the temporally interpolated picture reference frame to predict the encoded video frame.
[0014] According to an implementation of this disclosure, a non-transitory computer-readable medium stores an encoded bitstream, wherein the encoded bitstream is configured to be decoded by operations including: performing motion compensation on a block of an encoded video frame by applying a frame-level nonlinear motion offset to one or more motion vectors determined for a block of the encoded video frame; and outputting a reconstructed frame based on the output of the motion compensation for storage or display.
[0015] In some implementations of non-transitory computer-readable media, frame-level nonlinear motion offset is determined during the encoding of the current frame into an encoded video frame and transmitted as a signal within the encoded bitstream in the frame header associated with that encoded video frame.
[0016] In some implementations of non-transitory computer-readable media, frame-level nonlinear motion offsets are determined based on linear motion estimated for the current frame or block-level offsets determined for blocks of the current frame.
[0017] In some implementations of non-transitory computer-readable media, frame-level nonlinear motion offsets are added to a single motion vector predicted for an encoded video frame to perform motion compensation.
[0018] In some implementations of non-transitory computer-readable media, for each block in a block, a frame-level nonlinear motion offset is added to the block-level motion vector determined for that block.
[0019] According to an implementation of this disclosure, a system for decoding an encoded video frame using a frame-level nonlinear motion offset includes: one or more memories; and one or more processors configured to execute instructions stored in the one or more memories to: determine a frame-level nonlinear motion offset for the encoded video frame; apply the frame-level nonlinear motion offset to one or more motion vectors determined for blocks of the encoded video frame to predict the blocks; reconstruct the encoded video frame into a reconstructed frame based on the prediction of the blocks; and output the reconstructed frame.
[0020] In some implementations of this system, the frame-level nonlinear motion offset is determined during encoding the current frame into an encoded video frame based on the linear motion estimated for the current frame or the block-level offset determined for the blocks of the current frame.
[0021] In some implementations of this system, in order to apply frame-level nonlinear motion offsets to one or more motion vectors determined for blocks of encoded video frames to predict those blocks, one or more processors are configured to execute instructions to: add frame-level nonlinear motion offsets to a single motion vector predicted for the encoded video frame.
[0022] In some implementations of this system, in order to apply frame-level nonlinear motion offsets to one or more motion vectors determined for blocks of encoded video frames to predict those blocks, one or more processors are configured to execute instructions to: for each block in the block, add frame-level nonlinear motion offsets to the block-level motion vectors determined for that block.
[0023] These and other aspects of this disclosure are disclosed in the following detailed description of implementations, the appended claims, and the accompanying drawings. Attached Figure Description
[0024] The description herein refers to the accompanying drawings, in which the same reference numerals are used throughout the views to refer to the same parts.
[0025] Figure 1 This is a schematic diagram of an example video encoding and decoding system.
[0026] Figure 2 This is a block diagram of an example computing device that can implement a sending station or a receiving station.
[0027] Figure 3 This is a diagram of an example video stream that is to be encoded and decoded.
[0028] Figure 4 This is a block diagram of an example encoder.
[0029] Figure 5 This is a block diagram of an example decoder.
[0030] Figure 6 This is a diagram illustrating examples of different parts of a video frame.
[0031] Figure 7 This is a diagram of the linear motion estimated for bidirectional inter-frame prediction of the current frame and the frame-level nonlinear motion offset determined for the current frame.
[0032] Figure 8 This is a flowchart illustrating an example of a technique for encoding frames using frame-level nonlinear motion offsets.
[0033] Figure 9 This is a flowchart illustrating an example of a technique for decoding encoded frames using frame-level nonlinear motion offset. Detailed Implementation
[0034] Video compression schemes may include breaking down corresponding images or frames of a video stream into smaller parts, such as blocks, and generating an encoded bitstream by using encoding techniques to limit the information included in each block. The bitstream can be decoded to recreate the source frames from the limited information. Various techniques can be used to compress (i.e., encode) video streams to reduce the bandwidth required to send or store them. Similarly, various techniques can be used to decompress (i.e., decode) the compressed video stream from the bitstream to prepare it for viewing or further processing. Video stream compression typically utilizes the spatial and temporal correlations of the video signal through spatial and / or motion-compensated prediction. Motion-compensated prediction can also be referred to as inter-frame prediction. Inter-frame prediction uses one or more motion vectors to generate blocks (also called prediction blocks) that are similar to the current block encoded using previously encoded and decoded pixels. By encoding the motion vectors and the differences (i.e., residuals) between the two blocks, a decoder receiving the encoded signal can reconstruct the current block by generating the prediction blocks and adding the pixels of the prediction blocks to the decoded residual blocks.
[0035] Each motion vector used to generate a prediction block during inter-frame prediction references at least one reference frame (i.e., a frame other than the current frame containing the block being predicted). The reference frame can be located before or after the current frame in the sequence of the video stream and can be a frame reconstructed before being used as a reference frame. Specifically, the reference frame can be a forward reference frame (i.e., a frame used for forward prediction relative to the sequence) or a backward reference frame (i.e., a frame used for backward prediction relative to the sequence). One or more forward reference frames and / or backward reference frames can be used to encode or decode blocks. Specifically, because many conventional video compression and decompression schemes use pyramid code processing structures to achieve high compression efficiency, bidirectional prediction—such as using forward and backward reference frames—is used to encode and decode many frames. Bidirectional prediction using forward and backward reference frames has been shown to significantly improve prediction quality and thus significantly improve the overall compression performance of the target video stream.
[0036] Conventional videocode processing methods, utilizing bidirectional prediction, assume that the motion represented by the target motion vector is linear across the backward reference frame, the current frame, and the forward reference frame. This means that the motion of a given object is represented as having a constant velocity and direction across these frames. However, since motion is sometimes non-linear, such conventional methods may fail to accurately represent the motion. To address this issue, different methods have recently been proposed, where block-level motion offsets are used to more accurately represent block-level motion. These recent methods involve determining an offset for each individual block within which non-linear motion is determined and applying that offset to the linear motion vector. While using such block-level motion vector offsets can more accurately represent the motion of a given block, it also requires determining additional data for each applicable block and transmitting that additional data within the bitstream. Therefore, such block-level methods introduce increased computational and signal transmission costs, which can significantly increase the size of the bitstream. Furthermore, in many cases, the motion offset will be fixed across the target frame, meaning that the same motion offset applies to some or even all blocks of that frame. The implementation of this disclosure addresses these problems by using frame-level non-linear motion offsets for bidirectional prediction during encoding or decoding. Specifically, the implementation of this disclosure describes a method for determining a single nonlinear motion offset, which can be applied to motion vectors determined for some or all blocks of the current frame being encoded or decoded, thereby more accurately representing the nonlinear motion therein, while limiting additional computational and signal transmission overhead to a single offset value.
[0037] While this document references by way of example blocks, such as those commonly used in video codecs such as VP9, AV1, and the currently developing AV2, implementations of this disclosure can be used in conjunction with other video codec processing structures. In a particular, but non-limiting example, implementations of this disclosure can be used with CTUs, CUs, PUs, etc., commonly used in video codecs such as H.265, known as efficient video codec processing, and H.266, known as versatile video codec processing. Accordingly, references to specific video codec processing structures such as blocks herein should be considered as expressions of non-limiting example video codec processing structures that can be used in conjunction with implementations of this disclosure.
[0038] This paper initially refers to systems that can implement techniques for encoding or decoding using frame-level nonlinear motion offsets to describe further details of such techniques. Figure 1 This is a schematic diagram of an example of a video encoding and decoding system 100. The transmitting station 102 can be, for example, such as... Figure 2 The computer described has an internal hardware configuration. However, other implementations of the sending station 102 are possible. For example, the processing of the sending station 102 can be distributed among multiple devices.
[0039] Network 104 can connect sending station 102 and receiving station 106 for encoding and decoding of video streams. Specifically, the video stream can be encoded in sending station 102 and decoded in receiving station 106. Network 104 can be, for example, the Internet. Network 104 can also be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network, or any other component that transmits video streams from sending station 102 to (in this example) receiving station 106.
[0040] In one example, receiving station 106 could be such as Figure 2 The described computer has an internal hardware configuration. However, other suitable implementations of the receiving station 106 are possible. For example, the processing of the receiving station 106 can be distributed among multiple devices.
[0041] Other implementations of the video encoding and decoding system 100 are possible. For example, network 104 may be omitted in the implementation. In another implementation, the video stream may be encoded and then stored for later transmission to receiving station 106 or any other device with memory. In one implementation, receiving station 106 receives (e.g., via network 104, a computer bus, and / or some communication path) the encoded video stream and stores it for later decoding. In an example implementation, Real-Time Transport Protocol (RTP) is used for the transmission of the encoded video over network 104. In another implementation, a transport protocol other than RTP may be used (e.g., a video streaming protocol based on Hypertext Transfer Protocol (HTTP)).
[0042] When used in a video conferencing system, for example, sending station 102 and / or receiving station 106 may include the ability to both encode and decode video streams as described below. For example, receiving station 106 may be a video conferencing participant who receives an encoded video bitstream from a video conferencing server (e.g., sending station 102) for decoding and viewing, and further encodes his or her own video bitstream and sends it to the video conferencing server for other participants to decode and view.
[0043] In some implementations, the video encoding and decoding system 100 can alternatively be used to encode and decode data other than video data. For example, the video encoding and decoding system 100 can be used to process image data. Image data may include data blocks from an image. In such an implementation, the transmitting station 102 can be used to encode the image data, and the receiving station 106 can be used to decode the image data.
[0044] Alternatively, receiving station 106 may refer to a computing device, such as storing encoded image data for later use after receiving encoded or pre-encoded image data from transmitting station 102. As a further alternative, transmitting station 102 may refer to a computing device, such as decoding image data before sending decoded image data to receiving station 106 for display.
[0045] Figure 2 This is a block diagram illustrating an example of a computing device 200 that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement... Figure 1 One or both of the transmitting station 102 and the receiving station 106. The computing device 200 may be in the form of a computing system including multiple computing devices, or in the form of a single computing device (e.g., a mobile phone, tablet computer, laptop computer, notebook computer, desktop computer, etc.).
[0046] The processor 202 in the computing device 200 can be a conventional central processing unit. Alternatively, the processor 202 can be another type of device or multiple devices that exist now or are developed later and are capable of manipulating or processing information. For example, although the disclosed implementation can be practiced with a single processor (e.g., processor 202) shown, advantages in speed and efficiency can be achieved by using more than one processor.
[0047] In an implementation, the memory 204 in the computing device 200 may be a read-only memory (ROM) device or a random access memory (RAM) device. However, other suitable types of storage devices may be used as memory 204. Memory 204 may include code and data 206 accessed by the processor 202 using bus 212. Memory 204 may further include an operating system 208 and an application program 210, which includes at least one program that permits the processor 202 to execute the methods described herein. For example, application program 210 may include applications 1 to N, which further include encoding and / or decoding software that performs encoding or decoding using frame-level nonlinear motion offsets as described herein.
[0048] The computing device 200 may also include auxiliary storage 214, which may be, for example, a memory card used with a mobile computing device. Because video communication sessions may contain a considerable amount of information, it may be stored, in whole or in part, in the auxiliary storage 214 and loaded into the memory 204 for processing as needed.
[0049] The computing device 200 may also include one or more output devices, such as a display 218. In one example, the display 218 may be a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. The display 218 may be coupled to the processor 202 via a bus 212. In addition to or as an alternative to the display 218, other output devices may be provided that allow a user to program or otherwise use the computing device 200. When the output device is a display or includes a display, the display may be implemented in various ways, including via a liquid crystal display (LCD), a cathode ray tube (CRT) display, or a light-emitting diode (LED) display (such as an organic LED (OLED) display).
[0050] The computing device 200 may also include an image sensing device 220, such as a camera or any other existing or later-developed image sensing device 220 capable of sensing images (such as images of a user operating the computing device 200), or communicating with such image sensing device. The image sensing device 220 may be positioned such that it is pointed toward the user operating the computing device 200. In an example, the position and optical axis of the image sensing device 220 may be configured such that the field of view includes an area directly adjacent to and visible from the display 218.
[0051] The computing device 200 may also include a sound sensing device 222, such as a microphone or any other sound sensing device that can sense the present or future presence of sound in the vicinity of the computing device 200, or communicate with such sound sensing device. The sound sensing device 222 may be positioned such that it is directed toward a user operating the computing device 200, and may be configured to receive sounds, such as speech or other words, emitted by the user when the user operates the computing device 200.
[0052] although Figure 2 The processor 202 and memory 204 of computing device 200 are depicted as integrated into a single unit, but other configurations may be utilized. The operation of processor 202 can be distributed across multiple machines (where individual machines may have one or more processors), which may be directly coupled or coupled across a local area network or other network. Memory 204 can be distributed across multiple machines, such as network-based memory or memory across multiple machines performing the operations of computing device 200.
[0053] Although described here as a single bus, the bus 212 of the computing device 200 can consist of multiple buses. Furthermore, the auxiliary storage 214 can be directly coupled to other components of the computing device 200 or can be accessed via a network, and can include integrated units (such as memory cards) or multiple units (such as multiple memory cards). Therefore, the computing device 200 can be implemented in a wide variety of configurations.
[0054] Figure 3 This is a diagram illustrating an example of a video stream 300 to be encoded and decoded. The video stream 300 includes a video sequence 302. At the next level, the video sequence 302 includes multiple adjacent video frames 304. Although three frames are depicted as adjacent frames 304, the video sequence 302 can include any number of adjacent frames 304. Adjacent frames 304 can then be further subdivided into individual video frames, such as frame 306.
[0055] At the next level, frame 306 can be divided into a series of planes or segments 308. For example, segment 308 can be a subset of frames that allow for parallel processing. Segment 308 can also be a subset of frames that can separate video data into individual colors. For example, frame 306 of color video data can include a luminance plane and two chrominance planes. Segment 308 can be sampled at different resolutions.
[0056] Regardless of whether frame 306 is divided into segments 308, frame 306 can be further subdivided into blocks 310, which can contain data corresponding to, for example, N×M pixels in frame 306, where N and M can refer to the same integer value or different integer values. Block 310 can also be arranged to include data from one or more segments 308 of pixel data. Block 310 can be any suitable size, such as 4×4 pixels, 8×8 pixels, 16×8 pixels, 8×16 pixels, 16×16 pixels, or larger up to a maximum block size, which can be 128×128 pixels or another N×M pixel size.
[0057] Figure 4 This is a block diagram of an example encoder 400. As described above, encoder 400 can be implemented in transmitting station 102, such as by providing a computer software program stored in memory (e.g., memory 204). The computer software program may include machine instructions that, when executed by a processor such as processor 202, cause transmitting station 102 to... Figure 4 The video data is encoded in the manner described herein. The encoder 400 can also be implemented as dedicated hardware included, for example, in the transmitting station 102. In some implementations, the encoder 400 is a hardware encoder.
[0058] The encoder 400 has the following stages for performing various functions in the forward path (shown by solid connecting lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: intra / inter-frame prediction stage 402, transform stage 404, quantization stage 406, and entropy coding stage 408. The encoder 400 may also include a reconstruction path (shown by dashed connecting lines) for reconstructing frames for encoding future blocks. Figure 4 In the encoder 400, the following stages are used to perform various functions in the reconstruction path: dequantization stage 410, inverse transform stage 412, reconstruction stage 414, and loop filtering stage 416. Other structural variations of the encoder 400 can be used to encode the video stream 300.
[0059] In some cases, the functions performed by encoder 400 may occur after filtering of video stream 300. That is, before encoder 400 receives video stream 300, video stream 300 may undergo preprocessing according to one or more implementations of this disclosure. Alternatively, encoder 400 itself may continue to perform functions related to... Figure 4 Such preprocessing is performed on video stream 300 before the described functions—such as before processing video stream 300 at intra / inter-frame prediction stage 402.
[0060] When the video stream 300 is submitted for encoding after preprocessing, the corresponding adjacent frames 304, such as frame 306, can be processed in blocks. At the intra-frame / inter-frame prediction stage 402, the corresponding blocks can be encoded using either intra-frame prediction (also known as intra-prediction) or inter-frame prediction (also known as inter-prediction). In either case, prediction blocks can be formed. In the case of intra-frame prediction, prediction blocks can be formed from samples that have already been encoded and reconstructed in the current frame. In the case of inter-frame prediction, prediction blocks can be formed from samples in one or more previously constructed reference frames.
[0061] Next, the predicted block can be subtracted from the current block at the intra / inter-frame prediction stage 402 to produce a residual block (also known as the residual). The transform stage 404 uses a block-based transform to transform the residual into transform coefficients, for example, in the frequency domain. The quantization stage 406 uses a quantizer value or quantization level to convert the transform coefficients into discrete quantum values, referred to as quantized transform coefficients. For example, the transform coefficients can be divided by the quantizer value and truncated.
[0062] The quantized transform coefficients are then entropy encoded by entropy coding stage 408. The entropy-coded coefficients, along with other information for decoding the block (which may include, for example, syntax elements indicating the prediction type, transform type, motion vector, quantizer value, etc.), are then output to the compressed bitstream 420. The compressed bitstream 420 can be formatted using various techniques such as variable-length code processing or arithmetic code processing. The compressed bitstream 420 may also be referred to as an encoded video stream or an encoded video bitstream, and the terms will be used interchangeably herein.
[0063] The reconstruction path (shown by the dashed connection line) can be used to ensure encoder 400 and (see below for details) Figure 5 The decoder 500 (described below) uses the same reference frame to decode the compressed bitstream 420. The reconstruction path is performed in accordance with (see below regarding...)Figure 5 The functions described are similar to those that occur during the decoding process, including dequantizing the quantized transform coefficients at the dequantization stage 410 and performing an inverse transform on the dequantized transform coefficients at the inverse transform stage 412 to produce the derived residual block (also referred to as the derived residual).
[0064] At reconstruction stage 414, the predicted blocks already predicted at intra / inter-frame prediction stage 402 can be added to the derived residuals to create reconstructed blocks. Loop filtering stage 416 can apply intra-loop filters or other filters to the reconstructed blocks to reduce distortion, such as blocking artifacts. Examples of filters that can be applied at loop filtering stage 416 include, but are not limited to, deblocking filters, direction enhancement filters, and loop recovery filters.
[0065] Other variations of encoder 400 can be used to encode the compressed bitstream 420. In some implementations, for certain blocks or frames, a non-transform-based encoder can directly quantize the residual signal without the transform stage 404. In some implementations, the encoder may have a quantization stage 406 and a dequantization stage 410 combined in a common stage.
[0066] Figure 5 This is a block diagram of an example decoder 500. Decoder 500 can be implemented in receiving station 106, for example, by providing a computer software program stored in memory 204. The computer software program may include machine instructions that, when executed by a processor such as processor 202, cause receiving station 106 to... Figure 5 The video data is decoded in the manner described herein. The decoder 500 can also be implemented in hardware included in, for example, transmitting station 102 or receiving station 106. In some implementations, the decoder 500 is a hardware decoder.
[0067] Similar to the reconstruction path of encoder 400 discussed above, in one example, decoder 500 includes the following stages for performing various functions to produce output video stream 516 from compressed bitstream 420: entropy decoding stage 502, dequantization stage 504, inverse transform stage 506, intra / inter-frame prediction stage 508, reconstruction stage 510, loop filtering stage 512, and post-filtering stage 514. Other structural variations of decoder 500 can be used to decode compressed bitstream 420.
[0068] When the compressed bitstream 420 is presented for decoding, the data elements within the compressed bitstream 420 can be decoded by the entropy decoding stage 502 to produce a quantized set of transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by a quantizer value), and the inverse transform stage 506 performs an inverse transform on the dequantized transform coefficients to produce a derived residual, which can be the same as the derived residual created by the inverse transform stage 412 in the encoder 400. Using the header information decoded from the compressed bitstream 420, the decoder 500 can use the intra / inter-frame prediction stage 508 to create a prediction block identical to the prediction block created in the encoder 400 (e.g., at the intra / inter-frame prediction stage 402).
[0069] At reconstruction stage 510, predicted blocks can be added to the derived residuals to create reconstructed blocks. Loop filtering stage 512 can be applied to the reconstructed blocks to reduce blocking artifacts. Examples of filters that can be applied at loop filtering stage 512 include, but are not limited to, deblocking filters, directional enhancement filters, and loop recovery filters. Other filters can be applied to the reconstructed blocks. In this example, post-filtering stage 514 is applied to the reconstructed blocks to reduce blocking distortion, and the result is output as output video stream 516. Output video stream 516 can also be referred to as decoded video stream, and the terms will be used interchangeably herein.
[0070] Other variations of decoder 500 can be used to decode the compressed bitstream 420. In some implementations, decoder 500 may produce output video stream 516 without post-filtering stage 514, or otherwise omit post-filtering stage 514.
[0071] Figure 6 This is an illustration of examples of portions of video frame 600, which may be, for example, [the video frame could be...]. Figure 3Frame 306 is shown. Video frame 600 includes multiple 64×64 blocks 610, such as four 64×64 blocks 610 in a matrix or Cartesian plane in two rows and two columns, as shown. Each 64×64 block 610 may include up to four 32×32 blocks 620. Each 32×32 block 620 may include up to four 16×16 blocks 630. Each 16×16 block 630 may include up to four 8×8 blocks 640. Each 8×8 block 640 may include up to four 4×4 blocks 950. Each 4×4 block 950 may include 16 pixels, which may be represented in four rows and four columns in each corresponding block in the Cartesian plane or matrix. In some implementations, video frame 600 may include blocks larger than 64×64 and / or blocks smaller than 4×4. Video frame 600 may be partitioned into various block arrangements according to features and / or other criteria within video frame 600.
[0072] Pixels may include information representing an image captured in video frame 600, such as luminance information, color information, and positional information. In some implementations, a block of 16×16 pixels, such as the one shown, may include: a luminance block 660, which may include luminance pixels 662; and two chrominance blocks 670 and 680, such as a U or Cb chrominance block 670 and a V or Cr chrominance block 680. Chroma blocks 670 and 680 may include chrominance pixels 690. For example, luminance block 660 may include 16×16 luminance pixels 662, and each chrominance block 670 and 680 may include 8×8 chrominance pixels 690, as shown. Although one arrangement of blocks is shown, any arrangement can be used. Figure 6 An N×N block is shown, but in some implementations, an N×M block can be used, where N and M are different numbers. For example, 32×64 blocks, 64×32 blocks, 16×32 blocks, 32×16 blocks, or any other block size can be used. In some implementations, N×2N blocks, 2N×N blocks, or combinations thereof can be used.
[0073] In some implementations, coding of video frame 600 may include ordered block-level coding. Ordered block-level coding may include coding blocks of video frame 600 in an order such as raster scan order, wherein blocks may be identified and processed starting with the top-left block of video frame 600 or a portion of video frame 600, and proceeding along rows from left to right and from top to bottom, thereby sequentially identifying and processing each block. For example, a 64×64 block in the top left column of video frame 600 may be the first coded block, and a 64×64 block immediately to the right of the first block may be the second coded block. The second row from the top may be the second coded row, such that a 64×64 block in the left column of the second row may be coded after the 64×64 block in the rightmost column of the first row.
[0074] In some implementations, encoding the blocks of video frame 600 may include using quadtree encoding, which may involve encoding smaller block units within a block in raster scan order. For example, quadtree encoding may be used to encode the 64×64 blocks shown in the lower left corner of this portion of video frame 600, where the upper left 32×32 blocks may be encoded, then the upper right 32×32 blocks, then the lower left 32×32 blocks, and then the lower right 32×32 blocks. Quadtree encoding may be used to encode each 32×32 block, where the upper left 16×16 blocks may be encoded, then the upper right 16×16 blocks, then the lower left 16×16 blocks, and then the lower right 16×16 blocks. Quadtree coding can be used to code each 16×16 block, where the top-left 8×8 block can be coded, followed by the top-right 8×8 block, then the bottom-left 8×8 block, and finally the bottom-right 8×8 block. Alternatively, quadtree coding can be used to code each 8×8 block, where the top-left 4×4 block can be coded, followed by the top-right 4×4 block, then the bottom-left 4×4 block, and finally the bottom-right 4×4 block. In some implementations, the 8×8 blocks can be omitted from the 16×16 block, and quadtree coding can be used to code the 16×16 block, where the top-left 4×4 block can be coded, and then the other 4×4 blocks in the 16×16 block can be coded according to the raster scan order.
[0075] In some implementations, encoding video frame 600 may include encoding information included in the original version of the image or video frame by, for example, omitting some information from the original version of the image or video frame. For example, encoding may include reducing spectral redundancy, reducing spatial redundancy, or a combination thereof. Reducing spectral redundancy may include using a color model based on a luminance component (Y) and two chrominance components (U and V or Cb and Cr), which may be referred to as the YUV or YCbCr color model or color space. Using the YUV color model may include using a relatively large amount of information to represent the luminance component of a portion of video frame 600, and using a relatively small amount of information to represent each corresponding chrominance component of that portion of video frame 600. For example, a portion of video frame 600 may be represented by a high-resolution luminance component that may include a 16×16 pixel block and two lower-resolution chrominance components, where each chrominance component represents that portion of the image as an 8×8 pixel block. A pixel can indicate a value, for example, a value in the range of 0 to 255, and can be stored or transmitted using, for example, eight bits. Although this disclosure is described with reference to the YUV color model, another color model can be used. Reducing spatial redundancy can include transforming the block into the frequency domain using, for example, a discrete cosine transform. For example, units of an encoder can perform a discrete cosine transform using transform coefficient values based on spatial frequencies.
[0076] Although this document describes video frame 600 with reference to a matrix or Cartesian representation for clarity, video frame 600 can be stored, transmitted, processed, or combinations thereof in data structures that allow for efficient representation of pixel values for video frame 600. For example, video frame 600 can be stored, transmitted, processed, or any combination thereof in a two-dimensional data structure such as a matrix as shown, or in a one-dimensional data structure such as a vector array. Furthermore, although described herein as showing a chroma-subsampled image where U and V have half the resolution of Y, video frame 600 can have different configurations of its color channels. For example, still referring to the YUV color space, full resolution can be used for all color channels of video frame 600. In another example, the resolution of the color channels of video frame 600 can be represented using a color space other than the YUV color space.
[0077] Figure 7This is a diagram illustrating the estimated linear motion for bidirectional inter-frame prediction of the current frame 700 and the frame-level nonlinear motion offset determined for the current frame 700. The current frame 700 is in prediction during encoding (e.g., in intra / inter-frame prediction phase 402) or decoding (e.g., in intra / inter-frame prediction phase 510). Specifically, one or more blocks of the current frame 700 are being inter-frame predicted using backward reference frame 702 and forward reference frame 704. The estimated linear motion 706, shown in dashed lines, represents the linear motion of an object (shown by a black circle) estimated from backward reference frame 702, through the current frame 700, to forward reference frame 704. For example, the estimated linear motion 706 can be determined by estimating the motion field of a block in the current frame 700 by projecting existing motion vectors. In another example, the estimated linear motion 706 may correspond to a linear motion vector predictor 708 determined based on a motion search performed with respect to the current frame 700 (e.g., using reconstructed motion vectors from backward reference frame 702 and forward reference frame 704). The estimated linear motion 706 is fixed across the current frame 700.
[0078] During motion compensation performed as part of encoding or decoding the current frame 700, the encoder or decoder (as appropriate) determines the actual motion 710 of the object, represented as a curve from the backward reference frame 702 through the current frame 700 to the forward reference frame 704. Therefore, the actual motion 710 determined by the encoder or decoder via the motion compensation process is non-linear. As a result, the estimated position of the object within the current frame 700 based on the estimated linear motion 706 is inaccurate; instead, the actual position of the object within the current frame 700 is shown at the black circle drawn along the actual motion 710. The difference between the position of the object based on the estimated linear motion 706 and the position of the object along the actual motion 710 is processed as a motion vector offset 712, which represents a distance metric applied to predict the motion of the object within the current frame 700. Specifically, the motion vector offset 712 is a frame-level non-linear motion offset for some or all blocks of the current frame 700. Therefore, instead of determining and using individual block-level motion vector offsets (e.g., by adding to existing block-level motion vectors) for different blocks of the current frame 700, inter-frame prediction of the current frame 700 involves applying the same motion vector offset 712 to some or all of its blocks.
[0079] Further details are now described regarding techniques used for encoding or decoding using frame-level nonlinear motion offsets. Figure 8 This is a flowchart illustrating an example of a technique 800 for encoding frames using frame-level nonlinear motion offset. Figure 9This is a flowchart illustrating an example of technique 900 for decoding encoded frames using frame-level nonlinear motion offset. For example, technique 800 may be performed wholly or partially at the prediction phase of an encoder used to encode a video stream (e.g., intra / inter-frame prediction phase 402), while technique 900 may be performed wholly or partially at the prediction phase of a decoder used to decode a bitstream (e.g., intra / inter-frame prediction phase 508).
[0080] Techniques 800 and / or 900 can be implemented as software programs, for example, executable by a computing device such as transmitting station 102 or receiving station 106. For example, the software program may include machine-readable instructions that can be stored in memory such as memory 204 or auxiliary storage 214, and when executed by a processor such as processor 202, can cause the computing device to execute techniques 800 and / or 900. Techniques 800 and / or 900 can be implemented using dedicated hardware or firmware. For example, hardware components such as hardware code processors can be configured to execute techniques 800 and / or 900. As explained above, some computing devices may have multiple memories or processors, and multiple processors, memories, or both can be used to distribute the operations described in techniques 800 and / or 900. For ease of explanation, techniques 800 and / or 900 are each depicted and described herein as a series of steps or operations. However, the steps or operations according to this disclosure may occur in various orders and / or concurrently. Additionally, other steps or operations not presented and described herein may be used. Furthermore, it may not be necessary to implement all the steps or operations shown to achieve the technology according to the disclosed subject matter.
[0081] First refer to Figure 8 The diagram illustrates a technique 800 for encoding frames using frame-level nonlinear motion offsets. At 802, motion compensation is performed on blocks of a video frame to determine a frame-level nonlinear motion offset for the video frame, which may be, for example, the current frame 700. The frame-level nonlinear motion offset—for example, a motion vector offset 712—is based on a linear motion estimated between the video frame and each of the backward and forward reference frames, such as the estimated linear motion 706.
[0082] Determining frame-level nonlinear motion offsets can include determining the nonlinear motion of a video frame based on the position of an object associated with the nonlinear motion within that frame. For example, determining nonlinear motion can include performing a frame-level motion search over the entire video frame to identify individual motion vectors predicted for the entire frame (i.e., the predicted motion vectors). The frame-level nonlinear motion offset can then be determined based on the difference between the actual motion of the video frame, determined based on motion compensation, and the predicted motion vectors. In another example, determining nonlinear motion can include performing a block-level motion search over individual blocks of the video frame to identify block-level offsets. These block-level offsets can then be processed to determine the frame-level nonlinear motion offset. For example, a simple mean, weighted mean, or mode of the block-level offsets can be calculated and used as the frame-level nonlinear motion offset. In some such cases, the simple mean, weighted mean, or mode may only consider those blocks in which the target object resides.
[0083] At 804, blocks are predicted by applying a frame-level nonlinear motion offset to the motion vectors determined for the blocks of the video frame. The motion vectors determined for the blocks are derived from the output of motion compensation performed on the video frame. Applying the frame-level nonlinear motion offset to the motion vectors involves adding that frame-level nonlinear motion offset as a motion vector offset value to each such motion vector. The prediction residuals can then be determined based on the offset-adjusted motion vectors of the blocks of the video frame.
[0084] In some implementations, the blocks used to predict video frames by applying frame-level nonlinear motion offsets to motion vectors may alternatively include a determined temporally interpolated picture (TIP) reference frame. The TIP reference frame is generated by interpolating reference blocks from forward and backward reference frames. This TIP reference frame can be generated based on a motion field determined according to the frame-level nonlinear motion offsets, for example, by applying the frame-level nonlinear motion offsets to an initial motion field determined for the video frame. The updated motion field resulting from this process can then be used to generate the TIP reference frame, which can then be used as a spatially and temporally co-located reference frame for predicting the video frame during inter-frame prediction.
[0085] At 806, the frame-level nonlinear motion offset is encoded into the bitstream. For example, the frame-level nonlinear motion offset can be encoded into the frame header of a video frame within the bitstream. In some implementations, the offset-adjusted motion vector of the video frame, obtained by applying the frame-level nonlinear motion offset to the predicted motion vector of the video frame, can be transmitted in the bitstream as a signal instead of the frame-level nonlinear motion offset. This allows the decoder to use the offset-adjusted motion vector without further decoder-side computation.
[0086] Next reference Figure 9This illustrates a technique 900 for decoding encoded frames using frame-level nonlinear motion offsets. At 902, a frame-level nonlinear motion offset for the encoded video frame is decoded from the bitstream to which the encoded video frame is encoded, which may be, for example, the current frame 700. For example, the frame-level nonlinear motion offset can be decoded from the frame header of the encoded frame within the bitstream.
[0087] At position 904, motion compensation is performed on blocks of the encoded video frame. Performing motion compensation on blocks of the encoded video frame involves applying a frame-level nonlinear motion offset to the motion vector determined for the block. Applying the frame-level nonlinear motion offset to the motion vector involves adding the frame-level nonlinear motion offset as a motion vector offset value to the motion vector predicted for the entire encoded frame or to each individual motion vector determined for a block of the encoded frame. For example, the motion vector can be decoded from the bitstream or computed at the decoder performing the motion compensation. The predicted block can then be determined based on the offset-adjusted motion vector of the block of the video frame.
[0088] In some implementations, similar to technique 800, predicting a video frame by applying a frame-level nonlinear motion offset to a motion vector may alternatively include determining a TIP reference frame. This TIP reference frame may be generated based on a motion field determined according to the frame-level nonlinear motion offset, for example, by applying the frame-level nonlinear motion offset to an initial motion field determined for the video frame. The updated motion field generated by this process can then be used to generate the TIP reference frame, which can then be used as a spatial and temporal co-location reference frame for predicting the video frame during inter-frame prediction.
[0089] At 906, the encoded video frame is reconstructed into a reconstructed frame based on the motion-compensated output. For example, residual data corresponding to individual blocks of the encoded frame can be reconstructed from the motion-compensated output to determine the reconstructed blocks, and those reconstructed blocks can then be combined to produce the reconstructed frame.
[0090] At position 908, the reconstructed frame is output within the output video stream. For example, the reconstructed frame may be combined with other reconstructed frames generated based on and / or independently of other frame-level nonlinear motion offsets as the output video stream to represent a decoded version of the initial video stream encoded into the bitstream. The output video stream may be output, for example, for storage or display.
[0091] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it should be understood that when those terms are used in the claims, encoding and decoding may mean compressing data, decompressing data, transforming data, or any other processing or alteration of data.
[0092] The word “example” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as an “example” is not necessarily to be construed as preferred or advantageous over other aspects or designs. Rather, the use of the word “example” is intended to present a concept in a specific manner. As used in this application, the term “or” is intended to mean inclusive “or” rather than exclusive “or.” That is, unless otherwise specified or clearly indicated in the context, the statement “X comprises A or B” is intended to mean either of its natural inclusive arrangements. That is, if X comprises A; X comprises B; or X comprises both A and B, then “X comprises A or B” is satisfied under any of the above examples. Additionally, the articles “a” and “an” as used in this application and the appended claims should generally be interpreted as meaning “one or more” unless otherwise specified or clearly indicated in the context for the singular form. Furthermore, the use of the terms “implementation” or “an implementation” throughout this disclosure is not intended to refer to the same implementation unless so described.
[0093] The implementation of transmitting station 102 and / or receiving station 106 (and the algorithms, methods, instructions, etc. stored thereon and / or executed thereon (including by encoder 400 and decoder 500 or another encoder or decoder as disclosed herein) can be implemented in hardware, software, or any combination thereof. Hardware may include, for example, a computer, intellectual property (IP) core, application-specific integrated circuit (ASIC), programmable logic array, optical processor, programmable logic controller, microcode, microcontroller, server, microprocessor, digital signal processor, or any other suitable circuitry. In the claims, the term "processor" should be understood to cover any of the foregoing hardware individually or in combination. The terms "signal" and "data" are used interchangeably. Furthermore, portions of transmitting station 102 and receiving station 106 do not necessarily have to be implemented in the same manner.
[0094] Furthermore, in one aspect, for example, transmitting station 102 or receiving station 106 may be implemented using a general-purpose computer or general-purpose processor having a computer program that, when executed, performs any of the corresponding methods, algorithms, and / or instructions described herein. Alternatively or alternatively, for example, a special-purpose computer / processor may be utilized, which may include additional hardware for performing any of the methods, algorithms, or instructions described herein.
[0095] Sending station 102 and receiving station 106 can be implemented, for example, on a computer in a video conferencing system. Alternatively, sending station 102 can be implemented on a server, and receiving station 106 can be implemented on a device separate from the server (such as a handheld communication device). In this example, sending station 102 can encode content into an encoded video signal and send the encoded video signal to the communication device. The communication device can then decode the encoded video signal. Alternatively, the communication device can decode content stored locally on the communication device (e.g., content not sent by sending station 102). Other suitable sending and receiving implementations are available. For example, receiving station 106 can be a generally fixed personal computer instead of a portable communication device.
[0096] Furthermore, all or part of the implementations of this disclosure may take the form of a computer program product accessible from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium may be any means capable of, for example, tangibly containing, storing, transmitting, or transporting a program for use by or in conjunction with any processor. The medium may be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media are also available.
[0097] The above implementations and other aspects have been described to facilitate easy understanding of this disclosure and are not intended to limit it. Rather, this disclosure is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, and this scope should be given the broadest interpretation permitted under law to cover all such modifications and equivalent arrangements.
Claims
1. A method for decoding encoded video frames using frame-level nonlinear motion offset, the method comprising: Decode the frame-level nonlinear motion offset of the encoded video frame from the bitstream to which the encoded video frame is encoded; Performing motion compensation on blocks of the encoded video frame includes applying the frame-level nonlinear motion offset to one or more motion vectors determined for the block; Based on the output of the motion compensation, the encoded video frame is reconstructed into a reconstructed frame; as well as The reconstructed frame is output within the output video stream.
2. The method as described in claim 1, wherein, The frame-level nonlinear motion offset is determined during the encoding of the current frame into the encoded video frame.
3. The method as described in claim 2, wherein, The frame-level nonlinear motion offset is determined based on the linear motion estimated between each of the current frame and its backward reference frame and forward reference frame.
4. The method of claim 3, wherein, The linear motion is represented by a single motion vector predicted for the current frame, and the frame-level nonlinear motion offset is determined based on the difference between the nonlinear motion of the encoded video frame and the single motion vector.
5. The method of claim 4, wherein, The nonlinear motion of the encoded video frame is determined based on the position of the object associated with the nonlinear motion within the current frame.
6. The method of claim 2, wherein, The frame-level nonlinear motion offset is determined based on block-level offsets determined for multiple blocks of the current frame.
7. The method of claim 6, wherein, The frame-level nonlinear motion offset is one of the simple average, weighted average, or mode of the block-level offset.
8. The method of claim 2, wherein, Decoding the frame-level nonlinear motion offset of the encoded video frame includes: The frame-level nonlinear motion offset is decoded from the frame header associated with the encoded video frame within the bitstream.
9. The method according to any one of claims 1 to 8, wherein, Performing the motion compensation on the block of the encoded video frame includes applying the frame-level nonlinear motion offset to the one or more motion vectors determined for the block, including: The frame-level nonlinear motion offset is added to a single motion vector predicted for all coded video frames in the coded video frame.
10. The method according to any one of claims 1 to 8, wherein, Performing the motion compensation on the block of the encoded video frame includes applying the frame-level nonlinear motion offset to the one or more motion vectors determined for the block, including: For each block in the block, the frame-level nonlinear motion offset is added to the block-level motion vector determined for that block.
11. The method according to any one of claims 1 to 8, wherein, Performing the motion compensation on the block of the encoded video frame includes applying the frame-level nonlinear motion offset to the one or more motion vectors determined for the block, including: The updated motion field of the encoded video frame is determined by applying the frame-level nonlinear motion offset to the initial motion field; Determine the temporal interpolation image reference frame based on the updated motion field; and The time-interpolated image reference frame is used to predict the encoded video frame.
12. A non-transitory computer-readable medium storing an encoded bit stream, wherein, The encoded bitstream is configured for decoding by including the following operations: Motion compensation is performed on blocks of the encoded video frame by applying frame-level nonlinear motion offsets to one or more motion vectors determined for blocks of the encoded video frame; The reconstructed frame, generated based on the motion compensation output, is output for storage or display.
13. The non-transitory computer-readable medium of claim 12, wherein, The frame-level nonlinear motion offset is determined during the encoding of the current frame into the encoded video frame and is transmitted as a signal within the encoded bitstream in the frame header associated with the encoded video frame.
14. The non-transitory computer-readable medium of claim 13, wherein, The frame-level nonlinear motion offset is determined based on the linear motion estimated for the current frame or the block-level offset determined for the blocks of the current frame.
15. The non-transitory computer-readable medium as claimed in any one of claims 12 to 14, wherein, The frame-level nonlinear motion offset is added to a single motion vector predicted for the encoded video frame to perform the motion compensation.
16. The non-transitory computer-readable medium as claimed in any one of claims 12 to 14, wherein, For each block in the block, the frame-level nonlinear motion offset is added to the block-level motion vector determined for that block.
17. A system for decoding encoded video frames using frame-level nonlinear motion offset, the system comprising: One or more memory units; as well as One or more processors, the one or more processors being configured to execute instructions stored in the one or more memories to: Determine frame-level nonlinear motion offsets for the encoded video frames; The frame-level nonlinear motion offset is applied to one or more motion vectors determined for a block of the encoded video frame to predict the block; Based on the prediction of the block, the encoded video frame is reconstructed into a reconstructed frame; as well as Output the reconstructed frame.
18. The system of claim 17, wherein, The frame-level nonlinear motion offset is determined during encoding the current frame into the encoded video frame based on either the linear motion estimated for the current frame or the block-level offset determined for the blocks of the current frame.
19. The system as claimed in any one of claims 17 to 18, wherein, In order to apply the frame-level nonlinear motion offset to the one or more motion vectors determined for the block of the encoded video frame to predict the block, the one or more processors are configured to execute the instructions to: The frame-level nonlinear motion offset is added to a single motion vector predicted for the encoded video frame.
20. The system as claimed in any one of claims 17 to 18, wherein, In order to apply the frame-level nonlinear motion offset to the one or more motion vectors determined for the block of the encoded video frame to predict the block, the one or more processors are configured to execute the instructions to: For each block in the block, the frame-level nonlinear motion offset is added to the block-level motion vector determined for that block.