Video coding and decoding method and device, computer equipment and storage medium

The overlapping optical flow correction technology with adaptive sub-block size selection solves the problem of inaccurate motion vector correction in the prior art and improves the accuracy of video encoding and decoding and the compression effect.

CN120676165APending Publication Date: 2025-09-19TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510195735.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-14
Filing Date
2025-02-21
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In existing video coding and decoding technologies, the sub-block-based motion vector correction method has a small number of samples, resulting in inaccurate motion vector correction, which in turn affects the accuracy of coding and decoding.

Method used

The motion vector correction technology based on overlapping optical flow is adopted to improve the accuracy of motion vector by adaptively selecting sub-block size for optical flow correction.

Benefits of technology

It improves the accuracy of video encoding and decoding, reduces encoding and decoding errors, and improves video compression effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676165A_ABST
    Figure CN120676165A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video encoding and decoding method and device, computer equipment and a storage medium. The video decoding method includes receiving a video bitstream including a plurality of blocks. The method further includes deriving, for a current block of the plurality of blocks, a set of sub-block motion vectors, and deriving a set of modified sub-block motion vectors for the current block by applying optical flow modifications to the set of sub-blocks of the current block. Wherein respective sizes of the set of sub-blocks are adaptively selected for optical flow correction. The method further includes reconstructing the current block using the set of modified sub-block motion vectors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 566,810, filed on March 18, 2024, entitled “Overlapped Optical Flow-Based Motion Vector Refinement with Adaptive Subblock Size,” and U.S. Patent Application No. 18 / 805,299, filed on August 14, 2024, the contents of which are incorporated by reference in their entirety into this application. Technical Field

[0003] The present application relates to video coding and decoding technology, and in particular to a video coding and decoding method and apparatus, computer equipment, and storage medium. Background Art

[0004] Various electronic devices support digital video, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, and video streaming devices. Electronic devices transmit and receive digital video data via a communication network or otherwise transfer digital video data, and / or store digital video data on a storage device. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video data may be compressed using video coding according to at least one video coding standard before being transmitted or stored. Video coding can be performed by hardware and / or software on the electronic device / client device or on a server providing cloud services.

[0005] Video coding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) to exploit the inherent redundancy in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation in video quality. Various video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended to be the successor to HEVC. ITU-T and ISO / IEC published the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). The Alliance for Open Media Video 1 (AV1) is an open video coding format designed as a successor to HEVC. On January 8, 2019, the verified version 1.0.0 of the specification and Errata 1 were released.

[0006] Bilateral matching can be used for motion vector correction. In decoder-side motion vector correction based on bilateral matching, a distortion metric can be directly used to compare two candidate prediction blocks. The candidate block with the lowest distortion can then be used as the corrected block. Sub-block-based decoder-side motion vector correction can achieve finer granularity in motion vector correction. However, due to the smaller number of samples used to infer the matching cost, the matching cost may be less accurate, resulting in less accurate motion vector correction results and, in turn, less accurate encoding and decoding results. Summary of the Invention

[0007] This disclosure describes, among other things, a set of techniques for video (image) compression related to optical flow-based motion vector correction. These techniques include techniques for motion vector correction based on overlapping optical flows. For example, when calculating motion vectors for sub-blocks for optical flow correction, the sub-block size can be adaptively selected. Adapting the sub-block size to the optical flow correction has the advantage of allowing more samples to be used for optical flow correction, improving motion vector accuracy (e.g., at no additional cost), and thus improving codec accuracy.

[0008] According to some embodiments, a video decoding method includes: receiving a video stream (e.g., an encoded video sequence) including a plurality of blocks (e.g., corresponding to one or more pictures); (ii) deriving a set of sub-block motion vectors for a current block among the plurality of blocks; deriving a set of corrected sub-block motion vectors for the current block by applying optical flow correction to the set of sub-blocks of the current block, wherein corresponding sizes of the set of sub-blocks are adaptively selected for the optical flow correction; and reconstructing the current block using the set of corrected sub-block motion vectors.

[0009] According to some embodiments, a video encoding method includes: receiving video data (e.g., a source video sequence) including a plurality of blocks (e.g., corresponding to at least one picture), the plurality of blocks including a current block; identifying a set of sub-block motion vectors for the current block; identifying a set of corrected sub-block motion vectors for the current block by applying optical flow correction to the set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction; and encoding the current block using the set of corrected sub-block motion vectors.

[0010] According to some embodiments, a non-volatile computer-readable storage medium stores at least one set of instructions configured to be executed by a computing device having control circuitry and memory, the at least one set of instructions including instructions for the following operations: obtaining a source video sequence comprising a plurality of frames; and performing conversion between the source video sequence and a video stream of visual media data according to a format rule, wherein the video stream comprises a set of coded blocks; and wherein the format rule specifies: deriving a set of sub-block motion vectors for a current block in the set of coded blocks, and deriving a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction.

[0011] According to some embodiments, a video stream decoding device includes a processing circuit configured to execute the aforementioned video decoding method.

[0012] According to some embodiments, a video stream encoding device includes a processing circuit configured to execute the aforementioned video encoding method.

[0013] According to some embodiments, a computer device includes a memory for storing computer-readable instructions; a processor for reading the computer-readable instructions and executing the aforementioned video decoding method or video encoding method according to instructions of the computer-readable instructions.

[0014] According to some embodiments, a computer storage medium stores instructions, where the instructions can be executed by at least one processor to perform the aforementioned video encoding method, generate a code stream, and store the code stream.

[0015] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes a control circuit and a memory storing at least one set of instructions. The at least one set of instructions includes instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0016] According to some embodiments, a non-volatile computer-readable storage medium is provided. The non-volatile computer-readable storage medium stores at least one set of instructions for execution by a computing system. The at least one set of instructions includes instructions for performing any of the methods described herein.

[0017] Thus, devices and systems having methods for encoding and decoding video are disclosed. Such methods, devices, and systems can supplement or replace conventional methods, devices, and systems for video encoding / decoding.

[0018] The features and advantages described in the specification are not necessarily all-inclusive, and in particular, some additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, description, and claims provided in this disclosure. Furthermore, it should be noted that the language used in the specification is primarily selected for readability and instructional purposes and is not necessarily selected to describe or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order that the present disclosure may be understood in more detail, a more particular description may be given of the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings illustrate only the relevant features of the present disclosure and are therefore not necessarily to be considered limiting, as those skilled in the art will understand after reading this disclosure that the description may allow for other effective features.

[0020] Figure 1 is a block diagram illustrating an example communication system in accordance with some embodiments.

[0021] Figure 2A is a block diagram illustrating example elements of an encoder component according to some embodiments.

[0022] Figure 2B is a block diagram illustrating example elements of a decoder component according to some embodiments.

[0023] Figure 3 is a block diagram illustrating an example server system in accordance with some embodiments.

[0024] Figure 4A An example of deriving sub-block motion vectors according to some embodiments is illustrated.

[0025] Figure 4B Illustrated is an example of decoder-side motion vector modification according to some embodiments.

[0026] Figure 4C Illustrated is an example block partitioned into a set of sub-blocks in accordance with some embodiments.

[0027] Figure 4DIllustrated are example sub-blocks having adaptive sizes according to some embodiments.

[0028] Figure 5A An example video decoding process is illustrated in accordance with some embodiments.

[0029] Figure 5B An example video encoding process is illustrated in accordance with some embodiments.

[0030] In accordance with common practice, the various features illustrated in the drawings are not necessarily drawn to scale, and like reference numerals may be used to denote like features throughout the specification and drawings. DETAILED DESCRIPTION

[0031] The present disclosure describes video / image compression techniques that include temporal motion vector prediction (TMVP) and decoder-side motion vector correction (DMVR). TMVP techniques include sub-block-based TMVP techniques that use sub-block-level motion information from a co-located reference picture. DMVR techniques include applying bilateral matching to correct input motion vector pairs and using the corrected motion vector pairs for motion-compensated prediction of components. The techniques described herein also include bidirectional optical flow (BDOF) correction. BDOF correction can be used to correct a bidirectional prediction signal for a coding block at the sub-block level. The present disclosure further describes deriving a set of corrected sub-block motion vectors for a current block by applying optical flow correction (e.g., BDOF) to a set of sub-blocks of the current block. For this optical flow correction process, a corresponding size of the set of sub-blocks is adaptively selected. For example, the corresponding size can be predefined, derived, or signaled. Using adaptive sub-block sizes allows more samples to be used for optical flow correction, improving the accuracy of the corrected motion vectors and thereby improving codec accuracy.

[0032] Figure 1 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 through 120-m) communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, for example, for use with video applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0033] Source device 102 includes a video source 104 (e.g., a camera component or media storage space) and an encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates at least one encoded video stream from the video stream. The video stream from video source 104 may have a larger data size than encoded video stream 108 generated by encoder component 106. Because encoded video stream 108 has a lower data size (less data) than the video stream from the video source, encoded video stream 108 requires less bandwidth to transmit and less storage space to store than the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., is configured to transmit uncompressed video to network 110).

[0034] At least one network 110 represents any number of networks that communicate information between source device 102, server system 112, and / or electronic device 120, including, for example, cable (wired) and / or wireless communication networks. At least one network 110 can exchange data using circuit-switched and / or packet-switched channels. Representative networks include telecommunication networks, local area networks, wide area networks, and / or the Internet.

[0035] At least one network 110 includes a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from source device 102). Server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, encoder component 114 is configured to decode encoded video stream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video stream 108.

[0036] In some embodiments, server system 112 functions as a media-aware network element (MANE). For example, server system 112 can be configured to prune encoded video stream 108 to customize a potentially different stream for at least one electronic device 120. In some embodiments, the MANE is provided separately from server system 112.

[0037] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an output video stream that can be reproduced on a display or other type of presentation device. In some embodiments, at least one electronic device 120 does not include a display component (e.g., communicatively coupled to an external display device and / or includes media storage space). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0038] The source device and / or the plurality of electronic devices 120 are sometimes referred to as “end devices” or “user devices.” In some embodiments, the source device 102 and / or the at least one electronic device 120 are examples of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0039] In an example operation of the communication system 100, a source device 102 transmits an encoded video stream 108 to a server system 112. For example, the source device 102 may encode a stream of images captured by the source device. The server system 112 receives the encoded video stream 108 and may decode and / or encode the encoded video stream 108 using an encoder component 114. For example, the server system 112 may encode the video data to be more suitable for network transmission and / or storage. The server system 112 may transmit the encoded video data 116 (e.g., at least one encoded video stream) to at least one electronic device 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video images.

[0040] Figure 2Ais a block diagram illustrating example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives the video sequence from a remote video source (e.g., a video source belonging to a different device component than encoder component 106). Video source 104 can provide the source video sequence in the form of a digital video sample stream having any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device storing previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures that create a sense of motion when viewed sequentially. The image itself can be organized as a spatial array of pixels, where each pixel can have at least one sample, depending on the sampling structure used, the color space, etc. The relationship between pixels and samples can be easily understood by those skilled in the art.

[0041] The encoder component 106 is configured to encode and / or compress the pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is one of the functions of the controller 204. In some embodiments, the controller 204 controls other functional units described below and is functionally coupled to the other functional units. Parameters set by the controller 204 may include parameters related to rate control (e.g., picture skipping, lambda values ​​for quantizers and / or rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller 204 can be readily identified by one of ordinary skill in the art, as they may relate to encoder component 106 optimized for a particular system design.

[0042] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (when lossless compression is used between the symbols and the encoded video stream). The reconstructed sample stream (sample data) is input to a reference picture memory 208. Because decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory 208 are also bit-accurate between the local encoder and the remote encoder. This allows the encoder's prediction component to interpret the same sample values ​​as the decoder's prediction during decoding as reference picture samples.

[0043] The operation of decoder 210 may be the same as the operation of a remote decoder (such as decoder component 122), described below in conjunction with Figure 2B The decoder component 122 is described in detail. Figure 2B However, since the symbols are available and the encoding / decoding of the symbols into the encoded video sequence by the entropy encoder 214 and the parser 254 can be lossless, the entropy decoding portion of the decoder component 122 (including the buffer memory 252 and the parser 254) may not be fully implemented in the local decoder 210.

[0044] The decoder techniques described in this disclosure, except for parsing / entropy decoding, can be present in a corresponding encoder in essentially the same functional form. For this reason, this application focuses on the decoder operation. In addition, the description of the encoder techniques can be simplified because the encoder techniques can be reciprocal to the decoder techniques.

[0045] As part of its operation, source encoder 202 can perform motion-compensated predictive coding, predictively encoding an input frame with reference to at least one previously encoded frame from a video sequence designated as a reference frame. In this manner, encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of a reference frame that can be selected as a prediction reference for the input frame. Controller 204 can manage the encoding operations of source encoder 202, including, for example, setting parameters and subgroup parameters used to encode video data.

[0046] The decoder 210 may decode the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. Figure 2AWhen decoded at a remote video decoder (not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process that may be performed by the remote video decoder on the reference frames and may cause the reconstructed reference frames to be stored in reference picture memory 208. In this way, encoder component 106 locally stores copies of the reconstructed reference frames that have common content (absent transmission errors) with the reconstructed reference frames that will be obtained by the remote video decoder.

[0047] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new picture. The predictor 206 may operate on a pixel-by-pixel-block basis based on sample blocks to find a suitable prediction reference. Based on the search results obtained by the predictor 206, it may be determined that the input picture may have a prediction reference extracted from a plurality of reference pictures stored in the reference picture memory 208.

[0048] The outputs of all of the above functional units may undergo entropy encoding in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0049] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission via a communication channel 218, which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter can be configured to merge the encoded video data from the source encoder 202 with other data to be transmitted (e.g., encoded audio data and / or an auxiliary data stream (source not shown)). In some embodiments, the transmitter can send additional data along with the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include time / space / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), auxiliary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, etc.

[0050] Controller 204 can manage the operation of encoder component 106. During encoding, controller 204 can assign a certain coded picture type to each coded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predicted picture (P picture), or a bidirectionally predicted picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow for the use of different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those skilled in the art are familiar with these variations of I pictures and their respective applications and characteristics, and therefore will not be repeated here. Predicted pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block. Bidirectionally predicted pictures can be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indexes to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0051] A source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 168 samples each) and coded on a block-by-block basis. These blocks may be predictively coded with reference to other (already coded) blocks, determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictively coded (spatial prediction or intra-frame prediction) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be non-predictively coded by spatial prediction with reference to one previously coded reference picture, or by temporal prediction. Blocks of a B picture may be non-predictively coded by spatial prediction with reference to one or two previously coded reference pictures, or by temporal prediction.

[0052] The captured video may be presented as a temporal sequence of multiple source pictures (video pictures). Intra-picture prediction (often shortened to intra prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a previously encoded and buffered reference picture in the video, the block in the current picture can be encoded using a vector called a motion vector. The motion vector points to the reference block in a reference picture, and when multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0053] The encoder component 106 may perform encoding operations according to a predetermined video coding technique or standard, such as any of the techniques or standards described in this disclosure. In its operation, the encoder component 106 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video coding technique or standard being used.

[0054] Figure 2B is a block diagram illustrating example elements of decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter unit 256 and configured to transmit data to the display 124 (eg, via a wired or wireless connection).

[0055] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver can be configured to receive at least one encoded video sequence to be decoded by decoder component 122. In some embodiments, each encoded video sequence can be decoded independently of the other encoded video sequences. Each encoded video sequence can be received from channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The receiver can receive the encoded video data and other data (e.g., encoded audio data and / or ancillary data streams), which may be forwarded to its corresponding consuming entity (not depicted). The receiver can separate the encoded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data can be included as part of at least one encoded video sequence. Decoder component 122 can use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0056] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra-frame prediction unit 262, a motion compensated prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 can be implemented at least in part by software.

[0057] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to combat network jitter). When receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be required or may be very small. When used over a best-effort packet network (such as the Internet), buffer memory 252 may be required. Buffer memory 252 may be relatively large and / or have an adaptive size and may be implemented at least in part in an operating system or similar element external to decoder component 122.

[0058] Parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols may include, for example, information for managing the operation of decoder component 122 and / or information for controlling a presentation device, such as display 124. The control information for the presentation device may be in the form of, for example, a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). Parser 254 parses (entropy decodes) the encoded video sequence. The encoded video sequence may be encoded according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. Parser 254 may extract, from the encoded video sequence, a subgroup parameter set for at least one subgroup of pixels in a video decoder based on at least one parameter corresponding to the group. A subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), and the like. Parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and the like.

[0059] The reconstruction of symbol 270 may involve a number of different units, depending on the type of coded video picture or portion thereof (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information, which is parsed by parser 254 based on the coded video sequence. For clarity, the flow of such subgroup control information between parser 254 and the various units described below is not described.

[0060] The decoder component 122 can be conceptually subdivided into multiple functional units. In some embodiments, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the sake of clarity, the conceptual subdivision of the functional units is retained in this disclosure.

[0061] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform mode to use, block size, quantization factor, and / or quantization scaling matrix) from the parser 254 as at least one symbol 270. The scaler / inverse transform unit 258 may output a block comprising sample values, which may be input to an aggregator 268.

[0062] In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed picture, but can use predictive information from a previously reconstructed portion of the current picture. This predictive information can be provided by the intra prediction unit 262. The intra prediction unit 262 can use surrounding reconstructed information obtained from the current (partially reconstructed) picture in the current picture memory 264 to generate a block of the same size and shape as the block being reconstructed. The aggregator 268 can add the prediction information already generated by the intra prediction unit 262 to the output sample information as provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0063] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and potentially motion-compensated block. In this case, the motion-compensated prediction unit 260 can access the reference picture memory 266 to obtain samples for prediction. After motion compensation is performed on the obtained samples according to the symbols 270 belonging to the block, these samples can be added to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or residual signal) by the aggregator 268, thereby generating output sample information. The address within the reference picture memory 266 from which the motion-compensated prediction unit 260 obtains the predicted samples can be controlled by a motion vector. The motion vector can be provided to the motion-compensated prediction unit 260 in the form of symbols 270, which can have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​obtained from the reference picture memory 266, when, for example, sub-sample accurate motion vectors are used, and motion vector prediction mechanisms.

[0064] The output samples of aggregator 268 may be subjected to various loop filtering techniques in loop filter unit 256. The video compression techniques may include in-loop filtering techniques that are controlled by parameters included in the coded video bitstream and provided to loop filter unit 256 as symbols 270 from parser 254, but may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of a coded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0065] The output of the loop filter unit 256 may be a sample stream that may be output to a presentation device such as the display 124 and stored in the reference picture memory 266 for subsequent inter-picture prediction.

[0066] Once reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture is reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct a subsequent coded picture.

[0067] Decoder component 122 may perform decoding operations according to a predetermined video compression technique, which may be documented in a standard, such as any of the standards described herein. The encoded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard, as specified in the video compression technique documentation or standard, and specifically in the profile documents therein. Furthermore, to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within bounds, as defined by the level of the video compression technique or standard. In some cases, the level limits a maximum picture size, a maximum frame rate, a maximum reconstruction sample rate (e.g., measured in megasamples per second), a maximum reference picture size, and the like. In some cases, the limits set by the level may be further constrained by the Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the encoded video sequence.

[0068] Figure 33 is a block diagram illustrating a server system 112 according to some embodiments. Server system 112 includes control circuitry 302, at least one network interface 304, memory 314, a user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, control circuitry 302 includes at least one processor (e.g., a CPU, GPU, and / or DPU). In some embodiments, control circuitry includes at least one field programmable gate array (FPGA), a hardware accelerator, and / or at least one integrated circuit (e.g., an application specific integrated circuit).

[0069] At least one network interface 304 can be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical network). The communication network can be a local area network, a wide area network, a metropolitan area network, an in-vehicle network, an industrial network, a real-time network, a delay-tolerant network, and the like. Examples of communication networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, in-vehicle and industrial networks including CANBus, and the like. Such communication can be one-way receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using a local area digital network or a wide area digital network). Such communication can include communication to at least one cloud computing network.

[0070] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 may include at least one of the following: a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The at least one output device 308 may include at least one of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0071] Memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one magnetic disk storage device, optical disk storage device, flash memory device, and / or other non-volatile solid-state storage device). Memory 314 optionally includes at least one storage device located remotely from control circuitry 302. Memory 314, or optionally at least one non-volatile solid-state memory device within memory 314, includes a non-volatile computer-readable storage medium. In some embodiments, memory 314 or the non-volatile computer-readable storage medium of memory 314 stores the following programs, modules, instructions, and data structures, or a subset or superset thereof:

[0072] ● Operating system 316, including processes that handle various basic system services and perform hardware-related tasks;

[0073] • a network communications module 318 for connecting the server system 112 to other computing devices through at least one network interface 304 (eg, via wired and / or wireless connections);

[0074] Encoding module 320, for performing various functions related to encoding and / or decoding of data (such as video data). In some embodiments, encoding module 320 is an example of encoder component 114. Encoding module 320 includes, but is not limited to, at least one of the following:

[0075] o a decoding module 322 for performing various functions associated with decoding of encoded data, such as those previously described in connection with the decoder component 122; and

[0076] o Encoding module 340 for performing various functions related to data encoding, such as those previously described in connection with encoder component 106; and

[0077] Picture memory 352 for storing pictures and picture data, such as for use with encoding module 320. In some embodiments, picture memory 352 includes at least one of the following: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.

[0078] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described in connection with the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described in connection with the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described in connection with the motion compensated prediction unit 260 and / or the intra-frame prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described in connection with the loop filter unit 256).

[0079] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202, the encoding engine 212, and / or the entropy encoder 214) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0080] Each of the modules identified above and stored in memory 314 corresponds to a set of instructions for performing the functions described in the present disclosure. The modules identified above (e.g., multiple sets of instructions) do not need to be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, encoding module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform two sets of functions. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above.

[0081] although Figure 3 The server system 112 is shown according to some embodiments, but Figure 3 It is intended to be a functional description of various features that may be present in at least one server system rather than a schematic diagram of the structure of the embodiments described in this disclosure. In practice, items shown separately may be combined and some items may be separated. For example, Figure 3 Some items shown individually in the figure may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement server system 112, and how features are distributed among them, will vary from one embodiment to another and may depend in part on the amount of data traffic that the server system handles during peak usage periods and during average usage periods.

[0082] Example Codec Technology

[0083] The encoding and decoding processes and techniques described below may be performed at the devices and systems described above (eg, source device 102, server system 112, and / or electronic device 120). According to some embodiments, an inter prediction method with vector correction is described.

[0084] For example, decoder-side motion vector correction uses existing decoder-side information (e.g., reconstructed samples) to correct the motion vectors. As an example, the optical flow equation can be applied to express the least squares problem, whereby fine motion can be derived based on the gradients of the composite inter-frame prediction samples. Using these fine motions, the motion vector (MV) of each sub-block can be corrected within the prediction block, thereby improving the quality of inter-frame prediction. Decoder-side motion vector correction can be an extension of BDOF because the method supports MV correction when the two reference blocks have an arbitrary temporal distance from the current block. The gradient of the current entire block predictor can be pre-calculated, and then the corrected MV of the sub-block can be calculated based on the optical flow model.

[0085] Hereinafter, "prefetch samples" may refer to the maximum number of samples (from the reference picture) that can be used in the bilateral matching process, and "prefetch area" may refer to the area occupied by the prefetch samples. When performing the bilateral matching process, the greater the number of prefetch samples, the higher the memory bandwidth required.

[0086] As discussed above, some codecs (e.g., AV1 and VVC) operate on blocks of pixels. Each pixel block can be processed in a prediction-transform coding scheme, where the prediction value is obtained using reference pixels and / or motion compensation. For inter-frame predicted blocks, motion parameters such as motion vectors, reference picture indices, reference picture list usage indices and / or required additional information can be used to generate inter-frame predicted samples. The motion parameters can be signaled explicitly or implicitly. As discussed above, inter-frame predicted blocks can use temporal motion vectors and / or spatial motion vectors. In addition, sub-block level motion vector correction can be applied to extend the block level TMVP.

[0087] Figure 4A An example of deriving sub-block motion vectors according to some embodiments is illustrated. Figure 4A , the current picture 402 includes a current block 403, and the current block 403 is composed of sub-blocks 405-1 to 405-16. Figure 4A The number and size of sub-blocks and blocks in are only examples, and in other embodiments, different numbers and sizes of blocks and sub-blocks are used. Figure 4A Also shown is a reference picture 406 having a reference block 407 corresponding to the current block 403. Figure 4A In the example of , reference block 407 is identified using the motion shift derived from the motion in block A1. Figure 4A In FIG, the arrows in the sub-blocks illustrate motion vectors, where the dashed arrows correspond to motion from the L0 reference picture and the solid arrows correspond to motion from the L1 reference picture.

[0088] therefore, Figure 4A The figure shows an example of sub-block based TMVP (SbTMVP). SbTMVP can predict the motion vector of the sub-block in the current block in two steps. In the first step, the spatial neighbors are identified. Figure 4A In the example, A1 is represented by A1. If A1 has a motion vector that uses the same-position picture as its reference picture, then this motion vector is selected as the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift can be set to (0,0). Figure 4A The example in uses motion shifting based on the motion vector from block A1.

[0089] In the second step, the motion shift identified in the first step is applied (e.g., added to the coordinates of the current block) to obtain Figure 4A In the collocated picture shown, sub-block level motion information (motion vector and reference index) is obtained. Then, for each sub-block, the motion information of the sub-block is derived using the motion information of its corresponding block in the collocated picture (e.g., the minimum motion grid covering the center sample). After the motion information of the collocated sub-block is identified, it is converted into the motion vector and reference index of the current sub-block in a manner similar to the TMVP process, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current block.

[0090] As described above, SbTMVP allows motion information to be inherited from the co-located reference picture at the sub-block level. For example, each sub-block of a large-sized coding block (e.g., CU) can have its own motion information without explicitly transmitting the block partition structure or motion information. SbTMVP can obtain the motion information of each sub-block through three steps. The first step is to derive the displacement vector (DV) of the current coding block. In step two, the availability of SbTMVP candidates is accessed and the center motion is derived. In step three, the sub-block motion information is derived from the corresponding sub-block through the DV. Therefore, unlike TMVP candidate derivation (which always derives the temporal motion vector from the co-located block in the reference frame), SbTMVP can apply the DV derived from the MV of the left adjacent coding block of the current coding block to find the corresponding sub-block of each sub-block of the current CU in the co-located picture. In the case where the corresponding sub-block is not inter-coded, the motion information of the current sub-block can be set to the center motion.

[0091] In this way, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the coding block in the current picture. The same co-located picture used by TMVP can be used for SbTMVP. The difference between SbTMVP and TMVP is that TMVP predicts motion at the coding block level, while SbTMVP predicts motion at the sub-coding block level. In addition, TMVP obtains the temporal motion vector from the co-located block in the co-located picture (for example, the co-located block is the bottom right block or the center block relative to the current CU), while SbTMVP applies a motion shift before obtaining the temporal motion information from the co-located picture. The motion shift can be obtained from the motion vector of one of the spatially neighboring blocks from the current coding block.

[0092] Figure 4B illustrates an example of decoder-side motion vector modification according to some embodiments. Figure 4B In the example, a set of reference pictures 404 and 406 are used to derive corrected motion vectors (corrected MV0 and corrected MV1). The initial motion vectors MV0 and MV1 can be used to identify the initial reference block in each reference picture. Figure 4B 416 and 418 in the figure) are applied to Figure 4B For each motion vector in , a corrected motion vector is derived. Therefore, Figure 4B The figure shows an example of applying decoder-side motion vector correction (DMVR) to a coded block (e.g., in merge mode). The MV pairs obtained from the regular merge candidate can be used as input to the DMVR process. DMVR applies bilateral matching (BM) to correct the input MV pairs {mv L0 ,mv L1}, and the modified MV pair is used for motion compensation prediction (e.g., both the luma component and the chroma component). The output MV of DMVR (the modified MV pair) is defined in Equation Group 1:

[0093] mv refinedL0 =mv L0 +Δmv

[0094] mv refinedL1 =mv L1 -Δmv

[0095] Equation Group 1 – Corrected motion vector pair

[0096] In equation group 1, the motion vector difference Δmv is applied to the input MV pair to obtain a modified MV pair by using the MVD mirroring property (for example, because the input MV pair points to two different reference pictures, the picture order count (POC) differences of the two reference pictures from the current picture are equal, and the two reference pictures are in different temporal directions).

[0097] In the example DMVR process, the luma coding block is divided into 16×16 sub-blocks for the MV correction process. In the next two steps (integer precision motion search followed by fractional motion search), Δmv is derived independently for each sub-block. Finally, the corrected MV pair {mv refinedL0 ,mv refinedL1}, apply sub-block motion compensation (MC).

[0098] Integer sample offset search can be performed in DMVR. In an example embodiment, the search space includes MV pair candidates (e.g., 25 pairs of candidates), as shown in Equation Group 2:

[0099] mv L0(i,j) =mv L0(0,0) +(i,j)

[0100] mv L1(i,j) =mv L1(0,0) -(i,j)

[0101] Equation 2 – Search space for MV candidates

[0102] Where (i, j) represents the coordinates of the search points around the initial MV pair, and i and j are integer values ​​between -2 and 2, inclusive. The sum of absolute differences (SAD) of the initial MV pair is calculated as shown in the following equation group 3:

[0103]

[0104] Equation Group 3 – SAD Calculation

[0105] Where W and H are the width and height of the sub-block. If the SAD of the initial MV pair is less than a threshold, the integer sample phase of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase. In some embodiments, for example, to reduce the loss of DMVR correction uncertainty, the SAD between the reference blocks referenced by the initial MV candidate is reduced by 1 / 4 of the SAD value.

[0106] In some embodiments, the candidate MV pairs selected in the integer sample offset search step are further modified. For example, fractional sample modification can be derived using a parametric error surface equation (e.g., to save computational complexity) rather than performing an additional search using SAD comparisons. The fractional sample modification is conditionally invoked based on the output of the integer sample search phase. For example, the fractional sample modification is conditionally invoked based on the output of the integer sample search phase. As an example, the fractional sample modification is further applied when the integer sample search phase ends in the first or second iterative search with the center having the minimum SAD.

[0107] Optical flow correction (e.g., BDOF) can be used to correct the bidirectional prediction signal of a coding block (e.g., CU). For example, BDOF can be performed at the 4×4 sub-block level. BDOF can be applied to a coding block if it meets at least a subset of the following conditions: (i) the coding block is encoded using a bidirectional prediction mode, where one of the two reference pictures is arranged before the current picture in display order and the other is arranged after the current picture in display order; (ii) the distances (e.g., POC difference) of the two reference pictures to the current picture are the same; (iii) both reference pictures are short-term reference pictures; (iv) the coding block is not encoded using affine mode or SbTVMP merge mode; (v) the coding unit has more than 64 luma samples; (vi) the height and width of the coding unit are greater than or equal to 8 luma samples; (vii) the BCW weight index indicates equal weight; (viii) WP is not enabled for the current coding block; (ix) the current coding block does not use CIIP mode. In some embodiments, BDOF is applied only to the luma component.

[0108] The BDOF mode is based on the concept of optical flow and assumes that the motion of the object is smooth. For each sub-block (e.g., 4×4 sub-block), the motion correction (v x ,v y ). Motion correction can then be used to adjust the bidirectional prediction sample values ​​in the sub-block, as discussed below.

[0109] The horizontal gradient of the two prediction signals and vertical gradient It can be calculated by directly calculating the difference between two adjacent samples, as shown in the following equation group 4:

[0110]

[0111] Equation 4 – Gradient Calculation

[0112] Among them, I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k, k = 0, 1, shift1 is calculated based on the luma bit depth bitDepth, shift1 = max(6, bitDepth-6). Then, the autocorrelation and cross-correlation S1, S2, S3, S5 and S6 of the gradient are calculated, as shown in Equation Group 5:

[0113]

[0114] Equation Group 5 – Autocorrelation and Cross-correlation Calculations

[0115] where ψ and θ are defined by Equation 6:

[0116]

[0117] Equation 6

[0118] where Ω is a 6×6 window around the sub-block, n a and n b The values ​​of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively.

[0119] The motion correction (v) can then be derived using the cross-correlation and autocorrelation terms using Equation Group 7. x ,v y ):

[0120]

[0121] Equation Group 7 – Motion Correction Calculation

[0122] in, th′ BIO =2max(5,BD-7) . is the floor function,

[0123] Based on the motion correction and the gradient, the adjustment amount for each sample in the sub-block can be calculated using Equation Group 8:

[0124] Equation Group 8 – Motion Correction Adjustment Calculation

[0125] Then, the BDOF samples of the coding block can be calculated by adjusting the bidirectional prediction samples as follows:

[0126] pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o 偏移 )>>Shift Equation 9 – BDOF Calculation

[0127] These values ​​may be chosen such that the multiplier in the BDOF process does not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits.

[0128] To derive the gradient value, some prediction samples I outside the current coding block boundary in list k (k=0,1) can be identified and used (k) (i, j). For example, BDOF can use an extended row / column around the coding block boundary. For example, the prediction samples in the extended area can be generated by directly obtaining the reference samples at integer positions near the extended area (e.g., using a floor() operation on the coordinates) without interpolation (e.g., to control the computational complexity of generating prediction samples outside the boundary), and using a normal 8-tap motion compensated interpolation filter to generate the prediction samples within the coding block. These extended sample values ​​can be used for gradient calculations. For other steps in the BDOF process, if any samples and gradient values ​​outside the coding block boundary are needed, they can be filled (e.g., copied) from their nearest neighbors.

[0129] In some embodiments, multiple rounds of decoder-side motion vector correction are applied. For example, in the first round, bilateral matching (BM) is applied to the coding block, in the second round, BM is applied to each 16×16 sub-block within the coding block, and in the third round, the MV in each 8×8 sub-block is corrected by applying bidirectional optical flow (BDOF). The corrected MV can then be stored for subsequent spatial motion vector prediction and / or temporal motion vector prediction.

[0130] As discussed above, decoder-side motion vector correction uses existing decoder-side information (e.g., reconstructed samples) to correct motion vectors. In addition, the optical flow equation can be applied to express the least squares problem, whereby fine motion can be derived based on the gradients of the composite inter-frame prediction samples. Using these fine motions, the MV of each sub-block can be corrected within the prediction block, thereby improving the quality of inter-frame prediction. This embodiment is an extension of BDOF because the method supports MV correction when the two reference blocks have an arbitrary temporal distance from the current block. The gradient of the current entire block predictor can be pre-calculated, and then the corrected MV of the sub-block can be calculated based on the optical flow model.

[0131] Bilateral matching can be used for MV correction. In decoder-side motion vector correction based on bilateral matching, a distortion metric (such as the sum of absolute differences (SAD)) can be used directly to compare two candidate prediction blocks. The candidate block with the lowest distortion can then be used as the corrected block. Decoder-side motion vector correction can be performed at the block level or the sub-block level. At the sub-block level, the current block is divided into multiple sub-blocks, and bilateral matching is performed independently for each sub-block. Sub-block-based decoder-side motion vector correction can achieve motion vector correction at a finer granularity (e.g., sub-block level), but the matching cost may also be less accurate due to the smaller number of samples used to infer the matching cost.

[0132] Based on the above features and techniques, the following describes motion vector correction based on optical flow. For example, when calculating the sub-block MV for optical flow correction, the sub-block size can be adaptively selected. In this example, the final motion compensation is unaffected (e.g., using a fixed sub-block size).

[0133] Figure 4C A block 450 is illustrated as being partitioned into a set of sub-blocks 452 (eg, 452-1 through 452-n) in accordance with some embodiments. Figure 4C In the example of , the set of sub-blocks 452 has a uniform size (eg, square). In some embodiments, the set of sub-blocks can have varying sizes, shapes, and / or aspect ratios. Figure 4D Sub-blocks 454 (eg, 454-1 to 454-5) with adaptive sizes are illustrated in accordance with some embodiments. In some embodiments, sub-blocks 454 (with varying sizes) are used for optical flow correction and sub-blocks 452 are used for final motion compensation.

[0134] Figure 5Ais a flow chart illustrating a method 500 for decoding video according to some embodiments. Method 500 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 500 is performed by executing instructions stored in a memory of the computing system (e.g., memory 314).

[0135] The system receives (502) a video stream (e.g., an encoded video sequence), the video stream including a plurality of blocks corresponding to at least one picture. The system derives (504) a set of sub-block motion vectors for a current block of the plurality of blocks (e.g., using a TMVP technique). The system derives (506) a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein the corresponding sizes of the set of sub-blocks are adaptively selected (e.g., such as Figure 4D (as shown) for optical flow correction. The system uses the set of corrected sub-block motion vectors to reconstruct (508) the current block. In some embodiments, the sub-block size is adaptively selected when calculating the sub-block corrected MVs for optical flow correction. In some embodiments, the final motion compensation is not affected, for example, using a fixed sub-block size.

[0136] In some embodiments, at least one basic sub-block size is predefined, derived, or signaled. For example, in one coding mode (or specific condition), the sub-block size is defined as 8×8. If the sub-block is not located at the boundary of the entire block, the sub-block is further extended by N samples on each side to derive the corrected MV based on optical flow, for example, the initial interpolated samples, gradients, and corresponding operations are extended by N samples on each side (top, bottom, left, right). In one example, N is equal to 2. In this example, under the condition that the sub-block is not located at the boundary of the entire block, 12×12 interpolated samples and gradients are used to calculate the optical flow-based MV correction of the 8×8 sub-block. In this example, the final motion compensation (MC) is performed at the 8×8 sub-block.

[0137] In some embodiments, if the sub-block is located at the boundary (or corner) of the current full block, the sub-block is still expanded as described above, and the missing samples / gradient values ​​are filled in. In some embodiments, if the sub-block is located at the boundary (or corner) of the current full block, the sub-block is expanded by N1 samples (e.g., N1 is different from N) on each side, and the missing samples / gradient values ​​are filled in. In some embodiments, N1 is a value different from N. As an example, N is set to 2 and N1 is set to 1. In some embodiments, if the sub-block is located at the boundary (or corner) of the current full block, the sub-block is not expanded, and the correction amount is calculated using the basic (or predefined) sub-block size (e.g., 8×8), while another sub-block in the current full block that is not located at the boundary (or corner) is expanded as described above.

[0138] In some embodiments, if the sub-block is located at a corner of the current entire block (i.e., two edges of the sub-block are adjacent to the current entire block), then an N-sample expansion is performed on the other two sides of the sub-block located within the current entire block. For example, the current sub-block is 8×8 and the current entire block is 32×32. If the current sub-block is located at the upper left corner of the current entire block (e.g., the upper and left edges of the current sub-block are adjacent to the current entire block), then only the bottom and right sides of the current entire block are expanded (e.g., expanded by 2 samples) for use in optical flow-based MV correction derivation, which means that the sub-block size after expansion is 10×10.

[0139] In some embodiments, if the sub-block is located at the boundary (not the corner) of the current entire block (i.e., one edge of the sub-block is adjacent to the current entire block), then an N-sample expansion is performed on the three sides of the sub-block that are inside the current entire block. For example, the current sub-block is 8×8 and the current entire block is 32×32. If the current sub-block is located at the top boundary of the current entire block (e.g., only the top edge of the current sub-block is adjacent to the current entire block), then the bottom side, left side, and right side of the current entire block are expanded by 2 samples for optical flow-based MV correction derivation, which means that the sub-block size after expansion is 12×10.

[0140] In some embodiments, the number of extended samples (N) depends on the relative sub-block position within the coding block. For example, for sub-blocks closer to the coding block boundary, a smaller number of extended samples may be applied, and for sub-blocks farther from the coding block boundary, a larger number of extended samples may be applied.

[0141] Figure 5Bis a flow chart illustrating a method 550 for encoding a video according to some embodiments. Method 550 can be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, method 550 is performed by executing instructions stored in a memory of the computing system (e.g., memory 314). In some embodiments, method 550 is performed by the same system as method 500 described above.

[0142] The system receives (552) video data (e.g., a source video sequence) comprising a set of blocks, the set of blocks corresponding to at least one picture, and the set of blocks including a current block. The system identifies (554) a set of sub-block motion vectors for the current block. The system identifies (556) a set of corrected sub-block motion vectors for the current block by applying optical flow correction to the set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction. The system encodes (558) the current block using the set of corrected sub-block motion vectors. As previously described, the encoding process can mirror the decoding process (e.g., determining motion vectors) described herein. For the sake of brevity, these details are not repeated here.

[0143] although Figure 5A and Figure 5B The various logical stages are illustrated in a particular order, but stages that are not order-dependent may be reordered, and other stages may be combined or decomposed. Some reordering or other groupings not specifically mentioned will be apparent to one of ordinary skill in the art, and thus the ordering and groupings presented herein are not exhaustive. Furthermore, it should be appreciated that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0144] Turning now to some example embodiments.

[0145] (A1) In one aspect, some embodiments include a method for video decoding (e.g., method 500). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and at least one processor. In some embodiments, the method is performed at an encoding module (e.g., encoding module 320). The method includes: (i) receiving a video stream (e.g., an encoded video sequence) comprising a plurality of blocks (e.g., corresponding to at least one picture); (ii) deriving a set of sub-block motion vectors for a current block from the plurality of blocks (e.g., using a TMVP process); (iii) deriving a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein corresponding sizes of the set of sub-blocks are adaptively selected for the optical flow correction; and (iv) reconstructing the current block using the set of corrected sub-block motion vectors. For example, when calculating the sub-block corrected MVs for the optical flow correction, the sub-block size is adaptively selected. In this example, the final motion compensation is unaffected, i.e., a fixed sub-block size is used.

[0146] (A2) In some embodiments of A1, the corresponding sizes of the set of sub-blocks are predefined. For example, at least one basic sub-block size is predefined, derived, or signaled.

[0147] (A3) In some embodiments of A1, the respective sizes of the set of sub-blocks are derived or signaled. For example, at least one base size is predefined, and at least one adapted size is derived and / or signaled.

[0148] (A4) In some embodiments of any one of A1 to A3, the corresponding size includes a base size and a second size, wherein the second size corresponds to the base size with N additional columns and rows, where N is a non-negative integer. For example, for a particular coding mode (and / or under particular conditions), the sub-block size may be defined as 8×8. If the sub-block is not located at the boundary of the entire block, the sub-block is further extended by N samples on each side to derive a modified MV based on optical flow (e.g., initial interpolated samples, gradients, and corresponding operations are extended by N samples on each side (top, bottom, left, right)).

[0149] (A5) In some embodiments of A4, the value of N is determined based on the position of the subblock within the current block. For example, the number of extended samples depends on the relative position of the subblock within the coding block. In one example, a smaller number of extended samples may be applied to subblocks closer to the coding block boundary, while a larger number of extended samples may be applied to subblocks farther from the coding block boundary.

[0150] (A6) In some embodiments of A4 or A5, the value of N is selected from the following group: 1, 2, 3, and 4. For example, the base size may be 8×8, and N may be equal to 2. In this example, 12×12 interpolated samples and corresponding gradients are used to compute optical flow-based MV corrections for the 8×8 sub-block (e.g., under the condition that the sub-block is not located at the boundary of the entire block). The final motion compensation (MC) will be performed on the 8×8 sub-block.

[0151] (A7) In some embodiments of any one of A1 to A6, the method further includes padding an extended size of a sub-block located at a boundary of the current block with at least one value. For example, when the sub-block is located at a top boundary of the current block, padding is performed with top samples of the sub-block to obtain an extended sub-block. As another example, when the sub-block is located at a corner of the current block, padding is applied to both sides of the sub-block located at the boundary of the current block.

[0152] (A8) In some embodiments of any one of A1 to A7, the method further includes: expanding a first sub-block in the group of sub-blocks according to a first size among the corresponding sizes, including: (i) when the first sub-block is located at a boundary of the current block, expanding at least one first side of the first sub-block by N samples and filling at least one second side of the first sub-block; and (ii) when the first sub-block is not located at a boundary of the current block, expanding each side of the first sub-block by N samples. For example, if the sub-block is located at a boundary (or corner) of the current entire block, the sub-block can be expanded (for example, in the same manner as an internal sub-block) and the missing samples and / or gradient values ​​can be filled. For example, the first side is inside the current block and the second side is outside the current block.

[0153] (A9) In some embodiments of any one of A1 to A7, the method further includes: expanding a first sub-block in the group of sub-blocks according to a first size among the corresponding sizes, including: (i) when the first sub-block is located at a boundary of the current block, expanding at least one first side of the first sub-block by M samples and filling at least one second side of the first sub-block; and (ii) when the first sub-block is not located at a boundary of the current block, expanding each side of the first sub-block by N samples, where N is different from M. For example, if the sub-block is located at a boundary (or corner) of the current entire block, M samples may be expanded on each side of the sub-block, and missing samples and / or gradient values ​​may be filled. In this example, M is a number different from N. For example, N may be set to 2 and M may be set to 1.

[0154] (A10) In some embodiments of any one of A1 to A6, the method further includes: expanding a first subblock in the group of subblocks according to a first size among the corresponding sizes, including: (i) not expanding the first subblock when the first subblock is located at a boundary of the current block; and (ii) expanding each side of the first subblock by N samples when the first subblock is not located at a boundary of the current block. For example, if the subblock is located at a boundary (or corner) of the current entire block, the subblock is not expanded, and the subblock is calculated using a predefined subblock size (e.g., 8×8), while other subblocks in the current entire block that are not located at the boundary (or corner) are expanded to a larger size.

[0155] (A11) In some embodiments of any one of A1 to A6, the method further includes expanding a first sub-block in the group of sub-blocks according to a first size among the corresponding sizes, including: (i) when the first sub-block is located at a boundary of the current block, expanding the side of the first sub-block that is inside the current block, and not expanding the side of the first sub-block that is on the boundary of the current block; and (ii) when the first sub-block is not located at the boundary of the current block, expanding each side of the first sub-block by N samples. For example, if the sub-block is located at a corner of the current entire block (i.e., two edges of the sub-block are adjacent to the current entire block), the expansion of N samples is performed on the other two sides of the sub-block that are inside the current entire block. In one example, the current sub-block is 8×8 and the current entire block is 32×32. The current sub-block is located at the upper left corner of the current entire block, i.e., the top and left edges of the current sub-block are adjacent to the current entire block. In this case, only the bottom side and the right side of the current sub-block are extended (eg, extended by 2 samples) for optical flow-based MV correction derivation, which means the sub-block size after extension is 10×10.

[0156] As another example, if the sub-block is located at the boundary (not the corner) of the current entire block (i.e., one edge of the sub-block is adjacent to the current entire block), then an N-sample expansion is performed on the three sides of the sub-block that are inside the current entire block. In one example, the current sub-block is 8×8 and the current entire block is 32×32. The current sub-block is located at the top boundary of the current entire block, i.e., only the top edge of the current sub-block is adjacent to the current entire block. In this case, the bottom side, left side, and right side of the current sub-block are expanded (e.g., by 2 samples) for optical flow-based MV correction derivation, which means that the sub-block size after expansion is 12×10.

[0157] (B1) In another aspect, some embodiments include a method for video encoding (e.g., method 650). In some embodiments, the method is performed at a computing system (e.g., server system 112) having memory and control circuitry. In some embodiments, the method is performed at an encoding module (e.g., encoding module 320). The method includes: (i) receiving video data (e.g., a source video sequence) including a set of blocks (e.g., corresponding to at least one picture), the set of blocks including a current block; (ii) identifying a set of sub-block motion vectors for the current block; (iii) identifying a set of corrected sub-block motion vectors for the current block by applying optical flow correction to the set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction; and (iv) encoding the current block using the set of corrected sub-block motion vectors.

[0158] (C1) On the other hand, some embodiments include a method for visual media data processing. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and a control circuit. In some embodiments, the method is performed at an encoding module (e.g., encoding module 320). The method includes: (i) obtaining a source video sequence including a plurality of frames; and (ii) performing conversion between the source video sequence and a video stream of visual media data according to a format rule. The video stream includes a set of encoded blocks. The format rule specifies: (a) deriving a set of sub-block motion vectors for a current block in the set of encoded blocks, and (b) deriving a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein the corresponding sizes of the set of sub-blocks are adaptively selected for optical flow correction.

[0159] (C2) In some embodiments of C1, the respective sizes of the set of sub-blocks are predefined.

[0160] (C3) In some embodiments of C1, the respective sizes of the set of sub-blocks are derived or signaled.

[0161] (C4) In some embodiments of any one of C1 to C3, the corresponding size includes a base size and a second size, wherein the second size corresponds to the base size with N additional columns and rows, N being a non-negative integer.

[0162] On the other hand, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry system 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing at least one set of instructions configured to be executed by the control circuitry, the at least one set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A11, B1, and C1 to C4 above).

[0163] On the other hand, some embodiments include a non-volatile computer-readable storage medium storing at least one set of instructions for execution by control circuitry of a computing system, the at least one set of instructions including instructions for performing any of the methods described herein (e.g., A1 to A11, B1, and C1 to C4 above).

[0164] As used herein, N refers to a variable number. Unless explicitly stated otherwise, different instances of N may refer to the same number (eg, the same integer value, such as the number 2) or different numbers.

[0165] Unless otherwise specified, any syntax element described herein may be a high-level syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to a sequence level, a frame level, a slice level, or a tile level. As another example, an HLS element may be signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0166] It will be understood that although the terms "first", "second" etc. can be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "one", "an" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of at least one of the associated listed items. It will be further understood that when used in this specification, the terms "comprises" and / or "comprising" indicate the presence of stated features, wholes, steps, operations, elements and / or parts, but do not exclude the presence or addition of at least one other feature, whole, step, operation, element, part and / or its group.

[0167] As used herein, the term “if” may be interpreted to mean “when” or “at the time” or “in response to determining” or “upon determining” or “in response to detecting” that the stated precondition is true, depending on the context. Similarly, the phrase “if it is determined that [the stated precondition is true]” or “if [the stated precondition is true]” or “when [the stated precondition is true]” may be interpreted to mean “upon the determination” or “in response to the determination” or “upon determining” or “upon detecting” or “in response to detecting” that the stated precondition is true, depending on the context.

[0168] For purposes of explanation, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to understand.

Claims

1. A video decoding method, characterized in that: The method comprises: receiving a video stream comprising a plurality of blocks; deriving a set of sub-block motion vectors for a current block among the plurality of blocks; deriving a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction; and The current block is reconstructed using the set of corrected sub-block motion vectors.

2. The method according to claim 1, characterized in that The respective sizes of the set of sub-blocks are predefined.

3. The method according to claim 1, characterized in that The respective sizes of the set of sub-blocks are derived or signaled.

4. The method according to claim 1, wherein The corresponding size includes a basic size and a second size, wherein the second size corresponds to the basic size with N additional columns and rows, N being a non-negative integer.

5. The method according to claim 4, characterized in that The value of N is determined based on the position of the sub-block within the current block.

6. The method according to claim 4, characterized in that The value of N is selected from the following group: 1, 2, 3, and 4.

7. The method according to any one of claims 1 to 6, characterized in that Further including: The extended size of the sub-block located at the boundary of the current block is padded with at least one value.

8. The method according to any one of claims 1 to 6, characterized in that Further including: Expanding a first sub-block in the group of sub-blocks according to a first size in the corresponding sizes includes: When the first sub-block is located at the boundary of the current block, extending at least one first side of the first sub-block by N samples and padding at least one second side of the first sub-block; as well as When the first sub-block is not located at the boundary of the current block, each side of the first sub-block is extended by N samples.

9. The method according to any one of claims 1 to 6, characterized in that Further including: Expanding a first sub-block in the group of sub-blocks according to a first size in the corresponding sizes includes: When the first sub-block is located at the boundary of the current block, extending at least one first side of the first sub-block by M samples and padding at least one second side of the first sub-block; and When the first sub-block is not located at the boundary of the current block, each side of the first sub-block is extended by N samples, where N is different from M.

10. The method according to any one of claims 1 to 6, characterized in that Further including: Expanding a first sub-block in the group of sub-blocks according to a first size in the corresponding sizes includes: When the first sub-block is located at the boundary of the current block, not extending the first sub-block; as well as When the first sub-block is not located at the boundary of the current block, each side of the first sub-block is extended by N samples.

11. The method according to any one of claims 1 to 6, characterized in that: Further including: Expanding a first sub-block in the group of sub-blocks according to a first size in the corresponding sizes includes: When the first sub-block is located at the boundary of the current block, extending a side of the first sub-block inside the current block, and not extending a side of the first sub-block on the boundary of the current block; as well as When the first sub-block is not located at the boundary of the current block, each side of the first sub-block is extended by N samples.

12. A video encoding method, characterized in that: include: receiving video data comprising a plurality of blocks, the plurality of blocks including a current block; identifying a set of sub-block motion vectors for the current block; identifying a set of corrected sub-block motion vectors for the current block by applying optical flow correction to a set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction; as well as The current block is encoded using the set of corrected sub-block motion vectors.

13. The method according to claim 12, characterized in that The corresponding size includes a basic size and a second size, wherein the second size corresponds to the basic size with N additional columns and rows, N being a non-negative integer.

14. A non-volatile computer-readable storage medium, characterized in that: storing at least one set of instructions configured to be executed by a computing device having control circuitry and memory, the at least one set of instructions comprising instructions for: obtaining a source video sequence comprising a plurality of frames; and Perform conversion between the source video sequence and the video code stream of the visual media data according to the format rules, The video code stream includes a set of coded blocks; and Wherein, the format rules specify: deriving a set of sub-block motion vectors for a current block in the set of coded blocks, and A set of corrected sub-block motion vectors of the current block is derived by applying optical flow correction to a set of sub-blocks of the current block, wherein respective sizes of the set of sub-blocks are adaptively selected for the optical flow correction.

15. A video code stream decoding device, characterized in that: The method comprises a processing circuit configured to execute the method according to any one of claims 1 to 11.

16. A video code stream encoding device, characterized in that: The method comprises a processing circuit configured to perform the method according to any one of claims 12 to 13.

17. A computer device, characterized in that: The method comprises a memory for storing computer-readable instructions; a processor for reading the computer-readable instructions and executing the method according to any one of claims 1 to 11 and 12 to 13 as instructed by the computer-readable instructions.

18. A computer storage medium, characterized in that Instructions are stored, and the instructions can be executed by at least one processor to execute the video encoding method according to any one of claims 12 to 13, generate a code stream, and store it.