Systems and methods for decoder-side motion vector refinement
By using samples outside the current block/subblock on the decoder side for bilateral matching processing, the problem of inaccurate motion vector matching is solved, the accuracy and quality of video decoding are improved, and bandwidth requirements are reduced.
Patent Information
- Application Number
- CN202480004965.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2024-04-16
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video encoding technology has the problem of inaccurate motion vector matching in inter-frame prediction, resulting in a degradation in video decoding quality.
The decoder-side motion vector refinement technology is used to improve the accuracy of motion vector matching by using samples outside the current block/subblock for bilateral matching.
Improve the accuracy and quality of video decoding, and reduce the bandwidth requirement of video data in transmission and storage.
Smart Images

Figure CN120359749A_ABST
Abstract
Description
Related Applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 601,686, entitled "Overlapped Decoder-Side Motion Vector Refinement," filed on November 21, 2023, and this application is a continuation of and claims priority to U.S. Patent Application No. 18 / 622,842, entitled "Systems and Methods for Decoder-Side Motion Vector Refinement," filed on March 29, 2024. Technical Field
[0002] The disclosed embodiments generally relate to video coding and decoding, including but not limited to systems and methods for bilateral matching for motion vector refinement using samples outside of a current block / sub-block. Background Art
[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital imaging devices, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise convey digital video data across communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by a server providing cloud services or by hardware and / or software on an electronic / client device.
[0004] Video coding typically uses prediction methods that utilize the redundancy inherent in video data (e.g., inter-frame prediction, intra-frame prediction, etc.). Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H (Moving Picture Experts Group-H) project. The ITU-T (International Telecommunication Union-Telecommunication Standardization Sector) and ISO / IEC (International Organization for Standardization / International Electrotechnical Commission) released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard designed to succeed HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, the verified version 1.0.0 with Specification Errata 1 was released. Summary of the Invention
[0005] In addition, the present disclosure describes a set of techniques for video (image) compression related to inter-frame prediction and decoder-side motion vector refinement. For example, an overlapping decoder-side motion vector refinement method includes using bilateral matching processing of samples outside the current block / sub-block to improve the matching accuracy.
[0006] According to some embodiments, a method of video decoding includes: (i) receiving a video bitstream including a plurality of blocks; (ii) deriving a set of sub-block motion vectors for a current sub-block of a current block among the plurality of blocks; (iii) using bilateral matching based on one or more samples outside the current sub-block to derive a set of refined sub-block motion vectors for the current sub-block; and (iv) using the derived set of refined sub-block motion vectors to reconstruct the current sub-block.
[0007] According to some embodiments, a method of video encoding includes: (i) receiving video data including a plurality of blocks, the plurality of blocks including a current block; (ii) deriving a set of sub-block motion vectors for a current sub-block of the current block among the plurality of blocks; (iii) using bilateral matching based on one or more samples outside the current sub-block to derive a set of refined sub-block motion vectors for the current sub-block; and (iv) using the derived set of refined sub-block motion vectors to encode the current sub-block.
[0008] According to some embodiments, a method of processing visual media data includes: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of the visual media data, wherein the bitstream includes: a plurality of encoded pictures corresponding to a plurality of pictures, wherein the plurality of encoded pictures include one or more encoded sub-blocks encoded using a set of refined sub-block motion vectors, the set of refined sub-block motion vectors being derived using bilateral matching based on one or more samples outside the current sub-block.
[0009] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic device. The computing system includes control circuitry and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder). According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described herein.
[0010] Accordingly, apparatuses and systems that utilize methods for encoding and decoding video are disclosed. Such methods, apparatuses, and systems may supplement or replace conventional methods, apparatuses, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all inclusive, and in particular, given the drawings, specification, and claims provided in the present disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. Additionally, it should be noted that the language used in this specification has been selected primarily for readability and guidance purposes and is not necessarily selected to depict or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] For a more detailed understanding of the present disclosure, a more specific description can be made by referring to the features of various embodiments, some of which are shown in the drawings. However, the drawings only show relevant features of the present disclosure and are therefore not necessarily considered restrictive, as the specification may allow other valid features that those skilled in the art will understand upon reading the present disclosure.
[0012] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0013] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0014] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0015] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0016] Figure 4A shows an example of deriving a sub-block motion vector according to some embodiments.
[0017] Figure 4B shows an example of decoder-side motion vector refinement according to some embodiments.
[0018] Figure 4C shows an example of an interleaved (checkerboard) pattern according to some embodiments.
[0019] Figure 5A shows an example of a video decoding process according to some embodiments.
[0020] Figure 5B shows an example of a video encoding process according to some embodiments.
[0021] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, like reference numerals may be used to indicate like features. Detailed Description
[0022] The present disclosure describes video / image compression techniques including Temporal Motion Vector Prediction (TMVP) and Decoder Side Motion Vector Refinement (DMVR). The TMVP technique includes sub-block-based TMVP that uses motion information at the sub-block level from a collocated reference picture. The DMVR technique includes applying bilateral matching to refine an input motion vector pair and using the refined motion vector pair for component motion compensation prediction. The present disclosure also describes an overlapping DMVR refinement technique that includes using samples outside the current block (or sub-block) as input for the bilateral matching process of DMVR. The advantage of using samples outside the current block / sub-block for bilateral matching is improved video decoding accuracy (e.g., improved matching accuracy). Example systems and devices
[0023] Figure 1 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.
[0024] The source device 102 includes a video source 104 (e.g., a camera device component or a media storage device) and an encoder component 106. In some embodiments, the video source 104 is a digital camera device (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from the video source 104 may be of high data volume compared to the encoded video bitstream generated by the encoder component 106. Since the encoded video bitstream 108 is of lower data volume (less data) compared to the video stream from the video source, the encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder component 106 (e.g., configured to transmit uncompressed video to the network 110).
[0025] One or more networks 110 represent any number of networks for transferring information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wired (wired) and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0026] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is a streaming server or includes a streaming server (e.g., configured to store and / or distribute video content such as an encoded video stream from the source device 102). The server system 112 includes codec components 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec components 114 include an encoder component and / or a decoder component. In various embodiments, the codec components 114 are instantiated as hardware, software, or a combination of hardware and software. In some embodiments, the codec components 114 are configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0027] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., are communicatively coupled to an external display device and / or include a media storage device). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0028] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0029] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using the codec component 114. For example, the server system 112 may apply an encoding that is more optimized for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0030] Figure 2A is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCB or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0031] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204, as these functions may pertain to the encoder component 106 optimized for a specific system design.
[0032] In some embodiments, the encoder component 106 is configured to operate in an encoding / decoding loop. In a simplified example, the encoding / decoding loop includes a source encoder 202 (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (in the case where the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret when using prediction during decoding as reference picture samples.
[0033] The operation of the decoder 210 may be the same as that of a remote decoder such as the decoder component 122 described in detail below Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the encoding of the symbols into the encoded video sequence by the entropy encoder 214 and the decoding of the symbols by the parser 254 can be lossless, the entropy decoding part of the decoder component 122 including the buffer memory 252 and the parser 254 may not be fully implemented in the local decoder 210.
[0034] Except for parsing / entropy decoding, the decoder techniques described herein can exist in a form of substantially the same functionality in the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques can be simplified because encoder techniques can be inverse to decoder techniques.
[0035] As part of the operation of the source encoder 202, the source encoder 202 can perform motion compensated predictive coding that predictively encodes an input frame with reference to one or more previously encoded frames designated as reference frames from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of the reference frame, where the reference frame can be selected as a predictive reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0036] The decoder 210 decodes the encoded video data of frames that can be designated as reference frames based on symbols created by the source encoder 202. The operation of the encoding engine 212 can be advantageously lossy processing. When the encoded video data is decoded at a video decoder ( Figure 2A (not shown)), the reconstructed video sequence can be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that can be performed by a remote video decoder on a reference frame and can store the reconstructed reference frame in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame that has common content (in the absence of transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.
[0037] The predictor 206 can perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 can search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can be used as an appropriate prediction reference for the new picture. The predictor 206 can operate on a per pixel block basis of sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory 208.
[0038] The outputs of all the above-mentioned functional units can undergo entropy encoding in the entropy encoder 214. The entropy encoder 214 converts these symbols into an encoded video sequence by losslessly compressing the symbols generated by various functional units according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).
[0039] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence created by the entropy encoder 214 in preparation for transmission via a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR (Signal-to-Noise Ratio) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, and the like.
[0040] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures may be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those of ordinary skill in the art are familiar with those variations of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures may be encoded and decoded using inter prediction or intra prediction that uses at most one motion vector and a reference index to predict the sample values of each block. Bi-predictive pictures may be encoded and decoded using inter prediction or intra prediction that uses at most two motion vectors and a reference index to predict the sample values of each block. Similarly, multi-predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0041] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and encoded on a block-by-block basis. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined by the encoding assignments applied to the corresponding pictures of the blocks. For example, blocks of an I picture can be non-predictively encoded, or blocks of an I picture can be predictively encoded (spatial prediction or intra prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture can be non-predictively encoded with reference to a previously encoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture can be non-predictively encoded with reference to one or two previously encoded reference pictures via spatial prediction or via temporal prediction.
[0042] Video can be captured as a plurality of source pictures (video pictures) in a time series. Intra picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture in encoding / decoding, which is referred to as the current picture, is segmented into blocks. In a case where a block in the current picture is similar to a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0043] The encoder component 106 can perform encoding operations according to any predetermined video encoding technique or standard such as those described herein. In the operation of the encoder component 106, the encoder component 106 can perform various compression operations, including predictive encoding operations that utilize the temporal redundancy and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technique or standard used.
[0044] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data (e.g., via a wired connection or a wireless connection) to the display 124.
[0045] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data (e.g., via a wired or wireless connection) from channel 218. The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0046] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. Decoder component 122 may be implemented at least partially in software.
[0047] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to counter network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be required, or buffer memory 252 may be small. To make the best use of packet networks such as the Internet, buffer memory 252 may be required, buffer memory 252 may be relatively large and / or have an adaptive size, and may be implemented at least partially in the operating system or a similar element external to decoder component 122.
[0048] The parser 254 is configured to reconstruct symbols 270 from an encoded video sequence. The symbols may include, for example, information for managing the operation of decoder components 122 and / or information for controlling a rendering device such as display 124. The control information for the rendering device may be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0049] Depending on the type of the encoded video picture or a portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the manner in which they are involved may be controlled by the parser 254 through subgroup control information parsed from the encoded video sequence. For clarity, this subgroup control information flow between the parser 254 and the multiple units below is not depicted.
[0050] The decoder component 122 may be conceptually subdivided into multiple functional units, and in some implementations, these units interact closely with each other and may be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained herein.
[0051] The Scaler / Inverse Transform Unit 258 receives, from the Parser 254, the quantized transform coefficients as symbols 270, as well as control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix). The Scaler / Inverse Transform Unit 258 may output a block including sample values, which may be input into the Aggregator 268. In some cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the Intra Picture Prediction Unit 262. The Intra Picture Prediction Unit 262 may generate a block of the same size and shape as the block being reconstructed using surrounding reconstructed information obtained from the current (partially reconstructed) picture from the Current Picture Memory 264. The Aggregator 268 may add the predictive information that the Intra Picture Prediction Unit 262 has generated to the output sample information provided by the Scaler / Inverse Transform Unit 258, on a per-sample basis.
[0052] In other cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to inter-coded and potentially motion-compensated blocks. In such a case, the Motion Compensation Prediction Unit 260 may access the Reference Picture Memory 266 to obtain samples for prediction. After motion-compensating the obtained samples according to the symbols 270 belonging to the block, these samples may be added by the Aggregator 268 to the output of the Scaler / Inverse Transform Unit 258 (referred to as residual samples or residual signal in this case) to generate output sample information. The address within the Reference Picture Memory 266 from which the Motion Compensation Prediction Unit 260 obtains the prediction samples may be controlled by a motion vector. The motion vector may be available to the Motion Compensation Prediction Unit 260 in the form of symbols 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include, for example, interpolation of sample values obtained from the Reference Picture Memory 266, a motion vector prediction mechanism when using sub-sampled accurate motion vectors.
[0053] The output samples of aggregator 268 can undergo various loop filtering techniques in loop filter unit 256. Video compression techniques can include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream and available to loop filter unit 256 as symbols 270 from parser 254, but video compression techniques can also respond to meta-information obtained during decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values. The output of loop filter unit 256 can be a sample stream that can be output to a rendering device such as display 124, and stored in reference picture memory 266 for use in future inter-picture prediction.
[0054] Once reconstructed, some encoded pictures can be used as reference pictures for future prediction. Once an encoded picture has been reconstructed and the encoded picture has been identified (e.g., by parser 254) as a reference picture, the current reference picture can become part of reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct subsequent encoded pictures.
[0055] Decoder component 122 can perform decoding operations according to a predetermined video compression technique that can be recorded in a standard such as any of the standards described herein. As specified in a video compression technique document or standard and in particular in a profile thereof, an encoded video sequence can conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. In addition, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by a Hypothetical Reference Decoder (HRD) specification and metadata for signaling HRD buffer management in the encoded video sequence.
[0056] Figure 3is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes one or more processors (e.g., a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and / or a DPU (Data Processing Unit)). In some embodiments, the control circuitry includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application specific integrated circuit).
[0057] The network interface 304 may be configured to interface with one or more communication networks (e.g., a wireless network, a wired network, and / or an optical network). The communication network may be local, wide area, metropolitan area, vehicular and industrial, real-time, delay tolerant, etc. Examples of communication networks include: local area networks such as Ethernet, wireless LAN (Local Area Network); cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CANBus (Controller Area Network - BUS), etc. Such communication may be only unidirectional reception (e.g., broadcast TV), only unidirectional transmission (e.g., CANBus to certain CAN bus devices), or bidirectional (e.g., to other computer systems using a local digital network or a wide area digital network). Such communication may include communication to one or more cloud computing networks.
[0058] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input device 310 may include one or more of the following: a keyboard, a mouse, a touchpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera device, etc. The output device 308 may include one or more of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or a screen), etc.
[0059] The memory 314 may include high-speed random access memory (such as DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory), DDR RAM (Double Data Rate Random Access Memory), and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices located remotely from the control circuitry 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316 that includes procedures for handling various basic system services and for performing hardware-related tasks; · A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection); · A codec module 320 for performing various functions related to encoding and / or decoding data such as video data. In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: ο A decoding module 322 for performing various functions related to decoding encoded data, such as those functions previously described with respect to the decoder component 122; and ο An encoding module 340 for performing various functions related to encoding data, such as those functions previously described with respect to the encoder component 106; and · A picture memory 352, such as for storing pictures and picture data for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.
[0060] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform the various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).
[0061] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the modules shown in. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.
[0062] Each of the modules identified above stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., the instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform two sets of functions. In some embodiments, the memory 314 stores a subset of the modules and data structures identified above. In some embodiments, the memory 314 stores additional modules and data structures not described above.
[0063] Although Figure 3 a server system 112 is shown in accordance with some embodiments, Figure 3 it is intended more as a functional description of the various features that may exist in one or more server systems than as a structural schematic of the embodiments described herein. In practice, the items shown separately may be combined and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among the servers will vary depending on the implementation, and optionally, it depends in part on the amount of data traffic processed by the server system during peak usage periods as well as during average usage periods. Example encoding and decoding techniques
[0064] As discussed above, some codecs (e.g., AV1 and VVC) operate on pixel blocks. Each pixel block can be processed in a predictive transform coding scheme, where prediction is obtained using reference pixels and / or motion compensation. For an inter-predicted block, motion parameters such as motion vectors, reference picture indices, reference picture list use indices, and / or additional information required can be used for inter-predicted sample generation. The motion parameters can be signaled in an explicit or implicit manner. As discussed above, the inter-predicted block can use temporal motion vectors and / or spatial motion vectors. Additionally, sub-block level motion vector refinement can be applied to extend block-level TMVP.
[0065] Figure 4A An example of deriving sub-block motion vectors according to some embodiments is shown. In Figure 4A , the current picture 402 includes a current block 403 composed of sub-blocks 405-1 to 405-15. Figure 4A The number and size of sub-blocks and blocks in Figure 4A are only examples, and in other embodiments, blocks and sub-blocks of different numbers and sizes are used. Figure 4A A reference picture 406 is also shown, which has a reference block 407 corresponding to the current block 403. In the example of Figure 4A , a motion shift derived from the motion in block A1 is used to identify the reference block 407.
[0066] Therefore, Figure 4A An example of Subblock-based TMVP (SbTMVP) is shown. SbTMVP can predict the motion vectors of sub-blocks within the current block in two steps. In the first step, spatial neighbors are identified, which are represented as A1 in Figure 4A . If A1 has a motion vector using a collocated picture as its reference picture, that motion vector is selected as the motion shift (or displacement vector) to be applied. If no such motion is identified, the motion shift can be set to (0,0). Figure 4A The example in
[0067] uses a motion shift based on the motion vector from block A1. Figure 4AThe co-located picture shown obtains sub-block level motion information (motion vectors and reference indices). Then, for each sub-block, the motion information of its corresponding block (e.g., the smallest motion grid covering the central sample) in the co-located picture is used to derive the motion information of the sub-block. After identifying the motion information of the co-located sub-block, it is converted into the motion vectors and reference indices of the current sub-block in a manner similar to TMVP processing, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current block.
[0068] As described above, SbTMVP allows inheriting motion information at the sub-block level from the co-located reference picture. For example, each sub-block of a large-sized coded block (e.g., a CU) can have its own motion information without explicitly transmitting the block partition structure or motion information. SbTMVP can obtain the motion information of each sub-block in three steps. The first step is to derive the displacement vector (DV) of the current coded block. In the second step, the availability of SbTMVP candidates is accessed and the central motion is derived. In the third step, the sub-block motion information is derived from the corresponding sub-block through the DV. Thus, different from the TMVP candidate derivation that always derives the temporal motion vector from the co-located block in the reference frame, SbTMVP can apply the DV derived from the MV (Motion Vector, MV) of the left neighboring coded block of the current coded block to find the corresponding sub-block in the co-located picture for each sub-block of the current CU. In the case where the corresponding sub-block is not inter-frame coded, the motion information of the current sub-block can be set to the central motion.
[0069] In this way, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode for the coded blocks in the current picture. The same co-located picture used by TMVP can be used for SbTMVP. The difference between SbTMVP and TMVP is that TMVP predicts the motion at the coded block level, while SbTMVP predicts the motion at the sub-coded block level. Additionally, TMVP obtains the temporal motion vector from the co-located block in the co-located picture (e.g., the co-located block is the bottom-right or central block relative to the current CU), while SbTMVP applies a motion shift before obtaining the temporal motion information from the co-located picture. The motion shift can be obtained from the motion vector of one of the spatial neighboring blocks of the current coded block.
[0070] Figure 4B An example of decoder-side motion vector refinement according to some embodiments is shown. In Figure 4B it, a set of reference pictures 404 and 406 are used to derive the refined motion vectors (refined MV0 and refined MV1). The initial motion vectors MV0 and MV1 can be used to identify the initial reference blocks in each reference picture. The motion difference (by Figure 4BThe arrows 416 and 418 in) are applied to Figure 4B each motion vector in to derive a refined motion vector. Thus, Figure 4B shows an example of applying decoder-side motion vector refinement (DMVR) to an encoded block (e.g., in merge mode). The MV pair obtained from the regular merge candidates can be used as the input for DMVR processing. DMVR applies bilateral matching (BM) to refine the input MV pair {mv L0 , mv L1} and uses the refined MV pair for motion compensation prediction (e.g., motion compensation prediction for both the luminance component and the chrominance component). The output MV of DMVR - the refined MV pair - is defined in Equation Set 1: mV 经细化的L0 = mv L0 + Δmv mv 经细化的L1 = mv L1 - Δmv Equation Set 1 - Refined motion vector pair
[0071] In Equation Set 1, the motion vector difference Δmv is applied to the input MV pair to obtain a refined MV pair by using the MVD (Motion Vector Difference, MVD) mirroring property (e.g., because the input MV pair points to two different reference pictures, the two different reference pictures have an equal difference from the current picture in picture order count (POC), and the two reference pictures are at different temporal directions).
[0072] In an example DMVR process, the luminance encoded block is divided into 16×16 sub-blocks for MV refinement processing. Δmv is independently derived for each sub-block in the following two steps: integer-precision motion search, followed by fractional motion search. Finally, the sub-block motion compensation (MC) is applied using the refined MV pair {mv 经细化的L0 , mv 经细化的L1}.
[0073] Integer sample offset search can be performed in DMVR. In an example implementation, the search space includes MV pair candidates (e.g., 25 pairs of candidates), as shown in Equation Set 2: mv L0(i,j) = mv L0(0,0) + (i, j) mv L1(i,j) = mv L1(0,0) - (i, j) Equation Group 2 - Search Space for MV Candidates where (i, j) represents the coordinates of the search points around the initial MV pair, and i and j are integer values between -2 and 2 (including -2 and 2). The sum of absolute differences (SAD) of the initial MV pair is calculated as shown in Equation Group 3 below: diff m,n = abs(P0 i,j [m + i, 2n + j] - P1 i,j [m - i, 2n - j]) Equation Group 3 - SAD Calculation where W and H are the weight and height of the sub-block. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. In some embodiments, for example, to reduce the penalty of the uncertainty of DMVR refinement, the SAD between the reference blocks referenced by the initial MV candidate is reduced by 1 / 4 of the SAD value.
[0074] In some embodiments, the candidate MV pair selected in the integer sample offset search step is further refined. For example, fractional sample refinement can be derived by using a parametric error surface equation (e.g., to save computational complexity) instead of additional search by SAD comparison. The fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. For example, the fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. As an example, when the integer sample search stage terminates at the center with the minimum SAD in the first iterative search or the second iterative search, the fractional sample refinement is further applied.
[0075] BDOF (Bi-Directional Optical Flow) can be used to refine the bi-prediction signals of coding blocks (e.g., CUs). For example, BDOF can be performed at the 4×4 sub-block level. If a coding block satisfies at least one subset of the following conditions, BDOF can be applied to the coding block: (i) the coding block is encoded using the following bi-prediction mode, in which one of the two reference pictures is before the current picture in display order and the other reference picture is after the current picture in display order; (ii) the distances from the two reference pictures to the current picture (e.g., POC difference) are the same; (iii) the two reference pictures are short-term reference pictures; (iv) the coding block is not encoded using the affine mode or the SbTVMP merge mode; (v) the coding unit has more than 64 luma samples; (vi) the height and width of the coding unit are greater than or equal to 8 luma samples; (vii) the BCW (Bi-Prediction with CU-Level Weight) weight index indicates equal weights; (viii) WP (Weighted Prediction) is not enabled for the current coding block; and (ix) the CIIP (Combined Inter and Intra Prediction) mode is not used for the current coding block. In some implementations, BDOF is applied only to the luma component.
[0076] The BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each sub-block (e.g., 4×4 sub-block), the motion refinement (v x , v y ) can be calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. Then, the motion refinement can be used to adjust the bi-prediction sample values in the sub-block, as discussed in detail below.
[0077] The horizontal and vertical gradients of the two prediction signals can be calculated by directly computing the difference between two adjacent samples and as shown in Equation Set 4 below: Equation Set 4 - Gradient Calculation where I (k) (i, j) is the sample value at the coordinate (i, j) of the prediction signal in list k (k = 0, 1), and shift1 is calculated based on the luma bit depth bitDepth, such that shift1 = max(6, bitDepth - 6). Then, the autocorrelation and cross-correlation of gradients S1, S2, S3, S5, and S6 are calculated as shown in Equation Set 5: S5 = ∑ (i,j)∈Ω Abs(ψ y , (i, j)), S6 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ y , (i, j)) Equation set 5 - Autocorrelation and cross - correlation calculation where ψ and θ are defined by equation set 6: θ(i, j) = (I (1) (i, j) >> n b ) - (I (0) (i, j) >> n b ) Equation set 6 where Ω is a 6×6 window around the sub - block, and n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8) respectively.
[0078] Then, the motion refinement (v x , v y ) can be derived using the cross - correlation term and autocorrelation term with equation set 7: Equation set 7 - Motion refinement calculation where is the floor function, and
[0079] Based on the motion refinement and gradient, the adjustment can be calculated for each sample in the sub - block using equation set 8: Equation set 8 - Motion refinement adjustment calculation
[0080] Then, the BDOF samples of the coded block can be calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,, y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift Equation 9 - BDOF calculation
[0081] These values can be selected such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits.
[0082] To derive gradient values, some prediction samples I outside the current coded block boundary in the list k (k = 0, 1) can be identified and used. (k) (i, j). For example, BDOF can use an extended horizontal / vertical row around the coded block boundary. For example, prediction samples in the extended region can be generated by directly obtaining reference samples at nearby integer positions (e.g., using the floor() operation on the coordinates) without interpolation (e.g., to control the computational complexity of generating prediction samples outside the boundary), and a normal 8-tap motion compensation interpolation filter can be used to generate prediction samples within the coded block. These extended sample values can be used in the gradient calculation. For other steps in the BDOF process, if any samples and gradient values outside the coded block boundary are needed, the samples and gradient values can be filled (e.g., repeated) from the nearest neighbors of the samples and gradient values.
[0083] In some implementations, multiple-pass decoder-side motion vector refinement is applied. For example, in the first pass, bilateral matching (BM) is applied to the coded block, in the second pass, BM is applied to each 16×16 sub-block within the coded block, and in the third pass, the MVs in each 8×8 sub-block are refined by applying bidirectional optical flow (BDOF). Then, the refined MVs can be stored for subsequent spatial and / or temporal motion vector prediction.
[0084] As discussed above, decoder-side motion vector refinement uses existing decoder-side information such as reconstructed samples to refine the motion vectors. Additionally, optical flow equations can be applied to formulate a least squares problem, from which the fine motion can be derived from the gradients of the composite inter-prediction samples. Using these fine motions, the MVs can be refined sub-block by sub-block within the prediction block, thereby enhancing the inter-prediction quality. This implementation is an extension of BDOF because it supports MV refinement when the two reference blocks have an arbitrary temporal distance from the current block. The gradients of the current entire block predictor can be pre-computed, and then the refined MVs for the sub-blocks can be calculated according to the optical flow model.
[0085] Bilateral matching can be used for MV refinement. In decoder-side motion vector refinement based on bilateral matching, a distortion metric such as sum of absolute differences (SAD) can be used to directly compare two candidate predictor blocks. Then, the candidate block with the lowest distortion is used as the refined block. Decoder-side motion vector refinement can be performed at the block level or sub-block level. At the sub-block level, the current block is divided into multiple sub-blocks, and bilateral matching for each sub-block is performed independently. Sub-block-based decoder-side motion vector refinement can achieve motion vector refinement with a finer granularity (e.g., sub-block level), but the matching cost may also be less accurate because fewer samples are used to infer the matching cost.
[0086] The prefetch samples referred to in this paper refer to the restricted maximum number of samples from the reference picture that can be used in the bilateral matching process. The prefetch region referred to in this paper refers to the region occupied by the prefetch samples. When performing the bilateral matching process, the larger the number of prefetch samples, the higher the memory bandwidth required.
[0087] In some embodiments, subsampling is performed using a staggered pattern. Figure 4C An example staggered (plum blossom) pattern according to some embodiments is shown. In Figure 4C the example, the circles represent samples (e.g., pixels), where the solid circles represent the selected samples, and the unfilled circles represent the unselected (e.g., skipped) samples.
[0088] Based on the features and techniques described above, the overlapping DMVR technique is described below.
[0089] In the example, bilateral matching between two predictors is performed to find the refined MV based within a 5×5 search window. Sum of absolute differences (SAD) is used as the distortion metric for bilateral matching. In some embodiments, the (sub-)block size is extended by 2 samples on each side to perform bilateral matching, while the final motion compensation (MC) remains unchanged. In this way, coding gain can be achieved without increasing the memory bandwidth.
[0090] Sub - block - based MV refinement is a useful coding tool for bidirectional inter - prediction as it utilizes bilateral matching to further refine the current MV without additional signaling overhead. In some implementations, the coding block is partitioned into 16×16 (8×16 / 16×8 if the width or height is 8) sub - blocks to perform bilateral matching. For example, during the MV refinement search, a bilinear interpolation filter can be used, while in the final motion compensation of the refined (sub) - block, an 8 - tap interpolation filter and sample padding are used. In one example, 2 samples on each side of the MV refinement search stage - i.e., a 20×20 (12×20 / 20×12 if the width or height is 8) sub - block - are used to perform bilateral matching (e.g., to achieve better matching accuracy).
[0091] In some systems, considering optical flow enabled, the memory bandwidth is equal to (w + 2+7)x(h + 2+7) samples. Using MV refinement with 2 additional samples as described above, the memory bandwidth required to perform MV refinement is equal to (w + 4 (search range)+4 (extended samples)+1 (bilinear interpolation))×(h + 4 (search range)+4 (extended samples)+1 (bilinear interpolation)), which does not increase the worst - case memory bandwidth.
[0092] Figure 5A is a flowchart showing a method 500 for decoding video according to some embodiments. Method 500 can be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 500 is executed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0093] The system receives (502) a video bitstream including a plurality of blocks. The system derives (504) a set of sub - block motion vectors for a current sub - block of a current block among the plurality of blocks. The system derives (506) a set of refined sub - block motion vectors for the current sub - block using bilateral matching based on one or more samples outside the current sub - block. The system reconstructs (508) the current sub - block using the derived set of refined sub - block motion vectors. In some embodiments, samples outside the current block / sub - block are used to perform the refined bilateral matching process (e.g., to improve matching accuracy). In some embodiments, the size of the samples used in the refined bilateral matching process is larger than the size of the current block / sub - block. For example, the current block / sub - block is a subset of the samples used for performing bilateral matching.
[0094] In some embodiments, partial samples within the current block / sub - block are used together with selected samples outside the current block / sub - block to perform the refined bilateral matching process.
[0095] In some embodiments, an additional N rows of samples around the current block / sub-block are used for distortion calculation in bilateral matching. For example, if the current sub-block size for refinement is 16×16 and N equals 2, an 18×18 sample sub-block of the two candidate predictors is interpolated, and the distortion between the two blocks is calculated and used for distortion comparison. The distortion calculation technique and the interpolation filter can use SAD and a bilinear filter. As an example, the value of N can be hard-coded (e.g., using a default value). In another example, the number N is preselected and signaled in the bitstream (e.g., at the frame / sequence level or other higher levels). In this way, N is switchable at the frame / sequence level or other higher levels. In another example, N is limited to 1 (e.g., to limit the number of prefetch samples from off-chip memory to on-chip memory). In another example, only the additional N rows above the top block / sub-block boundary and / or the additional N rows below the bottom block / sub-block boundary are used to perform bilateral matching (e.g., without using the additional rows beyond the left and / or right block / boundaries to perform bilateral matching). In another example, when expanding the N rows and performing bilateral matching, subsampled blocks and reference regions are used.
[0096] In some embodiments, an additional N rows of samples around the current block / sub-block are used for distortion calculation in bilateral matching. For example, when the additional rows fall outside a predefined prefetch region, a padding process is performed (e.g., to avoid increasing memory bandwidth). In various embodiments, the padding process is performed at the entire block level or sub-block level. As an example, the prefetch region is defined as the size of the current entire block (or current sub-block size) plus seven samples (e.g., to perform 8-tap interpolation). If the additional N rows of the current candidate block / sub-block pointed to by the current MV + refinement MV fall outside the prefetch region, the external samples are padded.
[0097] When optical flow is enabled, the prefetch region can be defined as the current sub-block size plus M samples, where the value of M is different from N (e.g., M is set to 9).
[0098] In some embodiments, an additional row of samples around the current block / sub-block is used for distortion calculation in bilateral matching. For example, an asymmetric extension can be used for the additional rows, and a normalized SAD can be used for distortion calculation (e.g., to avoid increasing memory bandwidth).
[0099] As an example, if the current sub-block / block (e.g., 16×16) has only one additional row to the left and above within the prefetch region and has two or more additional rows to the right and bottom within the prefetch region, an asymmetric extension can be used. In this example, one row at the top / left and two rows at the right / bottom are additionally added to the candidate predictor, which provides a 19×19 predictor for bilateral matching. Then the SAD value can be normalized by the number of samples. In another example, an asymmetric extension is used for the additional rows, but the SAD is not normalized for the distortion calculation. In another example, an asymmetric extension is used for the additional rows, but the SAD is not normalized for the distortion calculation. Instead, samples outside the prefetch region are padded to calculate the SAD.
[0100] In some embodiments, when deriving the bilateral matching cost, a different interpolation filter is used for samples outside the block / sub-block region compared to the interpolation filter used for samples within the block / sub-block region.
[0101] In some embodiments, the interpolation filter used for samples within the block / sub-block region has longer filter taps than the interpolation filter used for samples outside the block / sub-block region.
[0102] In some embodiments, the bilateral matching distortion calculation is based on the processed initial predictor rather than the original predictor at the sub-block or the entire block. In some embodiments, the initial predictor is subsampled to reduce the complexity of the bilateral matching distortion calculation. Using the processed predictor for the bilateral matching distortion calculation reduces the distortion calculation complexity. As an example, horizontal subsampling is used, where every N horizontal rows are used for the distortion calculation in the bilateral matching. In another example, vertical subsampling is used. In horizontal subsampling and / or vertical subsampling, every N rows are used for the distortion calculation, and N is equal to or greater than 2. In various embodiments, N is predefined or signaled at the high-level syntax. As an example, when N is equal to 2, only even (or odd) indexed rows are used for the distortion calculation in the bilateral matching.
[0103] In some embodiments, progressive subsampling is used. For example, for the central part of the sub-block or the entire block, no downsampling is applied. For other parts of the sub-block or the entire block, row-based subsampling or interleaved subsampling is used. In some embodiments, interleaved (checkerboard pattern) subsampling is used. For example, if (i + j) % N == 1 (or if (i + j) % N == 0), the sample is not used for the distortion calculation, where i and j are the sample coordinates. In various embodiments, N is predefined or signaled at the high-level syntax.
[0104] In some embodiments, interleaved (checkerboard pattern) subsampling is used (e.g., as Figure 4Cas shown). For example, in even rows, even-indexed samples are used (or not used), while in odd rows, odd-indexed samples are used (or not used). As an example, if (i + j) % 2 == 1, the sample is not used for distortion calculation, where i and j are sample coordinates (e.g., corresponding to N equal to 2). In some embodiments, a filter-based subsampling method is used.
[0105] In some embodiments, multiple subsampling techniques are supported, and the choice of subsampling technique is signaled in a high-level syntax, which includes but is not limited to flags at the sequence level, picture level, sub-picture level, slice level, tile level, and maximum coding block level. In some embodiments, multiple subsampling techniques are supported, and the choice of subsampling technique is signaled in a block-level syntax. In some embodiments, multiple subsampling techniques are supported, and the choice of subsampling technique is implicitly derived based on coding information or any information known to both the encoder and the decoder, such as block size and / or block shape.
[0106] In some embodiments, multiple rounds of bilateral matching can be performed, and different subsampling methods are performed for different rounds of bilateral matching. As an example, for the initial round of bilateral matching, a subsampled predictor is used to perform bilateral matching, and N motion vector candidates with the lowest distortion (e.g., measured by the bilateral matching cost) are identified, and then another round of bilateral matching is performed among these N motion vector candidates based on a predictor without subsampling to derive the finally derived motion vector. As another example, for the initial round of bilateral matching, a more coarsely (e.g., using fewer samples in bilateral matching) subsampled predictor is used to perform bilateral matching, and N motion vector candidates with the lowest distortion (e.g., measured by the bilateral matching cost) are identified, and then another round of bilateral matching is performed among these N motion vector candidates based on a more finely (e.g., using more samples in bilateral matching) subsampled predictor to derive the final motion vector.
[0107] In some embodiments, the initial predictor is filtered before bilateral matching to achieve better matching. As an example, one or more low-pass filters (e.g., Gaussian filters) are applied on the initial predictor before bilateral matching (e.g., to reduce noise). In another example, a gradient filter (e.g., Sobel filter) is applied on the initial predictor before bilateral matching (e.g., to improve the accuracy of matching).
[0108] For example, when calculating the refined MV of a sub-block for optical flow refinement, additional samples and gradient values outside the current sub-block (rows and / or columns) are involved in deriving the refined MV of the sub-block, while the block size for the final motion compensation based on the optical flow remains unchanged. As an example, for a current sub-block of 4×4 or 8×8, where the current block is 16×16, 2 rows are extended to use the samples around the current block in the reference block. In this way, more samples are available for optical refinement, which improves the accuracy of the motion vector without additional hardware (e.g., memory) cost.
[0109] In some embodiments, additional rows and columns beyond the number N of the current sub-block are extended for optical flow-based sub-block MV refinement. The N additional rows and columns beyond the current sub-block can be adjacent or non-adjacent samples located at the top, bottom, left, and / or right of the current sub-block.
[0110] In some embodiments, before calculating the optical flow-based refinement of the sub-block, the interpolated predictors P0 and P1 and the corresponding gradients are padded (e.g., to keep the worst-case memory bandwidth the same as the current design). As an example, if the current block size is 16×16, the online size is limited to 16 + 7, and if the samples exceed 16 + 7, the extra samples are padded to avoid exceeding the limit. In one example, if N is greater than one sample (e.g., width or height), N - 1 samples are padded after obtaining the interpolated P0 and P1. For example, the padding can be performed at the entire block level. As another example, the padding can be performed at the sub-block level.
[0111] In some embodiments, a single N-tap interpolation filter is used instead of an 8-tap interpolation filter to generate the initial interpolated P0 and P1 and the corresponding gradient values (e.g., to keep the worst-case memory bandwidth the same as the current design), such that up to 3 rows / columns can be extended for optical flow-based sub-block MV refinement. For example, N can be less than or equal to 3. For example, N is set to 2, and the interpolation filter is a bilinear interpolation filter.
[0112] In some embodiments, a rounded integer initial MV is used to generate the initial interpolated P0 and P1 and the corresponding gradient values to avoid initial interpolation (e.g., to keep the worst-case memory bandwidth the same as the current design). In this way, up to N rows / columns can be extended for optical flow-based sub-block MV refinement. For example, N can be less than or equal to 3. In some embodiments, to generate the initial interpolated P0 and P1, a lower fractional precision MV such as a half-pixel or a quarter-pixel is used.
[0113] In some embodiments, initial interpolation is performed at the sub-block level, and depending on the sub-block position, different interpolation filters are selected (e.g., to keep the worst-case memory bandwidth the same as current). For example, if the sub-block is at the whole-block boundary, a 2-tap interpolation filter is used, while if the sub-block is entirely within the whole block without adjacent boundaries, an 8-tap interpolation filter is used.
[0114] In some embodiments, for different samples within a sub-block, the selection of the interpolation filter is different. For example, a shorter-tap interpolation filter (e.g., 2-tap) is used to perform initial interpolation on samples located at the boundary of the associated sub-block, and a longer-tap interpolation filter (e.g., 8-tap) is used to perform initial interpolation on samples located at the boundary of the associated sub-block.
[0115] In some embodiments, an asymmetric number of rows / columns is used to extend the optical-flow based refinement of the sub-block according to, for example, the position of the sub-block. For example, if the sub-block is at the whole-block boundary, 1 row (e.g., a horizontal or vertical row) is extended, while if the sub-block is entirely within the whole block without adjacent boundaries, 2 rows are extended. In some embodiments, N rows are extended, where N is a non-negative number.
[0116] In some embodiments, the use or enabling of the above embodiments is controlled by high-level syntax, which includes but is not limited to flags at the sequence level, picture level, sub-picture level, slice level, tile level, and maximum coding block level.
[0117] Figure 5B is a flowchart showing a method 550 for encoding video according to some embodiments. Method 550 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and a memory storing instructions for execution by the control circuitry. In some embodiments, method 550 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.
[0118] The system receives (552) video data including a plurality of blocks, the plurality of blocks including a current block. The system derives (554) a set of sub-block motion vectors for a current sub-block of the current block among the plurality of blocks. The system derives (556) a set of refined sub-block motion vectors for the current sub-block using bilateral matching based on one or more samples external to the current sub-block. The system encodes (558) the current sub-block using the derived set of refined sub-block motion vectors. As previously described, the encoding process may mirror the decoding process described herein. For the sake of brevity, these details are not repeated here.
[0119] Although Figure 5A and Figure 5BA number of logical stages are shown in a particular order, but stages that do not depend on the order can be reordered, and other stages can be combined or split. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the orderings and groupings presented herein are not exhaustive. Additionally, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.
[0120] Turning now to some example embodiments.
[0121] (A1) In one aspect, some embodiments include a method of video decoding (e.g., method 500). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a plurality of blocks; (ii) deriving a set of sub-block motion vectors for a current sub-block of the plurality of blocks; (iii) using bilateral matching based on one or more samples external to the current sub-block to derive a set of refined sub-block motion vectors for the current sub-block; and (iv) using the derived set of refined sub-block motion vectors to reconstruct the current sub-block. For example, samples external to the current block / sub-block are used to process the refined bilateral matching process to improve matching accuracy. As an example, decoder-side motion vector refinement uses bilateral matching (where a larger block provides a better match).
[0122] (A2) In some embodiments according to A1, the bilateral matching uses a first set of samples larger than a set of samples in the current sub-block. For example, the size of the samples for the refined bilateral matching process is greater than the current block / sub-block. That is, the current block / sub-block is a subset of the samples used to perform the bilateral matching.
[0123] (A3) In some embodiments according to A1 or A2, the bilateral matching uses a set of samples including one or more samples within the current sub-block and one or more samples external to the current sub-block. For example, some samples within the current block / sub-block are used together with selected samples external to the current block / sub-block to perform the refined bilateral matching process.
[0124] (A4) In some embodiments according to A3, one or more samples from within the current sub-block are obtained by subsampling the current sub-block. In some embodiments, the subsampled sub-block is derived by subsampling the current sub-block, and one or more samples from within the current sub-block are obtained from the subsampled sub-block. For example, when expanding N rows, a subsampled block and a reference region are used when performing bilateral matching.
[0125] (A5) In some embodiments according to A3 or A4, the method further includes: (i) applying a first interpolation filter to one or more samples within the current sub-block; and (ii) applying a second interpolation filter, different from the first interpolation filter, to one or more samples outside the current sub-block. For example, when deriving the bilateral matching cost, a different interpolation filter can be used for samples outside the block / sub-block region compared to the interpolation filter used for samples within the block / sub-block region. As an example, the interpolation filter for samples within the block / sub-block region has longer filter taps than the interpolation filter for samples outside the block / sub-block region.
[0126] (A6) In some embodiments according to any one of A1 to A5, one or more samples outside the current sub-block are selected from a set of N adjacent rows of the current sub-block, where N is a positive integer. For example, the additional N rows of samples around the current block / sub-block are used for the distortion calculation of bilateral matching. For example, N can be equal to 2, and the DMVR search range is the block size plus 2 pixels on each side. In some embodiments, the memory allocated in the decoder (e.g., buffer size) is set to the block size + 2 for the bilateral matching operation.
[0127] (A7) In some embodiments according to A6, one or more samples are used for the distortion calculation of bilateral matching. For example, if the current sub-block size for refinement is 16×16 and N is equal to 2. In this example, an 18×18 sample sub-block of two candidate predictors is interpolated, and the distortion between these two blocks is calculated and used for distortion comparison. The distortion calculation method and the interpolation filter can use the current SAD and the bilinear filter.
[0128] (A8) In some embodiments according to A6 or A7, N is equal to a preset value. For example, the value of N is hard-coded (e.g., set to be equal to 2). In another example, N is set to 1 (e.g., to limit the number of prefetch samples from off-chip memory to on-chip memory).
[0129] (A9) In some embodiments according to A6 or A7, the value of N is signaled in the video bitstream. For example, the number N is pre-selected and signaled in the bitstream at the frame / sequence level or other high levels (e.g., N is switchable at the frame / sequence level or other high levels).
[0130] (A10)In some embodiments according to any one of A6 to A9, a set of N adjacent rows includes the rows above and / or below the current sub-block. For example, bilateral matching is performed using only the additional N rows above the top block / sub-block boundary and / or the additional N rows below the bottom block / sub-block boundary. In this example, additional rows beyond the left block and / or the right block / boundary are not used to perform bilateral matching.
[0131] (A11)In some embodiments according to any one of A6 to A10, a set of N adjacent rows is asymmetric with respect to the current sub-block. For example, samples of the additional rows around the current block / sub-block are used for distortion calculation in bilateral matching. In this example, asymmetric extension is used for the additional rows, and normalized SAD is used for distortion calculation (e.g., to avoid increasing memory bandwidth).
[0132] (A12)In some embodiments according to A11, a set of N adjacent rows includes a first number of rows from above or to the left of the current sub-block and a second number of rows from below or to the right of the current sub-block, the second number being different from the first number. For example, in one case, the current sub-block / block (e.g., 16×16) has only one additional row to the left and above within the prefetch region, and two or more additional rows to the right and below within the prefetch region. In this case, one row at the top / left and two rows at the right / bottom are added to the candidate predictor, which provides a 19×19 predictor for bilateral matching. In this example, the SAD value is normalized by the number of samples. In another example, asymmetric extension is used for the additional rows, but the SAD is not normalized for distortion calculation. In another example, asymmetric extension is used for the additional rows, the SAD is not normalized for distortion calculation, but samples outside the prefetch region are padded to calculate the SAD.
[0133] (A13)In some embodiments according to any one of A6 to A12, the method further includes: (i) determining the sample position of a sample in one or more samples; (ii) obtaining a padding value of the sample when the sample position is outside a predefined acquisition region; and (iii) obtaining a value from the sample position when the sample position is inside a predefined acquisition region. For example, samples of the additional N rows around the current block / sub-block are used for distortion calculation in bilateral matching. When the additional rows fall outside the predefined prefetch region, a padding process is performed (e.g., to avoid increasing memory bandwidth). The padding process can be performed at the whole block level or sub-block level. In some embodiments, the padding value of the sample is obtained according to the determination that the sample position is outside the predefined acquisition region.
[0134] (A14)In some embodiments according to A13, the predefined acquisition region corresponds to the region of the current block plus a set of N samples for an N-tap interpolation technique, where N is a positive integer. For example, the prefetch region is defined as the current entire block size plus seven samples to perform 8-tap interpolation. If the additional N rows of the current candidate block / sub-block pointed to by the current MV + refined MV fall outside the prefetch region, padding is performed on the outside samples.
[0135] (A15)In some embodiments according to A13, the predefined acquisition region corresponds to the region of the current block plus a set of N samples for an N-tap interpolation technique, where N is a positive integer. For example, the prefetch region is defined as the current entire block size plus seven samples to perform 8-tap interpolation. If the additional N rows of the current candidate block / sub-block pointed to by the current MV + refined MV fall outside the prefetch region, padding is performed on the outside samples.
[0136] (A15)In some embodiments according to A13: (i) when the optical flow mode is disabled for the current sub-block, the predefined acquisition region corresponds to the region of the current sub-block plus a set of N samples, where N is a positive integer; and (ii) when the optical flow mode is enabled for the current sub-block, the predefined acquisition region corresponds to the region of the current sub-block plus a set of M samples, where M is a positive integer different from N. For example, in the case where optical flow is enabled, the prefetch region is defined as the current sub-block size plus M samples. The value of M is different from N. In one example, M is set to 9. As another example, DMVR is enabled and then refined (e.g., 0 to 2), and then optical flow can be applied to refine again (e.g., 1), where the total is 1 to 3.
[0137] (B1)On the other hand, some embodiments include a method of video encoding (e.g., method 550). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving video data including a plurality of blocks, the plurality of blocks including a current block; (ii) deriving a set of sub-block motion vectors for a current sub-block of the current block among the plurality of blocks; (iii) using bilateral matching based on one or more samples outside the current sub-block to derive a set of refined sub-block motion vectors for the current sub-block; and (iv) encoding the current sub-block using the derived set of refined sub-block motion vectors.
[0138] (B2)In some embodiments according to B1, the bilateral matching uses a first set of samples larger than a set of samples in the current sub-block.
[0139] (B3) In some embodiments according to B1 or B2, the method further comprises: (i) determining a sample location of a sample in one or more samples; (ii) obtaining a padding value of the sample when the sample location is outside a predefined acquisition region; and (iii) obtaining a value from the sample location when the sample location is inside the predefined acquisition region. In some embodiments, the padding value of the sample is obtained based on the determination that the sample location is outside the predefined acquisition region.
[0140] (C1) In another aspect, some embodiments include a method of processing visual media data. In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method comprises: (i) obtaining a source video sequence; and (ii) performing a conversion between the source video sequence and a bitstream of the visual media data, wherein the bitstream includes a plurality of encoded pictures corresponding to a plurality of pictures, and wherein the plurality of encoded pictures includes one or more encoded sub-blocks encoded using a set of refined sub-block motion vectors, the set of refined sub-block motion vectors being derived using bilateral matching based on one or more samples outside a current sub-block.
[0141] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B3, and C1 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by control circuitry of a computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15, B1 to B5, C1, D1, and E1).
[0142] As used herein, N refers to a variable number. Unless otherwise specified, different instances of N may refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.
[0143] Unless otherwise specified, any of the syntax elements described herein may be High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to the sequence level, the frame level, the slice level, or the tile level. As another example, HLS elements may be signaled in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), a slice header, a picture header, a tile header, and / or a CTU header.
[0144] It will be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are also intended to include the plural forms. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0145] As used herein, the term "if" may be interpreted, depending on the context, as meaning "when the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with the determination that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrase "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" may be interpreted, depending on the context, as meaning "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with the determination that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true".
[0146] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best illustrate the principles of operation and practical application, so as to enable others skilled in the art to implement them.
Claims
1. A method for video decoding performed at a computing system, the computing system having a memory and one or more processors, the method comprising: Receiving a video bitstream including a plurality of blocks; Deriving a set of sub-block motion vectors for a current sub-block of the current block among the plurality of blocks; Deriving a set of refined sub-block motion vectors for the current sub-block using bilateral matching based on one or more samples outside the current sub-block; And Reconstructing the current sub-block using the derived set of refined sub-block motion vectors.
2. The method according to claim 1, wherein, The bilateral matching uses a first set of samples larger than a set of samples in the current sub-block.
3. The method according to claim 1, wherein The bilateral matching uses a set of samples including one or more samples within the current sub-block and the one or more samples outside the current sub-block.
4. The method according to claim 3, wherein The one or more samples within the current sub-block are obtained by subsampling the current sub-block.
5. The method according to claim 3, further comprising: Applying a first interpolation filter to the one or more samples within the current sub-block; And Applying a second interpolation filter, different from the first interpolation filter, to the one or more samples outside the current sub-block.
6. The method according to claim 1, wherein, The one or more samples outside the current sub-block are selected from a set of N adjacent rows of the current sub-block, where N is a positive integer.
7. The method according to claim 6, wherein The one or more samples are used for distortion calculation in the bilateral matching.
8. The method according to claim 6, wherein N is equal to a preset value.
9. The method according to claim 6, wherein The value of N is signaled in the video bitstream.
10. The method according to claim 6, wherein, The set of N adjacent rows includes rows above and / or below the current sub-block.
11. The method according to claim 6, wherein, The set of N adjacent rows is asymmetric with respect to the current sub-block.
12. The method according to claim 11, wherein, The set of N adjacent rows includes a first number of rows from above or to the left of the current sub-block and a second number of rows from below or to the right of the current sub-block, the second number being different from the first number.
13. The method according to claim 1, further comprising: Determining a sample position of a sample among the one or more samples; Obtaining a padding value for the sample when the sample position is outside a predefined acquisition region; And Obtaining a value from the sample position when the sample position is within the predefined acquisition region.
14. The method according to claim 13, wherein, The predefined acquisition region corresponds to the region of the current block plus a set of N samples for an N-tap interpolation technique, where N is a positive integer.
15. The method according to claim 13, wherein, The predefined acquisition region corresponds to the region of the current sub-block plus a set of N samples for an N-tap interpolation technique, where N is a positive integer.
16. The method according to claim 13, wherein: When the optical flow mode is disabled for the current sub-block, the predefined acquisition region corresponds to the region of the current sub-block plus a set of N samples, N being a positive integer; And When the optical flow mode is enabled for the current sub-block, the predefined acquisition region corresponds to the region of the current sub-block plus a set of M samples, M being a positive integer different from N.
17. A computing system, comprising: Control circuitry; Memory; And One or more sets of instructions stored in the memory and configured to be executed by the control circuitry, the one or more sets of instructions including instructions for the following operations: Receiving video data including a plurality of blocks, the plurality of blocks including a current block; Deriving a set of sub-block motion vectors for a current sub-block of the current block; Deriving a set of refined sub-block motion vectors for the current sub-block using bilateral matching based on one or more samples external to the current sub-block; And Encoding the current sub-block using the derived set of refined sub-block motion vectors.
18. The computing system according to claim 17, wherein, The bilateral matching uses a first set of samples larger than a set of samples in the current sub-block.
19. The computing system according to claim 17, wherein The one or more sets of instructions further include instructions for the following operations: Determining a sample position of a sample in the one or more samples; Obtaining a padding value for the sample when the sample position is outside a predefined acquisition region; And Obtaining a value from the sample position when the sample position is inside the predefined acquisition region.
20. A non-transitory computer-readable storage medium storing one or more sets of instructions configured to be executed by a computing device having control circuitry and a memory, the one or more sets of instructions including instructions for the following operations: Obtaining a source video sequence including a plurality of pictures; and Perform the conversion between the source video sequence and the bit stream of the visual media data, wherein, The bitstream includes: A plurality of encoded pictures corresponding to the plurality of pictures, wherein the plurality of encoded pictures include one or more encoded sub-blocks encoded using a set of refined sub-block motion vectors, the set of refined sub-block motion vectors being derived using bilateral matching based on one or more samples external to each current sub-block.