System and method for adaptive motion vector prediction list construction

By constructing a motion vector predictor list and using encoded information to determine the scanning order, the problem of inaccurate motion vector prediction in inter-frame prediction in the prior art is solved, and the efficiency and accuracy of video decoding are improved.

CN120677703APending Publication Date: 2025-09-19TENCENT AMERICA LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480005369.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-20
Filing Date
2024-05-22
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing video coding technologies have difficulty in effectively utilizing coded information in inter-frame prediction to improve the accuracy and efficiency of motion vector prediction, resulting in poor video decoding results.

Method used

By constructing a motion vector predictor list, the scanning order is determined to improve the accuracy of the motion vector by utilizing the encoded information of the current block such as the motion vector, prediction mode and reference frame index of the neighboring block, including receiving a video bit stream, determining the scanning order, generating a motion vector list, identifying the motion vector predictor and decoding.

Benefits of technology

The efficiency and accuracy of video decoding are improved, and the performance of video encoding is improved by using motion vectors with lower index values ​​with higher probability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677703A_ABST
    Figure CN120677703A_ABST
Patent Text Reader

Abstract

An example method of video coding includes receiving a video bitstream including a plurality of blocks. The method further includes determining a scan order of a motion vector list of a first block of the plurality of blocks based on one or more of: a number of neighboring blocks of the current block having a corresponding temporal motion vector, a number of neighboring blocks of the current block encoded in an inter prediction mode, a mode of the current block, and a reference frame index of the current block. The method further includes generating a motion vector list according to the scan order, and identifying a motion vector predictor of the current block from the motion vector list. The method also includes decoding the current block using the identified motion vector predictor.
Need to check novelty before this filing date? Find Prior Art

Description

Related applications

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 546,909, filed on November 1, 2023, entitled “Adaptive Motion Vector PredictionList Construction,” and is a continuation of and claims priority to U.S. Patent Application No. 18 / 669,375, filed on May 20, 2024, entitled “Systems and Methods for Adaptive Motion Vector PredictionList Construction.” Technical Field

[0002] The disclosed embodiments relate generally to video coding and decoding, including but not limited to systems and methods for using inter-frame prediction modes and constructing motion vector lists. Background Art

[0003] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data across communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by hardware and / or software on the server or electronic / client device providing the cloud service.

[0004] Video coding typically uses prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that exploit the redundancy inherent in video data. Video coding aims to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H (Moving Picture Experts Group-H, MPEG-H) project. ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) and ISO / IEC (International Organization for Standardization / International Electrotechnical Commission, ISO / IEC) released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended as a successor to HEVC. The ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (Alliance for Open MediaVideo 1, AV1) is an open video coding format designed as an alternative to HEVC. On January 8, 2019, Verified Version 1.0.0 was released with Specification Errata 1. Summary of the Invention

[0005] In addition, the present disclosure describes a set of techniques for video (image) compression related to inter-frame prediction modes and derived motion vector predictors. Some embodiments include constructing a motion vector predictor (MVP) list for identifying the MVP of a current video block. The order of the MVP list is important for using more accurate motion vectors (MVs) with lower index values. The scanning order of the motion vector candidates of the MVP list can be used to set the order in the MVP list. Some embodiments include determining the scanning order based on encoded information (such as the prediction mode of neighboring blocks, reference frame information and / or current block attributes). Determining the scanning order of MVs based on encoded information enables more accurate MVs with lower index values ​​to be used with a higher probability, thereby improving the efficiency and accuracy of video decoding.

[0006] According to some embodiments, a method of video decoding includes: (i) receiving a video bitstream comprising a plurality of blocks; (ii) determining a scanning order of a motion vector list of a current block among the plurality of blocks based on one or more of: (a) the number of neighboring blocks of the current block having corresponding temporal motion vectors; (b) the number of neighboring blocks of the current block encoded in an inter-frame prediction mode; (c) the mode of the current block; and (d) a reference frame index of the current block; (iii) generating a motion vector list according to the scanning order; (iv) identifying a motion vector predictor of the current block from the motion vector list; and (v) decoding the current block using the identified motion vector predictor.

[0007] According to some embodiments, a method of video encoding includes: (i) receiving video data including multiple blocks, the multiple blocks including a current block; (ii) determining a scanning order of a motion vector list of the current block based on encoded information; (iii) generating a motion vector list according to the scanning order; (iv) identifying a motion vector of the current block from the motion vector list; and (v) encoding the current block using the identified motion vector.

[0008] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system or other electronic device. The computing system includes a control circuit system and a memory storing one or more groups of instructions. The one or more groups of instructions include instructions for executing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder). According to some embodiments, a non-transient computer-readable storage medium is provided. The non-transient computer-readable storage medium stores one or more groups of instructions for execution by a computing system. The one or more groups of instructions include instructions for executing any of the methods described herein.

[0009] Thus, devices and systems utilizing methods for encoding and decoding video are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for video encoding / decoding. The features and advantages described in the specification are not necessarily all-inclusive, and in particular, some additional features and advantages will be apparent to one of ordinary skill in the art in view of the drawings, description, and claims provided in this disclosure. Furthermore, it should be noted that the language used in this specification has been selected primarily for readability and instructional purposes and has not necessarily been selected to describe or limit the subject matter described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order that the present disclosure may be understood in more detail, a more specific description may be given by reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the drawings illustrate only the relevant features of the present disclosure and are therefore not necessarily to be considered limiting, as the description may allow for other effective features that will be understood by those skilled in the art upon reading the present disclosure.

[0011] Figure 1 is a block diagram illustrating an example communication system in accordance with some implementations.

[0012] Figure 2A is a block diagram illustrating example elements of an encoder component according to some implementations.

[0013] Figure 2B is a block diagram illustrating example elements of a decoder component according to some implementations.

[0014] Figure 3 is a block diagram illustrating an example server system according to some implementations.

[0015] Figures 4A to 4C An example of a motion vector scanning order according to some embodiments is shown.

[0016] Figure 4D Motion vector search points for two reference frames are shown according to some embodiments.

[0017] Figure 4E Example block positions for deriving a temporal motion vector predictor according to some embodiments are shown.

[0018] Figure 4F Example motion vector candidate generation for a single inter-prediction block is shown in accordance with some embodiments.

[0019] Figure 5A Example motion vector candidate pool processing is shown in accordance with some implementations.

[0020] Figure 5B An example motion vector predictor list construction order is shown according to some embodiments.

[0021] Figure 5C An example scanning order for motion vector predictor candidate blocks is shown according to some embodiments.

[0022] Figure 6A An example video decoding process is shown in accordance with some implementations.

[0023] Figure 6B An example video encoding process is shown in accordance with some implementations.

[0024] According to common practice, the various features shown in the drawings are not necessarily drawn to scale, and like reference numerals may be used to denote like features throughout the specification and drawings. DETAILED DESCRIPTION

[0025] The present disclosure describes video / image compression techniques that include inter-frame prediction and motion vector list construction. For example, when deriving the MVP (Motion Vector Predictor) of a current block, the scanning order of the spatial motion vector, the temporal motion vector, and the derived motion vector can depend on encoded information from the video bitstream. The encoded information can include the number of neighboring blocks of the current block with corresponding temporal motion vectors, the number of neighboring blocks of the current block encoded in inter-frame prediction mode, the mode of the current block, and / or the reference frame index of the current block. Determining the scanning order based on the encoded information increases the probability of selecting the most appropriate motion vector at a lower index value, which improves the efficiency and accuracy of video decoding. Example systems and devices

[0026] Figure 1 1 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 through 120-m) communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, for example, for use with video-enabled applications such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0027] Source device 102 includes a video source 104 (e.g., a camera component or a media storage device) and an encoder component 106. In some embodiments, video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). Encoder component 106 generates one or more encoded video bitstreams based on the video stream. The video stream from video source 104 can be high in data volume compared to the encoded video bitstream 108 generated by encoder component 106. Because encoded video bitstream 108 is lower in data volume (less data) than the video stream from the video source, encoded video bitstream 108 requires less bandwidth to transmit and less storage space to store than the video stream from video source 104. In some embodiments, source device 102 does not include encoder component 106 (e.g., configured to transmit uncompressed video to network 110).

[0028] One or more networks 110 represent any number of networks that transmit information between source device 102, server system 112, and / or electronic device 120, including, for example, wired and / or wireless communication networks. One or more networks 110 can exchange data using circuit switching channels and / or packet switching channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0029] One or more networks 110 include a server system 112 (e.g., a distributed / cloud computing system). In some embodiments, server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from source device 102). Server system 112 includes a codec component 114 (e.g., configured to encode and / or decode video data). In some embodiments, codec component 114 includes an encoder component and / or a decoder component. In various embodiments, codec component 114 is implemented as hardware, software, or a combination of hardware and software. In some embodiments, codec component 114 is configured to decode encoded video bitstream 108 and re-encode the video data using different encoding standards and / or methods to generate encoded video data 116. In some embodiments, server system 112 is configured to generate multiple video formats and / or encodings based on encoded video bitstream 108. In some embodiments, server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to prune the encoded video bitstream 108 to tailor a potentially different bitstream for one or more of the electronic devices 120. In some implementations, the MANE is provided separately from the server system 112.

[0030] Electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, decoder component 122 is configured to decode encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, one or more of electronic devices 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes media storage). In some embodiments, electronic device 120 is a streaming client. In some embodiments, electronic device 120 is configured to access server system 112 to obtain encoded video data 116.

[0031] The source device and / or the plurality of electronic devices 120 are sometimes referred to as “end devices” or “user devices.” In some embodiments, one or more of the electronic devices 120 and / or the source device 102 are examples of server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptops), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0032] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 can encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and can decode and / or encode the encoded video bitstream 108 using a codec component 114. For example, the server system 112 can apply a coding that is more optimized for network transmission and / or storage to the video data. The server system 112 can transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 can decode the encoded video data 116 and optionally display the video pictures.

[0033] Figure 2A is a block diagram illustrating example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives the video sequence from a remote video source (e.g., a video source that is a component of a device different from encoder component 106). Video source 104 can provide the source video sequence in the form of a digital video sample stream having any suitable bit depth (e.g., 8-bit, 10-bit, or 12-bit), any color space (e.g., BT.601 Y CrCB or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures that are given motion when viewed sequentially. The pictures themselves can be organized as a spatial array of pixels, where each pixel can include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples can be readily understood by those skilled in the art.

[0034] The encoder component 106 is configured to encode and / or compress the pictures of the source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform conversions between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to the other functional units described below. Parameters set by the controller 204 may include rate control-related parameters (e.g., picture skipping, lambda values ​​for quantizers and / or rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller 204 can be readily identified by one of ordinary skill in the art, as these functions may pertain to the encoder component 106 optimized for a particular system design.

[0035] In some embodiments, the encoder component 106 is configured to operate in a codec loop. In a simplified example, the codec loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on the input picture to be encoded and reference pictures) and a (local) decoder 210. The decoder 210 reconstructs the symbols to create sample data in a similar manner to the (remote) decoder (in the case where the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to a reference picture memory 208. Because the decoding of the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference picture memory 208 are also bit-accurate between the local encoder and the remote encoder. In this way, the prediction portion of the encoder interprets the same sample values ​​as the decoder would interpret when using prediction during decoding as reference picture samples.

[0036] The operation of decoder 210 can be combined with a remote decoder such as Figure 2B The operation of the decoder component 122 is the same as that described in detail. However, briefly referring to Figure 2B , since the symbols are available and encoding of the symbols into an encoded video sequence by the entropy encoder 214 and decoding of the symbols by the parser 254 can be lossless, the entropy decoding portion of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0037] In addition to parsing / entropy decoding, the decoder techniques described herein can be present in the corresponding encoder in the form of substantially the same functionality. For this reason, the disclosed subject matter focuses on the decoder operation. Additionally, the description of the encoder techniques can be simplified because the encoder techniques can be reciprocal to the decoder techniques.

[0038] As part of the operation of the source encoder 202, the source encoder 202 can perform motion-compensated predictive coding, which predictively encodes an input frame with reference to one or more previously encoded frames from a video sequence that are designated as reference frames. In this manner, the encoding engine 212 encodes the differences between pixel blocks of the input frame and pixel blocks of a reference frame that can be selected as a prediction reference for the input frame. The controller 204 can manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0039] The decoder 210 decodes the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. Figure 2A When decoded at a remote video decoder (not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. Decoder 210 replicates the decoding process that may be performed on the reference frame by the remote video decoder and may cause the reconstructed reference frame to be stored in reference picture memory 208. In this way, encoder component 106 locally stores a copy of the reconstructed reference frame that has common content (absent transmission errors) with the reconstructed reference frame that will be obtained by the remote video decoder.

[0040] The predictor 206 may perform a prediction search for the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that can be used as appropriate prediction references for the new picture. The predictor 206 may operate on a sample block by pixel block basis to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 208.

[0041] The outputs of all of the above-mentioned functional units may be subjected to entropy encoding in entropy encoder 214. Entropy encoder 214 converts the symbols, as generated by the various functional units, into an encoded video sequence by losslessly compressing them according to techniques known to those skilled in the art (e.g., Huffman encoding, variable length encoding, and / or arithmetic coding).

[0042] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter can be configured to buffer the encoded video sequence, as created by the entropy encoder 214, in preparation for transmission via a communication channel 218, which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter can be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown). In some embodiments, the transmitter can transmit additional data along with the encoded video. The source encoder 202 can include such data as part of the encoded video sequence. The additional data can include temporal / spatial / SNR (Signal-to-Noise Ratio) enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set segments, and the like.

[0043] Controller 204 can manage the operation of encoder component 106. During encoding, controller 204 can assign a specific coded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture can be assigned as an intra picture (I picture), a predictive picture (P picture), or a bidirectional predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are familiar with these variations of I pictures and their corresponding applications and features, and therefore will not be repeated here. Predictive pictures can be encoded and decoded using inter-prediction or intra-prediction, which uses at most one motion vector and reference index to predict sample values ​​for each block. Bidirectional predictive pictures can be encoded and decoded using inter-prediction or intra-prediction, which uses at most two motion vectors and reference indices to predict sample values ​​for each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.

[0044] A source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively) and encoded on a block-by-block basis. These blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded, or they may be predictively coded (spatially or intra-predicted) with reference to already coded blocks of the same picture. Pixel blocks of a P picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be non-predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0045] Video can be captured as multiple source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often referred to simply as intra prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture in encoding / decoding, referred to as the current picture, is divided into blocks. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case of using multiple reference pictures, the motion vector may have a third dimension that identifies the reference picture.

[0046] The encoder component 106 may perform encoding operations according to a predetermined video encoding technique or standard, such as any described herein. In operation of the encoder component 106, the encoder component 106 may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.

[0047] Figure 2B is a block diagram illustrating example elements of decoder component 122 according to some implementations. Figure 2B The decoder component 122 in the embodiment is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data to the display 124 (eg, via a wired connection or a wireless connection).

[0048] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver can be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, each encoded video sequence is decoded independently of the other encoded video sequences. Each encoded video sequence can be received from channel 218, which can be a hardware / software link to a storage device storing the encoded video data. The receiver can receive the encoded video data as well as other data, such as encoded audio data and / or ancillary data streams, which can be forwarded to their respective consuming entities (not depicted). The receiver can separate the encoded video sequence from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data can be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0049] According to some embodiments, the decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensated prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, the decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The decoder component 122 can be implemented at least partially in software.

[0050] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to combat network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 internal to decoder component 122 (e.g., configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to combat network jitter). When receiving data from a store / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be required, or buffer memory 252 may be small. To maximize the use of packet networks such as the Internet, buffer memory 252 may be required. Buffer memory 252 may be relatively large and / or have an adaptive size and may be implemented at least in part in an operating system or similar component external to decoder component 122.

[0051] Parser 254 is configured to reconstruct symbols 270 from the coded video sequence. The symbols may include, for example, information for managing the operation of decoder component 122 and / or information for controlling a rendering device, such as display 124. The control information for the rendering device may be in the form of, for example, a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). Parser 254 parses (entropy decodes) the coded video sequence. The coded video sequence may be encoded according to a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. Parser 254 may extract, from the coded video sequence, a subgroup parameter set for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. Subgroups may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and the like. The parser 254 may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and the like.

[0052] Depending on the type of coded video picture or portion thereof (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of symbol 270 may involve multiple different units. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by parser 254. For clarity, the flow of such subgroup control information between parser 254 and the following multiple units is not depicted.

[0053] The decoder component 122 may be conceptually subdivided into a plurality of functional units, and in some implementations, these units closely interact with each other and may be at least partially integrated with each other. However, for the sake of clarity, the conceptual subdivision of the functional units is maintained herein.

[0054] The scaler / inverse transform unit 258 receives the quantized transform coefficients as symbols 270 from the parser 254, along with control information (e.g., which transform to use, block size, quantization factor, and / or quantization scaling matrix). The scaler / inverse transform unit 258 may output a block comprising sample values, which may be input to an aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may use surrounding reconstructed information obtained from the current (partially reconstructed) picture from the current picture memory 264 to generate a block of the same size and shape as the block being reconstructed. The aggregator 268 may add the prediction information already generated by the intra-picture prediction unit 262 to the output sample information, as provided by the scaler / inverse transform unit 258, on a per-sample basis.

[0055] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and possibly motion-compensated block. In this case, the motion-compensated prediction unit 260 may access the reference picture memory 266 to obtain samples for prediction. After motion compensation is performed on the obtained samples according to the symbols 270 belonging to the block, these samples may be added to the output of the scaler / inverse transform unit 258 (in this case, referred to as residual samples or residual signal) by an aggregator 268 to generate output sample information. The address within the reference picture memory 266 from which the motion-compensated prediction unit 260 obtains the predicted samples may be controlled by a motion vector. The motion vector may be provided to the motion-compensated prediction unit 260 in the form of a symbol 270, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values, such as those obtained from the reference picture memory 266, when using sub-sample accurate motion vectors, and motion vector prediction mechanisms.

[0056] The output samples of the aggregator 268 may be subjected to various loop filtering techniques in the loop filter unit 256. The video compression techniques may include in-loop filtering techniques controlled by parameters included in the coded video bitstream and available to the loop filter unit 256 as symbols 270 from the parser 254, but the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of a coded picture or coded video sequence, as well as to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a rendering device, such as the display 124, and stored in the reference picture memory 266 for use in future inter-picture prediction.

[0057] Once reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture is reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.

[0058] Decoder component 122 can perform decoding operations according to a predetermined video compression technique, which can be documented in a standard, such as any of the standards described herein. As specified in the video compression technique document or standard, and particularly in the profiles therein, the encoded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard. Furthermore, to conform to some video compression techniques or standards, the complexity of the encoded video sequence can be within a range, as defined by the level of the video compression technique or standard. In some cases, the level limits maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further constrained by the Hypothetical Reference Decoder (HRD) specification and metadata used for HRD buffer management signaled in the encoded video sequence.

[0059] Figure 3is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes a control circuit system 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit system 302 includes one or more processors (e.g., a CPU (Central Processing Unit, CPU), a GPU (Graphics Processing Unit, GPU), and / or a DPU (Data Processing Unit, DPU)). In some embodiments, the control circuit system includes a field programmable gate array, a hardware accelerator, and / or an integrated circuit (e.g., an application-specific integrated circuit).

[0060] The network interface 304 can be configured to interface with one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication networks can be local, wide-area, metropolitan area networks, in-vehicle and industrial, real-time, delay-tolerant, and the like. Examples of communication networks include: local area networks such as Ethernet and wireless LANs; cellular networks including GSM (Global System for Mobile Communications), 3G (the Third Generation), 4G (the Fourth Generation), 5G (the Fifth Generation), LTE (Long Term Evolution), and the like; wired or wireless wide-area digital TV networks including cable TV, satellite TV, and terrestrial broadcast TV; in-vehicle and industrial networks including CANBus (Controller Area Network-BUS). Such communications can be one-way receive only (e.g., broadcast TV), one-way send only (e.g., CANBus to certain CANBus devices), or bidirectional (e.g., to other computer systems using local area digital networks or wide area digital networks). Such communications may include communications to one or more cloud computing networks.

[0061] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. The input devices 310 may include one or more of the following: a keyboard, a mouse, a trackpad, a touch screen, a data glove, a joystick, a microphone, a scanner, a camera, etc. The output devices 308 may include one or more of the following: an audio output device (e.g., a speaker), a visual output device (e.g., a display or monitor), etc.

[0062] The memory 314 may include high-speed random access memory (e.g., DRAM (Dynamic Random Access Memory, DRAM), SRAM (Static Random Access Memory, SRAM), DDR RAM (Double Data Rate Random Access Memory, DDR RAM), and / or other random access solid-state memory devices) and / or non-volatile memory (e.g., one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 may optionally include one or more storage devices located remotely from the control circuit system 302. The memory 314, or alternatively, the non-volatile solid-state memory device within the memory 314, includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or a subset or superset thereof: Operating system 316, which includes processes for handling various basic system services and for performing hardware-related tasks; A network communications module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via wired and / or wireless connections); A codec module 320 for performing various functions related to encoding and / or decoding data, such as video data. In some embodiments, the codec module 320 is an instance of the codec component 114. The codec module 320 includes, but is not limited to, one or more of the following: o a decoding module 322 for performing various functions related to decoding encoded data, such as those previously described with respect to the decoder component 122; and o an encoding module 340 for performing various functions related to encoding data, such as those previously described with respect to encoder component 106; and A picture memory 352 for storing pictures and picture data, for example for use with the codec module 320. In some embodiments, the picture memory 352 includes one or more of the following: the reference picture memory 208, the buffer memory 252, the current picture memory 264, and the reference picture memory 266.

[0063] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform the various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform the various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform the various functions previously described with respect to the motion compensated prediction unit 260 and / or the intra-picture prediction unit 262), and a filter module 330 (e.g., configured to perform the various functions previously described with respect to the loop filter 256).

[0064] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform the various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform the various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0065] Each of the modules identified above stored in memory 314 corresponds to an instruction set for performing the functions described herein. The modules identified above (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and therefore various subsets of these modules can be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include a separate decoding module and encoding module, but instead uses the same set of modules to perform two function sets. In some embodiments, memory 314 stores a subset of the modules and data structures identified above. In some embodiments, memory 314 stores additional modules and data structures not described above.

[0066] although Figure 3 The server system 112 is shown according to some embodiments, but Figure 3 It is intended more as a functional description of various features that may be present in one or more server systems rather than as a block diagram of the embodiments described herein. In practice, items shown separately may be combined and some items may be separated. For example, Figure 3 Some items shown individually in the figure may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement server system 112 and how features are distributed among the servers will vary from implementation to implementation and may depend in part on the amount of data traffic that the server system handles during peak usage periods and during average usage periods. Example Codec Technology

[0067] The encoding and decoding processes and techniques described below can be performed at the devices and systems described above (e.g., the source device 102, the server system 112, and / or the electronic device 120). As mentioned above, for inter-frame coded blocks (blocks using inter-frame prediction mode), one or two associated motion vectors are used. These motion vectors can be predicted using a dedicated motion vector predictor, and the difference between the current motion vector and its corresponding predictor can be conveyed within the bitstream. The motion vector predictor can be identified by an index corresponding to an entry in the constructed motion vector prediction list. The motion vector prediction list can be constructed based on motion vectors from spatial neighbors or temporal neighbors. As discussed in more detail below, spatial neighbors include adjacent spatial neighboring blocks (which are direct neighbors to the top and left of the current block) and non-adjacent spatial neighboring blocks (which are close to but not directly adjacent to the current block). The temporal MV (Motion Vector, MV) predictor can be derived using co-located blocks in the reference frame. For example, one way to generate a temporal MV predictor is to store the MVs of reference frames with reference indices associated with the corresponding reference frames, and then the MVs of the reference frames whose trajectories pass through each 8×8 block of the current frame are identified and stored in a temporal MV buffer together with the reference frame index. Thereafter, given a predefined block coordinate, the associated MV stored in the temporal MV buffer is identified and projected onto the current block to derive a temporal MV predictor pointing from the current block to its reference frame.

[0068] Hereinafter, the motion vector library refers to a local buffer that stores a set of motion vectors used by coded blocks. The motion vector library can be used to derive a predictor for the motion vector of the current block. The entries (motion vectors) stored in the motion vector library can be updated during encoding and decoding.

[0069] Figures 4A to 4C An example of a motion vector scanning order according to some embodiments is shown. Figure 4A An example vector scan order for the spatial motion vector predictor is shown. Figure 4A In the example, the spatial neighbor to the left of the current block (e.g., the leftmost bottom block) is scanned first (denoted as "1"), and the spatial neighbor to the left of the current block (e.g., the leftmost top block) is scanned next (denoted as "2"). The scanning order is as follows: Figure 4A Continue as shown until the spatial neighbor to the upper left of the current block is scanned (denoted as "7"). In some embodiments, only Figure 4A A subset of the spatial neighbors shown is scanned. For example, only spatial neighbors coded in inter-frame mode are scanned. In another example, the scan ends before scanning the top left spatial neighbor (e.g., based on one or more decoding settings and / or motion vectors of previously scanned spatial neighbors).

[0070] Figure 4B Another example vector scan order for the spatial motion vector predictor is shown in FIG. Figure 4B In some embodiments, an interleaved adjacent SMV predictor (SMVP) is used to scan the adjacent SMV candidates of the MVP list in the order of the top row, left column, and upper right corner block shown. Figure 4A For example, for the upper adjacent row and left adjacent column (as shown in Figure 4B The scan of the sparse matrix can be simplified to include only candidates 1, 2, 3 and 4 (as shown in Figure 4A ), they are inserted into the MVP list in an interleaved manner.

[0071] Figure 4C Another example vector scanning order for the spatial motion vector predictor of a non-square current block is shown. For example, in the case where the aspect ratio (w / h or h / w) is greater than or equal to 4:1, the middle position candidates along the long side can be additionally inserted into the MVP list (e.g., in Figure 4C In this way, MV candidates can be inserted more efficiently. In some embodiments, the upper left spatial neighbor (in Figure 4A and Figure 4C In some embodiments, the weighted and context modeled counts of the top-left spatial neighbor are unchanged relative to conventional methods.

[0072] In some embodiments, two or more MVP lists are constructed / used. In some embodiments, the MVP list is constructed in the following order: adjacent SMV candidates (e.g., reordered based on weights), TMV (Temporal Motion Vector, TMV) candidates, non-adjacent SMV candidates, derived candidates, and / or additional candidates. In some embodiments, the MVP list is constructed in the following order: adjacent SMV candidates (e.g., reordered based on weights), TMV (Temporal Motion Vector, TMV) candidates, non-adjacent SMV candidates, derived candidates, and / or additional candidates. Figure 5A The candidate (shown) is inserted into the end of the MVP list. In some embodiments, the reference MV library is updated after decoding each super block.

[0073] In some embodiments, the MVP list construction process changes based on the mode of the current block (e.g., one for skip mode and one for other inter-frame prediction modes). In some embodiments, temporal motion vector (TMV) candidates are scanned after adjacent spatial motion vector (SMV) candidates. In some embodiments, TMV candidates are scanned after scanning a subset of SMV candidates (e.g., the first 1, 2, or 3) for skip mode. In some embodiments, when skip mode is activated, TMV candidates are scanned after scanning position 2 SMV candidates. In some embodiments, when skip mode is not activated, the scanning order of TMV candidates is after scanning adjacent SMV candidates.

[0074] In some embodiments, the scanning order of adjacent MVP candidates is staggered (e.g. Figure 4A and Figure 4C In some embodiments, MV library candidates are conditionally inserted before the derived candidates. For example, based on other encoded information, MV library candidates are inserted into the MVP list before or after one or more derived candidates. For example, if the width and height of the current coded block are both less than 16 luma samples, the MV library candidate can be inserted before any derived mode. Otherwise, if the width or height is greater than or equal to 16 luma samples, the MV library candidate can be inserted after any derived mode.

[0075] In some embodiments, the MV library is updated at the block level (e.g., super block level). In some embodiments, the MV library is updated at the coding block level (e.g., rather than the super block level). For example, after decoding each coding block, the corresponding motion information is updated to the MV library.

[0076] Figure 4D , shows motion vector search points for two reference frames according to some embodiments. In some embodiments, the SMVP is derived from spatially neighboring blocks, including adjacent spatially neighboring blocks (which are direct neighbors to the top and left of the current block) and non-adjacent spatially neighboring blocks (which are close to but not directly adjacent to the current block). Figure 4D An example of a set of spatially neighboring blocks of a luma block (eg, where each spatially neighboring block is an 8×8 block) is shown in .

[0077] Spatially neighboring blocks may be checked to find one or more MVs associated with the same reference frame index as the current block. As an example, for the current block, Figure 4DThe numbers 1 to 8 in the indicate the search order for spatially adjacent 8×8 luminance blocks. In some embodiments, fewer spatially adjacent blocks are scanned (e.g., numbers 5 and 7 are skipped). As an example, first, check the upper adjacent row from left to right. Second, check the left adjacent column from top to bottom. Third, check the upper right adjacent block. Fourth, check the upper left adjacent block. Fifth, check the first upper non-adjacent row from left to right. Sixth, check the first left non-adjacent column from top to bottom. Seventh, check the second upper non-adjacent row from left to right. Eighth, check the second left non-adjacent column from top to bottom.

[0078] In some embodiments, adjacent candidates (e.g., Figure 4D 1 to 3 in the MVP list) before any temporal MVP (Temporal Motion Vector Predictor, TMVP) candidate. In some embodiments, non-adjacent (e.g., Figure 4D 4 to 8 in the MV predictor list after one or more TMVP candidates. In this example, all SMVP candidates have the same reference picture as the current block. If the current block has a single reference picture, the MVP candidates with a single reference picture should have the same reference picture. For blocks with composite reference pictures (e.g., 2 reference pictures), one of the reference pictures should be the same reference picture as the current block. If the current block has two reference pictures, only MVP candidates with two identical reference pictures are added to the MVP list.

[0079] Figure 4E Example block positions for deriving a temporal motion vector predictor according to some embodiments are shown. In addition to spatially neighboring blocks, co-located blocks in a reference frame can also be used to derive an MV predictor, which is called a temporal MV predictor. For example, to generate a temporal MV predictor, the MV of a reference frame is stored together with a reference index associated with the corresponding reference frame. Thereafter, for each 8×8 block of the current frame, the MV of the reference frame whose trajectory passes through the 8×8 block is identified and stored in a temporal MV buffer together with the reference frame index. For example, for inter-frame prediction using a single reference frame, regardless of whether the reference frame is a forward reference frame or a backward reference frame, the MV is stored in 8×8 units for performing temporal motion vector prediction for future frames. As another example, for composite inter-frame prediction, only the forward MV is stored in 8×8 units for performing temporal motion vector prediction for future frames.

[0080] In some embodiments, adjacent SMVP candidates, TMVP candidates, and / or non-adjacent SMVP candidates added to the MVP list are reordered. For example, the reordering process can be based on a weight assigned to each candidate. The weights of the candidates can be predefined based on the overlap area of ​​the current block and the candidate block. In some embodiments, the weights of non-adjacent (external) SMVP candidates and TMVP candidates are not considered during the reordering process (e.g., the reordering process only affects adjacent candidates).

[0081] The derived MVP candidates can include both the derived MVPs of the composite mode and the single reference picture. For single inter prediction, if the reference frame of the neighboring block is different from the reference frame of the current block, but they are in the same direction, a time scaling algorithm can be used to scale the MV to the reference frame to form the MVP of the motion vector of the current block. Figure 4E FIGURE 2 shows example motion vector candidate generation for a single inter-prediction block according to some embodiments. Figure 4E As shown, the MVP of the motion vector mv0 of the current block is derived using mv1 from the neighboring block A using time scaling.

[0082] Figure 4F An example motion vector candidate generation for a single inter prediction block according to some embodiments is shown.For composite inter prediction, the MVP of the current block is derived using combined MVs from different neighboring blocks, but the reference frame of the combined MV may need to be the same as the current block. Figure 4F FIG. 4 shows an example motion vector candidate generation for a composite prediction block according to some embodiments. Figure 4F As shown, the combined MVs (mv2, mv3) have the same reference frame as the current block, but come from different neighboring blocks.

[0083] like Figure 5A As shown, some embodiments include a reference motion vector candidate pool. For example, each buffer corresponds to a unique reference frame type, corresponding to a single reference frame or a pair of reference frames, covering single inter mode and composite inter mode, respectively. In some embodiments, all buffers are the same size. In some embodiments, when a new MV is added to a full buffer, the existing MV is evicted to make room for the new MV.

[0084] For example, in addition to the reference MV candidates generated using the previously described reference MV list, the coding block can refer to the MV candidate library to collect reference MV candidates. For example, after encoding a super block, the MV library is updated using the MV used by the coding block of the super block. Each tile can have an independent MV reference library used by all super blocks within the tile. For example, when encoding each tile, the corresponding library is cleared. Thereafter, when encoding each super block within the tile, the MV from the library can be used as an MV reference candidate. When the encoding of the super block is finished, the library is updated.

[0085] Figure 5A FIG. 4 shows an example motion vector candidate library process according to some embodiments. Figure 5A As shown, the library update process can be based on the super block. For example, after encoding the super block, the first (e.g., up to 64) candidate MVs used by each coding block within the super block are added to the library. Pruning can also be involved during the update. In other embodiments, the library update process can be based on the coding block level.

[0086] After performing the reference MV candidate scan as previously described, if there is an open slot in the candidate list, the system can refer to the MV candidate library (e.g., in a buffer with a matching reference frame type) to obtain additional MV candidates. For example, from the end of the buffer backward to the beginning of the buffer, if the MV in the library buffer is not already in the list, the MV in the library buffer is appended to the candidate list.

[0087] Figure 5B An example motion vector predictor list construction order is shown in accordance with some embodiments. Figure 5B In the example of , the MVP list is constructed (e.g., using pruning) in the following order: (i) adjacent SMVP candidates; (ii) reordering process for existing candidates; (iii) TMVP candidates; (iv) non-adjacent SMVP candidates; (v) derived candidates; (vi) additional MVP candidates; and (vii) candidates from the reference MV candidate library. In some embodiments, only one TMVP candidate can be added to the MVP candidate list. For example, once a TMVP candidate is added to the MVP list, the remaining TMVP candidate blocks are skipped. In some embodiments, a reverse horizontal scan order is used to check internal TMVP candidates. Figure 5C An example is shown in , where the scanning order of the inner TMVP candidates is B3->B2->B1->B0. In some embodiments, the outer TMVP candidates are not scanned.

[0088] In some embodiments, skip mode motion information acquisition for spatially neighboring blocks includes both reference picture indices and motion vectors. Information acquisition may include: (i) inserting MVs from adjacent spatially neighboring blocks; (ii) inserting temporal motion vector predictors from reference pictures; (iii) inserting MVs from non-adjacent spatially neighboring blocks; (iv) sorting MV candidates from adjacent spatially neighboring blocks; (v) inserting composite MVs when the existing list size is less than a predetermined threshold (e.g., less than 2 or 3); and / or (vi) inserting MVs from a reference MV library. The MVs and temporal motion vectors from the reference MV library can use pre-selected reference pictures. For example, the pruning process can take into account the reference picture index.

[0089] In some existing processes (e.g., AV2 (AOMedia Video 2, AV2)), the number of slots for MVP candidates is limited to four. In addition, at most one TMVP can be inserted into the MVP list regardless of the MVP list length, which may result in a loss of accuracy in the encoding / decoding process.

[0090] In the following, the term "mode 1" is used to refer to a coding mode that inherits the motion vectors of neighboring blocks, and the term "mode 2" is used to refer to a coding mode that signals the motion vector difference relative to a motion vector predictor selected from spatial or temporal neighboring blocks or a given derived motion vector (e.g., a global motion vector).

[0091] Figure 6A 6 is a flow chart illustrating a method 600 for decoding a video according to some embodiments. The method 600 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, the method 600 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0092] The system receives (602) a video bitstream including a plurality of blocks (e.g., corresponding to a current picture). The system determines (604) a scan order for a motion vector list for a current block of the plurality of blocks based on one or more of the following: the number of neighboring blocks of the current block having corresponding temporal motion vectors, the number of neighboring blocks of the current block encoded in an inter-frame prediction mode, a mode of the current block (e.g., whether the current block uses a first mode in which the current block inherits motion vectors from neighboring blocks or a second mode in which motion vector differences relative to motion vector predictors selected from spatial or temporal neighboring blocks are signaled in the video bitstream), and a reference frame index of the current block. The system generates (606) a motion vector list according to the scan order. The system identifies (607) a motion vector predictor for the current block from the motion vector list. The system decodes (608) the current block using the identified motion vector predictor. For example, when deriving a motion vector predictor for a current block, the scanning order of the spatial motion vector, the temporal motion vector, and the derived motion vector may depend on encoded information from the bitstream, including but not limited to inter-frame prediction modes and reference frame indices from the current block and its neighboring blocks.

[0093] In some embodiments, the temporal motion vector is scanned before all spatial motion vectors for mode 1, and at least one spatial motion vector is scanned before the temporal motion vector for mode 2. For example, the temporal motion vector is scanned before all spatial motion vectors for mode 1 (e.g., a coding mode that inherits the motion vectors of neighboring blocks), and at least one spatial motion vector is scanned before the temporal motion vector for mode 2 (e.g., a coding mode that signals the motion vector difference relative to a motion vector predictor selected from spatial or temporal neighboring blocks). As an example, mode 1 is skip mode, and mode 2 is an inter-frame prediction mode other than skip mode.

[0094] In some embodiments, the temporal motion vector is scanned before at least one of the adjacent spatial motion vectors for mode 1. In some embodiments, the maximum number of temporal motion vectors added depends on the count of spatial neighbors included in the motion vector prediction list. In some embodiments, the maximum number of temporal motion vectors added decreases as the count of spatial neighbors included in the motion vector prediction list increases. In some embodiments, when the count of spatial neighbors included in the motion vector prediction list is less than or equal to a threshold value T1, the maximum number of temporal motion vectors added is N1. Otherwise, when the count of spatial neighbors included in the motion vector prediction list is greater than the threshold value T1, the maximum number of temporal motion vectors added is N2. In one example, both N1 and N2 are positive integers, but N2 is less than N1. In another example, N1 is set to 2 and N2 is set to 1.

[0095] In some embodiments, more than one temporal motion vector is scanned to generate a motion vector prediction list, wherein one of the temporal motion vectors is scanned before at least one neighboring spatial motion vector and at least one of the temporal motion vectors is scanned after all neighboring spatial motion vectors. In some embodiments, the maximum number of added temporal motion vectors depends on the inter-frame prediction mode of the neighboring blocks. In some embodiments, the maximum number of added temporal motion vectors depends on the number of neighboring blocks using the temporal motion vector as a predictor. In one example, if multiple neighboring blocks are using the temporal motion vector as a predictor, the maximum number of added temporal motion vectors can be increased. In some embodiments, the maximum number of added temporal motion vectors depends on the temporal layer ID (Identifier, ID).

[0096] In some embodiments, the motion vectors stored in the library are not scanned for mode 1 to generate a motion vector prediction list, while the motion vectors in the library are scanned for modes other than mode 1. In some embodiments, the derived motion vectors are not scanned for mode 1 to generate a motion vector prediction list, while the motion vectors in the library are scanned for modes other than mode 1. In some embodiments, when deriving a motion vector predictor for a current block, a selection of a scanning order for spatial motion vectors, temporal motion vectors, and derived motion vectors is signaled using a high-level syntax including, but not limited to, a sequence header, a picture header, a sub-picture header, a slice header, and a tile header. In some embodiments, a set of fixed scanning orders is predefined, and an index of the selected scanning order is signaled.

[0097] Figure 6B 6 is a flow chart illustrating a method 650 for encoding a video according to some embodiments. The method 650 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having control circuitry and memory storing instructions for execution by the control circuitry. In some embodiments, the method 650 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0098] The system receives (652) video data comprising a plurality of blocks (e.g., a plurality of blocks corresponding to a current picture), the plurality of blocks including a current block. The system determines (654) a scan order for a motion vector list for the current block based on encoded information. In some embodiments, the encoded information includes one or more of: (i) the number of neighboring blocks of the current block having corresponding temporal motion vectors; (ii) the number of neighboring blocks of the current block encoded in an inter-frame prediction mode; (iii) a first mode for the current block in which the current block inherits one or more motion vectors from the neighboring blocks; and (iv) a second mode for the current block in which a motion vector difference relative to a motion vector predictor selected from a spatial or temporal neighboring block is used. The system generates (656) a motion vector list according to the scan order. The system identifies (658) a motion vector for the current block from the motion vector list. The system encodes (660) the current block using the identified motion vector. As previously described, the encoding process can mirror the decoding process described herein (e.g., motion vector list construction and use). For the sake of brevity, these details are not repeated here.

[0099] although Figure 6A and Figure 6B The various logical stages are shown in a particular order, but stages that are not relevant to the order may be reordered, and other stages may be combined or split. Some reordering or other groupings not specifically mentioned will be apparent to one of ordinary skill in the art, and thus the ordering and groupings presented herein are not exhaustive. Furthermore, it should be appreciated that the stages may be implemented in hardware, firmware, software, or any combination thereof.

[0100] Turning now to some example implementations.

[0101] (A1) In one aspect, some embodiments include a method for video decoding (e.g., method 500). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving a video bitstream (e.g., an encoded video sequence) including a plurality of blocks (e.g., corresponding to one or more pictures); (ii) determining a scanning order for a motion vector list of a first block of the plurality of blocks based on one or more of the following: (a) the number of neighboring blocks of the current block having a corresponding temporal motion vector; (b) the number of neighboring blocks of the current block encoded in an inter-frame prediction mode; (c) the mode of the current block; and (d) the reference frame index of the current block; (iii) generating a motion vector list according to a scanning order; (iv) identifying a motion vector predictor for the current block from the motion vector list; and (v) decoding the current block using the identified motion vector predictor. For example, when deriving a motion vector predictor for a current block, the scanning order of the spatial motion vector, the temporal motion vector, and the derived motion vector may depend on encoded information from the bitstream, including but not limited to inter-frame prediction modes and reference frame indices from the current block and its neighboring blocks.

[0102] (A2) In some embodiments of A1, when the mode of the current block includes inheriting one or more neighboring block motion vectors: (i) the scanning order includes scanning a set of spatial motion vectors; and (ii) The scanning order includes scanning at least one temporal motion vector before scanning all spatial motion vectors in a set of spatial motion vectors. For example, when the current block is in a coding mode that inherits motion vectors of neighboring blocks, the temporal motion vector is scanned before all spatial motion vectors. In some embodiments, based on a determination that the mode of the current block includes inheriting motion vectors of one or more neighboring blocks, the scanning order includes scanning a set of spatial motion vectors; and the scanning order includes scanning at least one temporal motion vector before scanning the entire set of spatial motion vectors. In some embodiments, based on a determination that the mode of the current block includes inheriting motion vectors of one or more neighboring blocks, the scanning order includes scanning a set of spatial motion vectors; and the scanning order includes scanning at least one temporal motion vector before scanning the entire set of spatial motion vectors.

[0103] (A3) In some embodiments of A2, the mode of the current block includes a skip mode.

[0104] (A4) In some embodiments of A2 or A3, at least one temporal motion vector is scanned before at least one adjacent spatial motion vector. For example, when the current block is in a coding mode that inherits the motion vector of a neighboring block, the temporal motion vector is scanned before at least one of the adjacent spatial motion vectors.

[0105] (A5) In some embodiments of any one of A1 to A4, when the mode of the current block includes using motion vector differences: (i) the scanning order includes scanning a set of spatial motion vectors; and (ii) the scanning order includes scanning at least one spatial motion vector before scanning a temporal motion vector. For example, when the current block is in a coding mode that signals a motion vector difference relative to a motion vector predictor selected from spatial or temporal neighboring blocks or a given derived motion vector (e.g., a global motion vector), at least one spatial motion vector is scanned before the temporal motion vector. In some embodiments, based on a determination that the mode of the current block includes using motion vector differences: the scanning order includes scanning a set of spatial motion vectors; and the scanning order includes scanning at least one spatial motion vector before scanning the temporal motion vector.

[0106] (A6) In some embodiments of A4 or A5, the mode of the current block includes a non-skipped inter prediction mode.

[0107] (A7) In some embodiments of any one of A1 to A6, the maximum number of temporal motion vectors in the motion vector list depends on the count of spatial neighbor motion vectors included in the motion vector list. For example, the maximum number of temporal motion vectors added depends on the count of spatial neighbors included in the motion vector prediction list. For example, the maximum number of temporal motion vectors added decreases as the count of spatial neighbors included in the motion vector prediction list increases. As another example, when the count of spatial neighbors included in the motion vector prediction list is less than or equal to a threshold value T1, the maximum number of temporal motion vectors added is N1. Otherwise, when the count of spatial neighbors included in the motion vector prediction list is greater than the threshold value T1, the maximum number of temporal motion vectors added is N2. As an example, both N1 and N2 are positive integers, but N2 is less than N1 (for example, N1=2 and N2=1).

[0108] (A8) In some embodiments of any one of A1 to A7, the maximum number of temporal motion vectors in the motion vector list depends on the count of neighboring blocks with inter prediction mode. For example, the maximum number of temporal motion vectors added depends on the inter prediction mode of the neighboring blocks.

[0109] (A9) In some embodiments of A8, the count of neighboring blocks in inter-frame prediction mode includes the count of neighboring blocks using a temporal motion vector as a predictor. For example, the maximum number of added temporal motion vectors depends on the number of neighboring blocks using a temporal motion vector as a predictor. For example, if multiple neighboring blocks are using a temporal motion vector as a predictor, the maximum number of added temporal motion vectors may be increased.

[0110] (A10) In some embodiments of A8 or A9, the maximum number of temporal motion vectors in the motion vector list also depends on the temporal layer identifier. For example, the maximum number of temporal motion vectors added depends on the temporal layer ID.

[0111] (A11) In some embodiments of any one of A1 to A10, the scanning order comprises: (i) scanning a first temporal motion vector before scanning a spatial motion vector corresponding to a neighboring block; and (ii) scanning a second temporal motion vector after scanning the spatial motion vector. For example, more than one temporal motion vector may be scanned to generate a motion vector prediction list, wherein one of the temporal motion vectors is scanned before at least one neighboring spatial motion vector, and at least one of the temporal motion vectors is scanned after all neighboring spatial motion vectors.

[0112] (A12) In some embodiments of any one of A1 to A11, the method further comprises: (i) when the mode of the current block includes the first mode, the scanning order does not include scanning the motion vector library; and (ii) when the mode of the current block includes the second mode, the scanning order includes scanning the motion vector library. For example, when the mode of the current block includes inheriting one or more neighboring block motion vectors, the motion vectors stored in the library are not scanned to generate the motion vector prediction list, while the motion vectors in the library are scanned for other modes.

[0113] (A13) In some embodiments of any one of A1 to A12, the method further comprises: (i) when the mode of the current block includes the first mode, the scanning order does not include scanning the derived motion vector; and (ii) when the mode of the current block includes the second mode, the scanning order includes scanning the derived motion vector. For example, when the mode of the current block includes inheriting one or more neighboring block motion vectors, the motion vectors stored in the library are not scanned to generate the motion vector prediction list, while the motion vectors in the library are scanned for other modes.

[0114] (A14) In some embodiments of any one of A1 to A13, the scanning order is further based on a signaled indicator from the video bitstream. For example, when deriving a motion vector predictor for the current block, a selection of a scanning order for the spatial motion vector, the temporal motion vector, and the derived motion vector is signaled in a high-level syntax.

[0115] (A15) In some embodiments of A14, the scanning order is selected from a predefined set of scanning orders based on a signaled index. For example, a set of fixed scanning orders is predefined, and the index of the selected scanning order is signaled.

[0116] (B1) In another aspect, some embodiments include a method for video encoding (e.g., method 550). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and one or more processors. In some embodiments, the method is performed at a codec module (e.g., codec module 320). The method includes: (i) receiving video data including a plurality of blocks (e.g., a plurality of blocks corresponding to one or more pictures), the plurality of blocks including a current block; (ii) determining a scanning order for a motion vector list of the current block based on encoded information; (iii) generating a motion vector list according to the scanning order; (iv) identifying a motion vector for the current block from the motion vector list; and (v) encoding the current block using the identified motion vector.

[0117] (B2) In some embodiments of B1, the scanning order includes scanning at least one temporal motion vector before scanning the spatial motion vector.

[0118] (B3) In some embodiments of B1 or B2, the scanning order is based on whether the coding mode of the current block includes using motion vector differences.

[0119] (B4) In some embodiments of any one of B1 to B3, the scanning order is based on whether the coding mode of the current block includes a skip mode.

[0120] In another aspect, some embodiments include a computing system (e.g., server system 112) including a control circuit system (e.g., control circuit system 302) and a memory (e.g., memory 314) coupled to the control circuit system, the memory storing one or more sets of instructions configured to be executed by the control circuit system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15 and B1 to B4 above). In yet another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more sets of instructions for execution by the control circuit system of the computing system, the one or more sets of instructions including instructions for performing any of the methods described herein (e.g., A1 to A15 and B1 to B4 above).

[0121] Unless otherwise specified, any syntax element described herein may be a High-Level Syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS may correspond to a sequence level, a frame level, a slice level, or a tile level. As another example, HLS elements may be signaled in a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), a Picture Parameter Set (PPS), an Adaptation Parameter Set (APS), a slice header, a picture header, a tile header, and / or a CTU header.

[0122] It will be understood that although the terms "first", "second" etc. can be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. The terms used in this article are only for the purpose of describing a specific embodiment and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" as used in this article refer to and encompass any and all possible combinations of one or more associated listed items in the associated listed items. It will also be understood that the terms "including" and / or "comprising" specify the presence of stated features, integers, steps, operations, elements and / or parts when used in this specification, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or their groups.

[0123] As used herein, the term "when" may be interpreted, depending on the context, to mean "if the precondition is true" or "after the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "in response to detecting that the precondition is true". Similarly, the phrase "if it is determined that [the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" may be interpreted, depending on the context, to mean "after determining that the precondition is true" or "in response to determining that the precondition is true" or "in accordance with determining that the precondition is true" or "after detecting that the precondition is true" or "in response to detecting that the precondition is true". As used herein, N refers to a variable number. Unless explicitly stated, different instances of N may refer to the same number (e.g., the same integer value such as the number 2) or different numbers. For purposes of illustration, the foregoing description has been described with reference to specific embodiments. However, the above illustrative discussions are not intended to be exhaustive or to limit the present claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. These embodiments were chosen and described in order to best explain the principles of operation and practical application, thereby enabling others skilled in the art to implement them.

Claims

1. A method of video decoding performed at a computing system having a memory and one or more processors, the method comprising: receiving a video bitstream comprising a plurality of blocks in a current picture; A scanning order of a motion vector list of a current block among the plurality of blocks is determined based on one or more of the following: the number of neighboring blocks of the current block having a corresponding temporal motion vector; The number of neighboring blocks of the current block that are encoded in inter-frame prediction mode; a first mode for the current block, in which the current block inherits one or more motion vectors from a neighboring block; a second mode for the current block, in which a motion vector difference relative to a motion vector predictor selected from spatially or temporally neighboring blocks is signaled in the video bitstream; as well as The reference frame index of the current block; constructing the motion vector list according to the scanning order; identifying a motion vector predictor for the current block from the motion vector list; as well as The current block is decoded using the identified motion vector predictor.

2. The method according to claim 1, wherein: The scanning order includes scanning a set of spatial motion vectors; and The scanning order includes scanning at least one temporal motion vector before scanning all spatial motion vectors in the set of spatial motion vectors.

3. The method according to claim 2, wherein: The first mode of the current block includes a skip mode.

4. The method according to claim 2, wherein: The at least one temporal motion vector is scanned before at least one adjacent spatial motion vector.

5. The method according to claim 1, wherein: The scanning order includes scanning a set of spatial motion vectors; and The scanning order includes scanning at least one spatial motion vector before scanning a temporal motion vector.

6. The method according to claim 4, wherein: The second mode of the current block includes a non-skipped inter prediction mode.

7. The method according to claim 1, wherein The maximum number of temporal motion vectors in the motion vector list depends on the count of spatial neighbor motion vectors included in the motion vector list.

8. The method according to claim 1, wherein The maximum number of temporal motion vectors in the motion vector list depends on the count of neighboring blocks with inter prediction mode.

9. The method according to claim 8, wherein The count of neighboring blocks having the inter prediction mode includes a count of neighboring blocks using a temporal motion vector as a predictor.

10. The method according to claim 8, wherein The maximum number of temporal motion vectors in the motion vector list also depends on the temporal layer identifier.

11. The method according to claim 1, wherein The scanning sequence includes: scanning a first temporal motion vector before scanning spatial motion vectors corresponding to neighboring blocks; and The second temporal motion vector is scanned after scanning the spatial motion vector.

12. The method according to claim 1, further comprising: When the mode of the current block includes the first mode, the scanning order does not include scanning a motion vector library; as well as When the mode of the current block includes the second mode, the scanning order includes scanning the motion vector library.

13. The method according to claim 1, further comprising: When the mode of the current block includes the first mode, the scanning order does not include a scan-derived motion vector; as well as When the mode of the current block includes the second mode, the scanning order includes scanning the derived motion vector.

14. The method according to claim 1, wherein The scanning order is also based on a signaled indicator from the video bitstream.

15. The method according to claim 14, wherein The scanning order is selected from a predefined set of scanning orders based on the signaled index.

16. A computing system comprising: control circuit system; Memory; as well as one or more sets of instructions stored in the memory and configured for execution by the control circuitry, the one or more sets of instructions comprising instructions for: receiving video data comprising a plurality of blocks corresponding to a current picture, the plurality of blocks including a current block; determining a scanning order of a motion vector list of the current block based on the encoded information; generating the motion vector list according to the scanning order; identifying a motion vector of the current block from the motion vector list; as well as The current block is encoded using the identified motion vector.

17. The computing system of claim 16, wherein: The scanning order includes scanning at least one temporal motion vector before scanning the spatial motion vector.

18. The computing system of claim 16, wherein: The scanning order is based on whether a coding mode of the current block includes using motion vector differences.

19. The computing system of claim 16, wherein: The scanning order is based on whether a coding mode of the current block includes a skip mode.

20. A non-transitory computer-readable storage medium storing one or more sets of instructions configured for execution by a computing device having control circuitry and memory, the one or more sets of instructions comprising instructions for: receiving a video bitstream comprising a plurality of blocks; A scanning order of a motion vector list of a current block among the plurality of blocks is determined based on one or more of the following: the number of neighboring blocks of the current block having a corresponding temporal motion vector; The number of neighboring blocks of the current block that are encoded in inter-frame prediction mode; the mode of the current block; and The reference frame index of the current block; generating the motion vector list according to the scanning order; identifying a motion vector predictor for the current block from the motion vector list; as well as The current block is decoded using the identified motion vector predictor.

Citation Information

Cited By

  • Video coding pre-analysis method and device

    CN121567873A