Motion vector prediction with motion vector prediction look ahead / look behind
By using look-ahead and look-back methods to construct motion vector prediction values in video coding, the problem of low efficiency in motion vector prediction in existing technologies is solved, achieving more efficient inter-frame predictive coding and improving the transmission quality and compression effect of video data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2025-01-17
- Publication Date
- 2026-07-24
AI Technical Summary
Existing video coding techniques suffer from inefficiency in motion vector prediction, especially in inter-frame prediction where it is difficult to effectively utilize temporal and spatial redundancy for efficient coding.
The derivation method of motion vector prediction values is improved by adopting a look-ahead and/or look-back approach. Candidate motion vector prediction values are constructed by determining the sum of intermediate vectors and initial vectors among multiple reference images. The look-ahead and look-back techniques are used to construct an MVP candidate list to improve prediction accuracy.
It improves the efficiency of video coding, enhances the coding performance of inter-frame prediction, and improves the quality and compression effect of video data transmission.
Smart Images

Figure CN122460069A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Patent Application No. 19 / 026,158, filed January 16, 2025, entitled "MOTION VECTOR PREDICTOR BY USING MOTION VECTOR PREDICTOR LOOKAHEAD / LOOKBEHIND," which claims priority to U.S. Provisional Application No. 63 / 622,534, filed January 18, 2024, entitled "MOTION VECTOR PREDICTOR BY USING MOTION VECTOR PREDICTOR LOOKAHEAD / LOOKBEHIND." The entire disclosure of the earlier applications is incorporated herein by reference. Technical Field
[0002] This disclosure describes aspects of video coding in general. Background Technology
[0003] The background description provided herein is intended to provide a general overview of the context of this disclosure. Nothing described in this background section, the work of the currently named inventors, or any aspect of this specification that was not prior art at the time of filing is expressly or implied to be considered prior art relative to this disclosure.
[0004] Image / video compression helps transmit image / video data between different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the currently reconstructed image to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can use motion compensation to predict samples in the current image based on previously reconstructed images. Motion compensation can be indicated by motion vectors (MV). Summary of the Invention
[0005] Various aspects of this disclosure include bitstreams, methods and apparatus for video encoding / decoding. In some examples, the apparatus for video encoding / decoding includes processing circuitry.
[0006] According to one aspect of this disclosure, a method for video decoding is provided. In this method, a video stream is received, the video stream including encoded information of a current block in a current image and encoded information of a plurality of reference images in a reference list. A plurality of intermediate vectors associated with one of the plurality of reference images are determined. The plurality of intermediate vectors includes an initial vector and a plurality of intermediate motion vectors (MVs). The initial vector is associated with the current image. Each of the plurality of intermediate MVs is defined between two corresponding reference images in the plurality of reference images. A candidate motion vector prediction value (MVP) for the current block is determined based on the sum of the plurality of intermediate vectors. The current block is reconstructed based on an MVP candidate list including the candidate MVP.
[0007] According to another aspect of the present invention, a video encoding method is provided. In this method, a plurality of intermediate vectors for a current block in a current image are determined. These intermediate vectors are associated with one of a plurality of reference images in a reference list. The plurality of intermediate vectors includes an initial vector and a plurality of intermediate motion vectors (MVs). The initial vector is associated with the current image. Each of the plurality of intermediate MVs is defined between two corresponding reference images in the plurality of reference images. A candidate MVP for the current block is determined based on the sum of the plurality of intermediate vectors. The current block is encoded based on an MVP candidate list including the candidate MVPs.
[0008] According to another aspect of this disclosure, a method for processing visual media data is provided. In this method, the bitstream of the visual media data is processed according to format rules. In an example, the bitstream includes encoding information of a current block in a current image and encoding information of multiple reference images in a reference list for the current image. The format rules specify that multiple intermediate vectors associated with one of the multiple reference images are determined. These multiple intermediate vectors include an initial vector and multiple intermediate vector values (MVs). The initial vector is associated with the current image. Each of the multiple intermediate MVs is defined between two corresponding reference images in the multiple reference images. The format rules specify that a candidate MVP for the current block is determined based on the sum of the multiple intermediate vectors. The format rules specify that the current block is processed based on an MVP candidate list including the candidate MVP.
[0009] This disclosure also provides an apparatus for video decoding. The apparatus for video decoding includes processing circuitry configured to implement any of the described methods for video decoding.
[0010] This disclosure also provides an apparatus for video encoding. The apparatus for video encoding includes processing circuitry configured to implement any of the described methods for video encoding.
[0011] This disclosure also provides a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding / processing.
[0012] The technical solution disclosed herein includes a method and apparatus for improving the derivation of motion vector prediction values (or merging candidates) for a current block by using look-ahead and / or look-back methods. In an example, a video stream is received, the video stream including encoding information of the current block in the current image and encoding information of multiple reference images in a reference list. Multiple intermediate vectors associated with one of the multiple reference images are determined. These multiple intermediate vectors include an initial vector and multiple intermediate MVs. The initial vector is associated with the current image. Each of the multiple intermediate MVs is defined between two corresponding reference images in the multiple reference images. A candidate MVP for the current block is determined based on the sum of the multiple intermediate vectors. The current block is reconstructed based on an MVP candidate list including the candidate MVP. By using look-ahead and / or look-back methods, the derivation of motion vector prediction values (or merging candidates) for the current block is improved. Attached Figure Description
[0013] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein: Figure 1 This is a schematic diagram of an example block diagram of a communication system (100); Figure 2 This is a schematic diagram of an example block diagram of a decoder; Figure 3 This is a schematic diagram of an example block diagram of an encoder; Figure 4 This is a schematic diagram of the first example of constructing look-ahead motion vector predictions for different reference indices in the forward reference list; Figure 5 This is an example diagram illustrating the location of the sports field in the construction of forward motion vector prediction values; Figure 6 This is an example diagram illustrating the construction of a look-ahead motion vector prediction based on the scaling of the relevant pointing motion vector; Figure 7 This is an example diagram illustrating the construction of look-ahead motion vector predictions for different reference indices in the forward reference list; Figure 8 A flowchart outlining some aspects of the decoding process according to this disclosure is shown; Figure 9 A flowchart outlining some aspects of the coding process according to this disclosure is shown; Figure 10 It is a schematic diagram of a computer system based on one aspect. Detailed Implementation
[0014] Figure 1 Block diagrams of some example video processing systems (100) are shown. The video processing system (100) is an application example of the disclosed subject matter, namely a video encoder and video decoder located in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0015] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101) such as a digital camera, which creates, for example, an uncompressed video image stream (102). In one example, the video image stream (102) includes samples captured by a digital camera. The video image stream (102), depicted as a thick line to emphasize its high data volume, may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or encoded video stream), depicted as a thin line to emphasize its lower data volume, may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as... Figure 1 Client subsystems (106) and (108) can access a streaming server (105) to retrieve copies (107) and (109) of encoded video data (104). Client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video stream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed topics are applicable in the context of VVC.
[0016] It should be noted that electronic devices (120) and (130) may include other components (not shown). For example, electronic device (120) may include a video decoder (not shown), and electronic device (130) may also include a video encoder (not shown).
[0017] Figure 2 An example block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., receiving circuitry). The video decoder (210) may be used in place of... Figure 1 The video decoder (110) in the example.
[0018] The receiver (231) may receive, for example, one or more encoded video sequences included in the bitstream that will be decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. Encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive encoded video data and other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective user entities (not depicted). The receiver (231) may separate the encoded video sequences from other data. To prevent network jitter, a buffer (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser (220)"). In some applications, the buffer (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located external to the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be placed externally to the video decoder (210) to, for example, prevent network jitter; additionally, another buffer memory (215) may be placed internally to, for example, handle playback timing. When the receiver (231) receives data from a store / forward device with sufficient bandwidth and controllability or from an isochronous synchronization network, the buffer memory (215) may not be necessary, or the buffer memory (215) may be made smaller. For use on packet-switched networks such as the Internet, the buffer memory (215) may be required; the buffer memory (215) may be relatively large, advantageously having an adaptive size, and may be implemented at least partially in the operating system or in a similar element (not depicted) external to the video decoder (210).
[0019] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. These symbols may include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to it, such as... Figure 2 As shown. The control information used for the presentation device can be in the form of Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be based on video coding techniques or standards and can follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set of at least one subgroup of pixels in the subgroup of pixels in the encoded video sequence for use in the video decoder based on at least one parameter corresponding to a group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0020] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0021] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser (220) through subgroup control information parsed from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser (220) and the various units described below is not depicted.
[0022] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with one another. However, for the purposes of describing the disclosed subject matter, it is appropriate to conceptually subdivide it into multiple functional units as described below.
[0023] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantization transform coefficients as symbols (221) from the parser (220) and control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block containing sample values, which can be input into the aggregator (255).
[0024] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed image, but can use prediction information from a previously reconstructed portion of the current image. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from the current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the prediction information generated by the intra-picture prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0025] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to blocks of inter-frame coding and potential motion compensation. In this case, the motion compensation prediction unit (253) may access the reference image memory (257) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The extraction of prediction samples by the motion compensation prediction unit (253) from the address in the reference image memory (257) may be controlled by motion vectors, which may be available to the motion compensation prediction unit (253) in the form of symbols (221), which may have, for example, X components, Y components, and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (257) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0026] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). The video compression technique may include an in-loop filtering technique controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), which can be used by the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0027] The output of the loop filter unit (256) can be a sample stream that can be output to the presentation device (212) and stored in the reference image memory (257) for future inter-frame image prediction.
[0028] Once fully reconstructed, some of the encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (220)) are identified as reference images, the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0029] The video decoder (210) performs decoding operations according to a predetermined video compression technology or standard, such as ITU-T Recommendation H.265. The encoded video sequence may conform to the syntax specified by the video compression technology or standard in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the configuration file recorded in the video compression technology or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technology or standard as the only tools available under that configuration file. For compliance, it may also be necessary that the complexity of the encoded video sequence be within the limits defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference image size, etc. In some cases, the limitations set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0030] On one hand, the receiver (231) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0031] Figure 3 An example block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the example.
[0032] The video encoder (303) can obtain data from the video source (301) (not...). Figure 3 In one example, an electronic device (320) receives video samples, and a video source (301) can capture video images that will be encoded by a video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0033] A video source (301) can provide a sequence of source video samples to be encoded by a video encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays. Depending on the sampling structure, color space, etc., used, each pixel can include one or more samples. The following focuses on describing the samples.
[0034] According to one aspect, the video encoder (303) can encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding rate is a function of the controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units described below. For clarity, the coupling is not depicted in the figures. Parameters set by the controller (350) may include rate control-related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions related to the video encoder (303) optimized for a particular system design.
[0035] In some respects, the video encoder (303) is configured to operate within an encoding loop. As an oversimplification, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (333) embedded within the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference image memory (334). Since decoding the symbol stream produces bit-accurate results regardless of the decoder's location (local or remote), the contents of the reference image memory (334) also correspond bit-accurately between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction section are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related techniques.
[0036] The operation of the “local” decoder (333) can be combined with, for example, the operations already described above. Figure 2 The video decoder (210) described in detail is the same as the "remote" decoder. However, a further brief reference is provided. Figure 2 Since symbols are available and the entropy encoder (345) and parser (220) are able to encode / decode symbols into encoded video sequences without loss, the entropy decoding portion of the video decoder (210), which includes the buffer (215) and parser (220), may not be fully implemented in the local decoder (333).
[0037] On the one hand, aside from the parsing / entropy decoding present in the decoder, the decoder techniques exist in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter disclosed focuses on decoder operation. The description of the encoder techniques can be simplified, as the encoder techniques are inverses of the fully described decoder techniques. In certain areas, more detailed descriptions are provided below.
[0038] During operation, in some examples, the source encoder (330) may perform motion-compensated predictive coding, which predictively encodes the input image by referencing one or more previously encoded images from the video sequence designated as "reference images." In this way, the encoding engine (332) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0039] The local video decoder (333) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (330). The operation of the encoding engine (332) can advantageously be a lossy process. When the encoded video data can be decoded by the video decoder (333), Figure 3 When decoded at (not shown), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0040] The predictor (335) can perform a prediction search against the encoding engine (332). That is, for a new image to be encoded, the predictor (335) can search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can be used as appropriate prediction references for the new image. The predictor (335) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input image may have prediction references obtained from multiple reference images stored in the reference image memory (334).
[0041] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0042] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, and arithmetic coding.
[0043] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission via a communication channel (360), which may be a hardware / software link to a storage device capable of storing the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0044] The controller (350) manages the operation of the video encoder (303). During encoding, the controller (350) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types: Intra-frame pictures (I-pictures) are pictures that can be encoded and decoded without using any other pictures in the sequence as prediction sources. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures; Predictive images (P-images) can be images that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses a motion vector and a reference index to predict sample values for each block. Bidirectional predictive images (B-images) can be images that can be encoded and decoded using intra-frame or inter-frame prediction, which uses two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can be used to reconstruct a single block using more than two reference images and associated metadata.
[0045] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 samples), and each block is encoded sequentially. These blocks can be predictively coded with reference to other (already coded) blocks, which are determined by the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or blocks of an I-image can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.
[0046] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as ITU-T H.265 Recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations utilizing temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0047] On one hand, the transmitter (340) can transmit additional data while transmitting encoded video. The source encoder (330) can include such data as part of the encoded video sequence. Additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0048] The captured video can be presented as multiple source images (video images) in a time-series format. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In one example, a specific image being encoded / decoded is segmented into blocks; this specific image being encoded / decoded is called the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. The motion vector points to the reference block in the reference image, and when using multiple reference images, the motion vector can have a third dimension that identifies the reference image.
[0049] In some respects, bidirectional prediction techniques can be used for inter-frame image prediction. According to this technique, two reference images are used, such as a first reference image and a second reference image that precede the current image in the video in decoding order (but may be past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. The block can be predicted using a combination of the first and second reference blocks.
[0050] In addition, merging mode techniques can be used for inter-frame image prediction to improve coding efficiency.
[0051] According to some aspects of this disclosure, predictions such as inter-frame picture prediction and intra-frame picture prediction are performed on a block-by-block basis. For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU consists of three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively partitioned into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be partitioned into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. CUs are further partitioned into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one respect, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Using a luma prediction block as an example, a prediction block comprises a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0052] It should be noted that any suitable technology can be used to implement the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210). On one hand, the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210), can be implemented using one or more integrated circuits. On the other hand, the video encoder (103) and video encoder (303), as well as the video decoder (110) and video decoder (210), can be implemented using one or more processors that execute software instructions.
[0053] Various aspects of this disclosure include methods and systems for deriving motion vector predictions using look-ahead and / or look-back techniques.
[0054] Video coding has been widely used in broadcasting, video recording, video streaming, and many other applications. Various emerging video coding standards, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1, have been released and widely adopted in these video applications. For example, hybrid video codecs include multiple coding modules, such as intra-frame prediction, inter-frame prediction, transform coding, quantization, entropy coding, and post-processing loop filters. In inter-frame predictive coding, the final motion vector can be derived based on spatial / temporal information or is the sum of the motion vector difference reported by the signal and the derived or selected motion vector prediction value. Motion vector prediction candidates can be generated based on motion vectors from spatial, temporal, non-adjacent, or historically based coded blocks. This disclosure provides a method for constructing motion vector predictions in inter-frame predictive coding to improve the coding efficiency of inter-frame predictive coding.
[0055] In this disclosure, the first motion vector prediction (MVP) for the current block can be derived from the motion vector of one of the multiple spatially adjacent blocks of the current block. (Regarding the reference index...) RefIdx i When deriving MVP candidates, if the image order count (POC, also known as the display order count) of the MVP candidate is similar to the reference index for the current block, then... RefIdx j If the POC of the reference image is not equal to that of the current block, a scaled motion vector can be used as the MVP. Alternatively, if the POC of the MVP candidate is equal to that of the reference image with index [missing information], then [missing information]. RefIdx i If the POC of the reference image is not equal, then the MVP candidate is unavailable. For example, if the reference block indicated by the MVP candidate is not located at reference index 1... RefIdx i In the reference image, the MVP candidate will not be selected.
[0056] In one aspect of this disclosure, MVP candidates are constructed using a first MVP candidate construction method with lookahead / lookbehind. Figure 4 This demonstrates an example of constructing MVP candidates using a lookahead-based first MVP candidate construction approach within a forward reference list (such as reference list 0). It should be noted that... Figure 4 This is just an example. MVP candidate building can also be done in a backward reference list (such as reference list 1) by using a first MVP candidate building approach with a look-back.
[0057] like Figure 4 As shown, the current block (402) is included in the current picture (404). The current picture (404) has multiple reference pictures in a reference list (such as forward reference list 0), such as reference pictures identified by reference indices 0-2. Reference indices 0-2 can indicate different reference pictures in the reference list or different directions in the reference list. Reference pictures identified by reference indices 0-2 can be adjacent or non-adjacent. For example, reference indices 0 and 1 can indicate two adjacent reference pictures or two non-adjacent reference pictures in the reference list.
[0058] For non-merging mode, the predicted motion vector of the current block (402) mv L0(0) The MVP can be derived based on spatial or temporal information. The derived MVP ( mv L0(0) This can be considered as the initial vector (or initial MVP). The derived MVP ( mv L0(0) The current image (404) can be referenced to reference index 0, for example, to a reference block (406) in the reference image identified by reference index 0. A lookahead MVP for reference index 1 can be achieved by... mv L0(0) Related pointing motion vector (or intermediate motion vector) mv L0(1) The derivation is done through summation. Related pointing motion vectors. mv L0(1) It can be limited to pointing from reference block (406) in reference index 0 to reference block (408) in reference index 1.
[0059] Still referencing Figure 4 In some respects, the lookahead MVP for reference index 1 can be derived using block vectors (BV). The BV can be defined as a pointer from reference block (408) to reference block (410) in reference index 1. Therefore, the derived lookahead MVP for reference index 1 can be equal to: mv L0(0) + mv L0(1) + bv (1) Similarly, for the lookahead MVP of reference index 2, the derived MVP can be obtained by... mv L0(0) ), related recursive pointing motion vector ( mv L0(1) and mv L0(2) ) and block vector bv (1) The derivation is done by summation. In Figure 4 In the example, the relevant pointer is the motion vector. mv L0(2) It can be limited to pointing from reference block (410) in reference index 1 to reference block (412) in reference index 2.
[0060] In one aspect of this disclosure, multiple MVPs can be derived for the current block in different reference pictures. For example, in addition to the direct MVP in the reference picture indicated by reference index 0 ( mv L0(0) In addition, the new propagated lookahead MVP for reference index 1 (e.g., MV(refIdx1)) can be derived as follows: mv L0(0) + mv L0(1) + bv (1) Similarly, a new propagational look-ahead MVP (e.g., MV(refIdx2)) for the indicated reference image is derived as follows: mv L0(0) + mv L0(1) + mv L0(2) + bv (1) .
[0061] In one respect, the relevant pointing motion vectors from reference index i to reference index j can be derived by checking the availability of motion vectors in the corresponding motion field.
[0062] In one aspect, the availability of motion vectors in the motion field is checked by scanning the availability of motion vectors at one or more locations. If multiple locations are checked, the scanning order can be a predefined order. Figure 5 Examples of candidate locations for a sports field are shown. Figure 5 As shown, the relevant pointing motion vectors from reference index 0 to reference index 1 are derived based on the availability of motion vectors at the five motion field locations. mv L0(1) For example, these 5 sports field locations were chosen by the MVP ( mvL0(0) Defined in the reference block (506) indicated by ). MVP ( mv L0(0) () can point from the current block (502) in the current image (504) to the reference block (506) in reference index 0. In Figure 5 In the example, the scan sequence starts from the center of the reference block (506) and checks the four corners of the reference block (506) in turn. The first available motion vector from reference index 0 to reference index 1 is selected, such as the first available unscaled motion vector. mv L0(1) Based on the derived MV ( mv L0(1) ), determine the reference block (508) in reference index 1.
[0063] On the one hand, in the MVP candidate list, the look-ahead / look-back MVP for each reference index is inserted after the available unscaled motion vectors from spatially adjacent coded blocks.
[0064] On the one hand, when the POC (also known as the explicit order count) of an MVP candidate (such as a scaled MVP or a look-ahead / look-back MVP) is compared with the POC for the current block, it has... RefIdx j When the POCs of the selected reference images are not equal, insert the look-ahead / look-back MVP from the selected reference images before the scaled MVP.
[0065] On the one hand, firstly implement Figure 5 Availability check of motion vectors. If no motion vectors are available after scanning all positions, check the availability of block vectors.
[0066] On one hand, a signaling notification will be sent to a flag. This flag can be signaled in high-level syntax (such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture Header, Slice Header) to indicate block vectors (such as... Figure 4 In bv (1) Whether it is used for constructing look-ahead motion vector predictions.
[0067] exist Figure 4 In the example, when the block vector is determined to be used for lookahead MVP construction, the motion vector prediction value for reference index 1 is equal to: mv L0(0) + mv L0(1) + bv (1) .
[0068] exist Figure 4In the example, when it is determined that the block vector is not used for look-ahead MVP construction, the predicted motion vector value for reference index 1 is equal to: mv L0(0) + mv L0(1) .
[0069] On one hand, when the MVP candidate list in the non-merging mode has at least one prospective MVP candidate, template matching reordering is employed. This involves reordering the MVP candidate list in the non-merging mode according to a predefined order (e.g., ascending order) using the template matching costs of each MVP candidate. In one example, the prospective motion vector of the prospective MVP candidate is used as the motion vector during the template matching process.
[0070] On the one hand, Figure 4 The depth of MVP propagation shown is limited.
[0071] exist Figure 4 In the example, when adding BV, the increase in propagation depth is not counted. In one example, when the propagation depth is set to a predefined value (e.g., 1), MV(refIdx1) (e.g., mv L0(0) + mv L0(1) + bv (1) ) is available, but MV(refIdx2) (for example, mv L0(0) + mv L0(1) + mv L0(2) + bv (1) (Not available)
[0072] In the example, when adding BV, the BV is factored into the increase in propagation depth. In one example, when the propagation depth is set to 1, MV(refIdx1) (e.g., mv L0(0) + mv L0(1) + bv (1) (Not available, but) mv L0(0) + mv L0(1) Available.
[0073] In one aspect, when reference index p is located between reference index i and reference index j in the time direction, the scaled correlation pointing motion vector from reference index i to reference index p is derived by using the correlation pointing motion vector from reference index i to reference index j and employing an appropriate time scaling factor. Figure 6 An example of scaling relative to the motion vector is shown. For example... Figure 6 As shown, the scaled correlated pointing motion vectors from reference index 0 to reference index 1 mv L0(0→1) This can be achieved by using derived motion vectors from reference index 0 to reference index 2. mv L0(0→2) The results were derived by applying appropriate time scaling.
[0074] In one aspect of this disclosure, when a spatial / temporal coding block is encoded in a coding mode as (or as) a block vector, a block vector (BV) and an associated pointing motion vector are used to construct an MVP candidate. Figure 7 An example is shown of constructing the proposed look-back / look-forward MVP candidates for reference list 0 and / or reference list 1 by using block vector predictions derived from the conventional MVP candidate construction process and associated pointing motion vectors.
[0075] like Figure 7 As shown, the current block (702) is included in the current image (704). For non-merging mode, the BV is derived based on MVP / BVP construction according to spatial or temporal information. MVP BV MVP A reference block (706) in the current image (704) can be pointed to from the current block (702). BV can be used. MVP and in reference index 0 mv L0(0) This is used to derive the lookahead MVP in the reference list (e.g., the forward reference list). For example, the lookahead MVP at reference index 0 is defined as: mv L0(0) +BV MVP Motion vector mv L0(0) Reference block (706) can be used to point to reference block (708) in the reference image identified by reference index 0. The lookahead MVP for reference index 1 can be achieved by... BV MVP With the relevant pointing motion vector mv L0(0) and mv L0(1) The summation is used to derive the definition. Therefore, the lookahead MVP for reference index 1 is defined as: mv L0(0) +BV MVP +mv L0(1) Block vectors (BV) can also be used to derive the lookahead MVP for reference index 1, and the derived lookahead MVP for reference index 1 can be equal to: BVMVP +mv L0(0) +mv L0(1) +bv (1) BV ( bv (1) Reference block (712) can be pointed to from reference block (710) in reference index 1. The relevant pointing motion vector... mv L0(1) Reference block (708) can be referenced to reference block (710). Similarly, the lookahead MVP for reference index 2 can be referenced to the derived MVP / BVP (e.g., BV). MVP ), and related recursive pointing motion vectors (such as mv L0(0) , mv L0(1) and mv L0(2) ) and block vector bv (1) The result is derived by summation. Therefore, the lookahead MVP for reference index 2 is defined as... BV MVP +mv L0(0) +mv L0(1) +mv L0(2) +bv (1) The relevant pointing motion vector (or intermediate motion vector) mv L0(2) Reference block (712) in reference index 1 can be pointed to reference block (714) in reference index 2.
[0076] In one aspect, MVP candidates based on look-ahead and / or look-back are inserted after available unscaled motion vectors from spatially adjacent coded blocks.
[0077] On the one hand, during the construction of the motion vector predictor, the motion vector predictor is unavailable when the forward motion vector does not point to the target reference image.
[0078] On the one hand, for each reference index, the look-back / look-forward MVP (such as...) Figure 7 The look-ahead / look-back MVP shown is directly inserted into the MVP candidate list after the available unscaled motion vectors from spatially adjacent coded blocks.
[0079] On the one hand, when the POC (also known as the display order count) of an MVP candidate (e.g., a scaled MVP or a look-ahead / look-back MVP) is compared with the current block, having RefIdx jWhen the POCs of the selected reference images are not equal, insert the look-ahead / look-back MVP from the selected reference images before scaling the MVP (e.g., Figure 7 The MVP shown is a forward / backward view.
[0080] In one aspect of this disclosure, when a spatial / temporal coded block is encoded using a coding mode as (or as) a block vector, a merging candidate is constructed using the block vector and its associated pointing motion vector. An example of this construction is shown below. Figure 7 As shown. Figure 7 As shown, block vector predictions derived from the MVP candidate construction process (e.g.) can be used. BV MVP The candidate look-back / look-forward merging is constructed for reference list 0 and / or reference list 1, along with the relevant pointing motion vectors. For non-merging modes, MVP / BVP is derived based on spatial or temporal information. BV MVP The look-ahead merging motion vector can be obtained by analyzing the derived MVP / BVP (e.g., ...). BV MVP ), and related recursive pointing motion vectors (e.g. mv L0(0) , mv L0(1) and mv L0(2) ) and block vectors (e.g., bv (1) It is derived by summing.
[0081] On one hand, the maximum look-ahead / look-back depth is used to constrain the tracking depth of look-ahead / look-back. For example, when the maximum depth, i.e., the look-ahead depth, equals a predefined value (e.g., 1), the combined look-ahead motion vector equals: BV MVP + mv L0(0) .
[0082] On the one hand, when the final look-ahead merging motion vector does not point to a reference image in reference list 0 and / or reference list 1, the merging candidate is unavailable.
[0083] Figure 8A flowchart outlining a process (800) according to one aspect of this disclosure is shown. This process (800) can be used with a video decoder. In various aspects, the process (800) is executed by processing circuitry, such as processing circuitry that performs the functions of video decoder (110), processing circuitry that performs the functions of video decoder (210), etc. In some aspects, the process (800) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (800). The process begins at (S801) and proceeds to (S810).
[0084] At (S810), a video stream is received. The stream includes the encoding information of the current block in the current image and the encoding information of multiple reference images in the reference list.
[0085] At (S820), a plurality of intermediate vectors are determined that are associated with one of a plurality of reference images. The plurality of intermediate vectors includes an initial vector and a plurality of intermediate vectors (MVs). The initial vector is associated with the current image. Each of the plurality of intermediate vectors is defined between two corresponding reference images in the plurality of reference images.
[0086] At (S830), the candidate MVP for the current block is determined based on the sum of multiple intermediate vectors.
[0087] At (S840), the current block is reconstructed based on the list of MVP candidates that includes the candidate MVP.
[0088] On one hand, two corresponding reference images are two corresponding adjacent reference images in the reference list. On the other hand, two corresponding reference images are two corresponding non-adjacent reference images in the reference list.
[0089] On one hand, the initial vector is determined as the first MV from the current block to the first reference block within the first reference picture among the plurality of reference pictures. The first intermediate MV among the plurality of intermediate MVs is determined as the second MV from the first reference block within the first reference picture among the plurality of reference pictures to the second reference block within the second reference picture among the plurality of reference pictures.
[0090] In one aspect, multiple intermediate vectors are determined to include a BV. This BV is defined as extending from a second reference block within a second reference image in a plurality of reference images to a third reference block within the second reference image in a plurality of reference images.
[0091] On one hand, multiple intermediate vectors are determined to include a BV. This BV is defined as the distance from the current block to the first reference block in the current image. The first intermediate MV among the multiple intermediate MVs is determined as the MV from the first reference block in the current image to the second reference block within the first reference image among the multiple reference images.
[0092] In one aspect, the first intermediate MV among multiple intermediate MVs is defined as a transition from a first reference block within a first reference image among multiple reference images to a second reference block within a second reference image among multiple reference images. The second reference block is a prediction block of the first reference block.
[0093] On one hand, multiple candidate motion field locations are determined within a first reference image among multiple reference images. These multiple candidate motion field locations are scanned to determine multiple candidate motion field locations leading to a second reference image among the multiple reference images. A first intermediate motion field (MV) among multiple intermediate MVs is determined as a first candidate MV among the multiple candidate MVs; this first candidate MV is an unscaled MV.
[0094] On one hand, the multiple candidate sports field locations include: the center position of the reference block within the first reference image of the multiple reference images, as indicated by the initial vector, and the four corners of the reference block within the first reference image of the multiple reference images.
[0095] On one hand, an MVP candidate list is constructed based on multiple MVP candidates. These multiple MVP candidates include: (i) unscaled MVPs from spatially adjacent coding blocks of the current block, (ii) candidate MVPs following the unscaled MVP, and (iii) scaled MVPs following the candidate MVPs. The POC associated with a scaled MVP is not equal to the POC of the corresponding reference images of the multiple reference images associated with that candidate MVP. The multiple MVP candidates are then reordered based on their template costs.
[0096] On one hand, the time frame (MV) from a first reference image to a third reference image among a plurality of reference images is determined. The scaled MV is determined by scaling this MV using a time scaling factor. Based on this scaled MV, the MV from the first reference image to a second reference image among the plurality of reference images is derived.
[0097] On one hand, the reference list is one of the forward reference list and the backward reference list relative to the current image.
[0098] On the one hand, the total number of intermediate vectors defined between the current image and a corresponding reference image among multiple reference images associated with multiple intermediate vectors is determined based on the maximum tracking depth.
[0099] Then, the process proceeds to (S899) and ends.
[0100] The procedure (800) can be adjusted as appropriate. Steps in the procedure (800) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0101] Figure 9 A flowchart outlining a process (900) according to one aspect of this disclosure is shown. The process (900) can be used with a video encoder. In various aspects, the process (900) is executed by processing circuitry, such as processing circuitry that performs the functions of a video encoder (103), processing circuitry that performs the functions of a video encoder (303), etc. In some aspects, the process (900) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (900). The process begins at (S901) and proceeds to (S910).
[0102] At (S910), multiple intermediate vectors for the current block in the current image are determined. These multiple intermediate vectors are associated with one of multiple reference images in a reference list. The multiple intermediate vectors include an initial vector and multiple intermediate vector values (MVs). The initial vector is associated with the current image. Each of the multiple intermediate MVs is defined between two corresponding reference images in the multiple reference images.
[0103] At (S920), the candidate MVP for the current block is determined based on the sum of multiple intermediate vectors.
[0104] At (S930), the current block is encoded based on the MVP candidate list that includes the candidate MVP.
[0105] On one hand, the initial vector is determined as the first MV from the current block to the first reference block within the first reference picture among the plurality of reference pictures. The first intermediate MV among the plurality of intermediate MVs is determined as the second MV from the first reference block within the first reference picture among the plurality of reference pictures to the second reference block within the second reference picture among the plurality of reference pictures.
[0106] Multiple intermediate vectors are identified as including a BV, which is from a second reference block within a second reference image in a plurality of reference images to a third reference block within a second reference image in a plurality of reference images.
[0107] Multiple intermediate vectors are identified as including a BV, which is from the current block to the first reference block in the current image. The first intermediate MV among multiple intermediate MVs is identified as the MV from the first reference block in the current image to the second reference block within the first reference image among multiple reference images.
[0108] In one aspect, the first intermediate MV among multiple intermediate MVs is defined as a transition from a first reference block within a first reference image among multiple reference images to a second reference block within a second reference image among multiple reference images. The second reference block is a prediction block of the first reference block.
[0109] In one aspect, multiple candidate motion field locations are determined within a first reference image among multiple reference images. These multiple candidate motion field locations are scanned to determine multiple candidate motion field locations leading to a second reference image among the multiple reference images. A first intermediate motion field location among multiple intermediate motion field locations is determined as a first candidate motion field location among multiple candidate motion field locations; this first candidate motion field location is an unscaled motion field location.
[0110] On one hand, the multiple candidate sports field locations include: the center position of the reference block within the first reference image of the multiple reference images, as indicated by the initial vector, and the four corners of the reference block within the first reference image of the multiple reference images.
[0111] On one hand, an MVP candidate list is constructed based on multiple MVP candidates. These multiple MVP candidates include: (i) unscaled MVPs from spatially adjacent coding blocks of the current block, (ii) candidate MVPs following the unscaled MVPs, and (iii) scaled MVPs following the candidate MVPs. The POC associated with a scaled MVP is not equal to the POC of the corresponding reference image among the multiple reference images associated with that candidate MVP. The multiple MVP candidates are then reordered based on their template costs.
[0112] Then, the process proceeds to (S999) and ends.
[0113] The procedure (900) can be adjusted as appropriate. Steps in the procedure (900) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be used.
[0114] In one aspect, a method for processing visual media data (or video data) is provided, comprising: processing a bitstream of the visual media data according to format rules. For example, the bitstream may be a bitstream decoded / encoded in any of the decoding and / or encoding methods described herein. Format rules may specify one or more constraints on the bitstream and / or one or more processes to be performed by the decoder and / or encoder.
[0115] In the example, the bitstream includes encoding information for the current block within the current image and encoding information for multiple reference images in the reference list. The format rules specify: determining multiple intermediate vectors associated with one of the multiple reference images; these intermediate vectors include an initial vector and multiple intermediate MVs, the initial vector being associated with the current image, and each of the multiple intermediate MVs being defined between two corresponding reference images. The format rules specify: determining a candidate MVP for the current block based on the sum of the multiple intermediate vectors. The format rules specify: processing the current block based on an MVP candidate list that includes this candidate MVP.
[0116] The above-described technology can be implemented as computer software that uses computer-readable instructions and is physically stored in one or more computer-readable media. For example, Figure 10 A computer system (1000) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0117] Computer software can be coded using any suitable machine code or computer language. Any suitable machine code or computer language can be assembled, compiled, linked, or similarly processed to create code containing instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), or through interpretation, microcode execution, etc.
[0118] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0119] Figure 10 The components of the computer system (1000) shown are examples and are not intended to impose any limitation on the scope of use or functionality of computer software implementing the aspects of this disclosure. The configuration of the components should also not be construed as having any dependency or requirement relating to any component or combination of components shown in the exemplary aspects of the computer system (1000).
[0120] The computer system (1000) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movement), audio input (e.g., speech, clapping), visual input (e.g., gestures), and olfactory input (not depicted). The human-machine interface device may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, images captured from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0121] Human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard (1001), mouse (1002), touchpad (1003), touch screen (1010), data glove (not shown), joystick (1005), microphone (1006), scanner (1007), camera (1008).
[0122] The computer system (1000) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of a touchscreen (1010), a data glove (not shown), or a joystick (1005), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (1009), headphones (not depicted)), visual output devices (e.g., screens (1010) including CRT screens, LCD screens, plasma screens, OLED screens, each screen may or may not have touchscreen input functionality, each screen may or may not have tactile feedback functionality, some of which are capable of outputting two-dimensional or more three-dimensional visual outputs through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays and smoke boxes (not depicted), and printers (not depicted).
[0123] The computer system (1000) may also include user-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1020) having media such as CD / DVD (1021), finger drives (1022), removable hard disk drives or solid-state drives (1023), conventional magnetic media such as magnetic tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0124] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.
[0125] The computer system (1000) may also include an interface (1054) leading to one or more communication networks (1055). The network may be, for example, a wireless network, a wired network, or an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a latency-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., cable or wireless wide area digital television networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (1000)) attached to some general-purpose data port or peripheral bus (1049); other network interfaces are typically integrated into the core of the computer system (1000) by being attached to a system bus as described below (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (1000) can use any of these networks to communicate with other entities. Such communication can be one-way receiving (e.g., broadcast television), one-way transmitting (e.g., a CANBus connected to certain CANBus devices), or bidirectional, such as connecting to other computer systems using a local area network (LAN) or wide area network (WAN) digital network. Certain protocols and protocol stacks can be used on each of those networks and network interfaces as described above.
[0126] The human-machine interface devices, user-accessible storage devices and network interfaces mentioned above can be attached to the kernel (1040) of the computer system (1000).
[0127] The core (1040) may include one or more central processing units (CPU) (1041), graphics processing units (GPUs) (1042), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1043), hardware accelerators (1044) for certain tasks, graphics adapters (1050), etc. These devices, as well as read-only memory (ROM) (1045), random access memory (1046), and internal mass storage (1047) such as internal non-user-accessible hard disk drives, SSDs, etc., may be connected via a system bus (1048). In some computer systems, the system bus (1048) may be connected in the form of one or more physical plugs to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1048) or attached to the core's system bus (1048) via a peripheral bus (1049). In one example, a screen (1010) may be connected to a graphics adapter (1050). Peripheral bus architectures include PCI, USB, etc.
[0128] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) can execute certain instructions that can be combined to form the computer code mentioned above. This computer code can be stored in ROM (1045) or RAM (1046). Temporary data can also be stored in RAM (1046), while permanent data can be stored, for example, in internal mass storage (1047). Fast storage and retrieval to any storage device can be achieved by using a cache, which can be closely associated with one or more CPUs (1041), GPUs (1042), mass storage (1047), ROM (1045), RAM (1046), etc.
[0129] Computer-readable media may have computer code thereon that performs various computer-implemented operations. The media and computer code may be media and computer code specifically designed and constructed for the purposes of this disclosure, or the media and computer code may be of a type known and available to those skilled in the art of computer software.
[0130] As an example, and not a limitation, a computer system (1000) having an architecture, particularly a kernel (1040), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with user-accessible mass storage as described above, as well as memory of certain non-transitory kernels (1040), such as internal kernel mass storage (1047) or ROM (1045). Software implementing aspects of this disclosure can be stored in such devices and executed by the kernel (1040). Depending on specific needs, the computer-readable media may include one or more storage devices or chips. The software can cause the kernel (1040), and in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes described herein or specific portions of such processes, including defining data structures stored in RAM (1046) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system provides functionality, through hard-wired or otherwise embodied logic in the circuitry (e.g., the accelerator (1044)), that the circuitry may replace or operate with the software to perform a particular process described herein or to perform a particular portion of the particular process described herein. Where appropriate, references to software may include logic, and vice versa. Where appropriate, references to computer-readable media may include circuitry storing software for execution (e.g., an integrated circuit (IC)), circuitry embodying logic for execution, or both. This disclosure includes any suitable combination of hardware and software.
[0131] As used in this disclosure, "at least one of..." or "one of..." is intended to include any one or a combination of the listed elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B, and to one of A and B, are intended to include either A or B or (A and B). Where applicable, the use of "one of..." does not exclude any combination of the listed elements, for example, when the elements are not mutually exclusive.
[0132] While several examples of various aspects have been described in this disclosure, there are modifications, substitutions, and various equivalent alternatives that fall within the scope of this disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design numerous systems and methods that, while not expressly shown or described herein, embody the principles of this disclosure and thus fall within its spirit and scope.
[0133] The above disclosure also includes the features mentioned below. Features can be combined in various ways, and are not limited to the combinations mentioned below.
[0134] (1) A video decoding method, the method comprising: receiving a video stream, the video stream including encoding information of a current block in a current image and encoding information of a plurality of reference images in a reference list; determining a plurality of intermediate vectors associated with one of the plurality of reference images, the plurality of intermediate vectors including an initial vector and a plurality of intermediate motion vectors (MVs), the initial vector being associated with the current image, each of the plurality of intermediate MVs being defined between two corresponding reference images in the plurality of reference images; determining a candidate motion vector prediction value (MVP) for the current block based on the sum of the plurality of intermediate vectors; and reconstructing the current block based on an MVP candidate list including the candidate MVPs.
[0135] (2) The method according to feature (1), wherein determining the plurality of intermediate vectors includes: determining the initial vector as a first MV from the current block to a first reference block in a first reference image among the plurality of reference images, and determining the first intermediate MV among the plurality of intermediate MVs as a second MV from the first reference block in a first reference image among the plurality of reference images to a second reference block in a second reference image among the plurality of reference images.
[0136] (3) The method according to feature (2), wherein determining the plurality of intermediate vectors includes: determining that the plurality of intermediate vectors include block vectors (BV) from the second reference block in the second reference image of the plurality of reference images to the third reference block in the second reference image of the plurality of reference images.
[0137] (4) The method according to any one of features (1) to (3), wherein determining the plurality of intermediate vectors comprises: determining that the plurality of intermediate vectors include a block vector (BV) from the current block to a first reference block in the current image; and determining the first intermediate MV of the plurality of intermediate MVs as the MV from the first reference block in the current image to a second reference block within the first reference image of the plurality of reference images.
[0138] (5) The method according to any one of features (1) to (4), wherein: the first intermediate MV of the plurality of intermediate MVs is defined as a first reference block in a first reference picture in the plurality of reference pictures to a second reference block in a second reference picture in the plurality of reference pictures, and the second reference block is a prediction block of the first reference block.
[0139] (6) The method according to any one of features (1) to (5), wherein determining the plurality of intermediate vectors further comprises: determining a plurality of candidate motion field positions within a first reference image of the plurality of reference images; scanning the plurality of candidate motion field positions to determine a plurality of candidate motion field positions from the plurality of candidate motion field positions to a second reference image of the plurality of reference images; and determining a first intermediate motion field among the plurality of intermediate motion field positions as a first candidate motion field among the plurality of candidate motion field positions, wherein the first candidate motion field is an unscaled motion field.
[0140] (7) The method according to feature (6), wherein the plurality of candidate sports field locations include: the center position of a reference block in a first reference image of the plurality of reference images, indicated by the initial vector, and the four corners of the reference block in the first reference image of the plurality of reference images.
[0141] (8) The method according to any one of features (1) to (7) further includes: constructing the MVP candidate list based on a plurality of MVP candidates; the plurality of MVP candidates includes: an unscaled MVP from a spatially adjacent coding block of the current block, the candidate MVP located after the unscaled MVP, and a scaled MVP located after the candidate MVP, wherein the POC associated with the scaled MVP is not equal to the POC of the corresponding reference image of the plurality of reference images associated with the candidate MVP; and reordering the plurality of MVP candidates based on the template cost of the plurality of MVP candidates.
[0142] (9) The method according to any one of features (1) to (8), wherein determining the plurality of intermediate vectors further comprises: determining the MV from a first reference image among the plurality of reference images to a third reference image among the plurality of reference images; determining a scaled MV by scaling the MV with a time scaling factor; and deriving the MV from the first reference image among the plurality of reference images to a second reference image among the plurality of reference images based on the scaled MV.
[0143] (10) The method according to any one of features (1) to (9), wherein the reference list is one of a forward reference list and a backward reference list relative to the current image.
[0144] (11) The method according to any one of features (1) to (10), wherein the total number of the plurality of intermediate vectors defined between the current image and the corresponding reference images in the plurality of reference images associated with the plurality of intermediate vectors is determined according to the maximum tracking depth.
[0145] (12) A video coding method, the method comprising: determining a plurality of intermediate vectors for a current block in a current image, the plurality of intermediate vectors being associated with one of a plurality of reference images in a reference list, the plurality of intermediate vectors including an initial vector and a plurality of intermediate motion vectors (MVs), the initial vector being associated with the current image, each of the plurality of intermediate MVs being defined between two corresponding reference images in the plurality of reference images; determining a candidate motion vector prediction value (MVP) for the current block based on the sum of the plurality of intermediate vectors; and encoding the current block based on an MVP candidate list including the candidate MVPs.
[0146] (13) The method according to feature (12), wherein determining the plurality of intermediate vectors includes: determining the initial vector as a first MV from the current block to a first reference block within a first reference image in the plurality of reference images, and determining the first intermediate MV among the plurality of intermediate MVs as a second MV from the first reference block within a first reference image in the plurality of reference images to a second reference block within a second reference image in the plurality of reference images.
[0147] (14) The method according to feature (13), wherein determining the plurality of intermediate vectors includes: determining that the plurality of intermediate vectors include block vectors (BV) from the second reference block in the second reference image of the plurality of reference images to the third reference block in the second reference image of the plurality of reference images.
[0148] (15) The method according to any one of features (12) to (14), wherein determining the plurality of intermediate vectors comprises: determining that the plurality of intermediate vectors include a block vector (BV) from the current block to a first reference block in the current image; and determining a first intermediate MV of the plurality of intermediate MVs as an MV from the first reference block in the current image to a second reference block within a first reference image of the plurality of reference images.
[0149] (16) The method according to any one of features (12) to (15), wherein: the first intermediate MV of the plurality of intermediate MVs is defined as a first reference block in a first reference picture of the plurality of reference pictures to a second reference block in a second reference picture of the plurality of reference pictures, wherein the second reference block is a prediction block of the first reference block.
[0150] (17) The method according to any one of features (12) to (16), wherein determining the plurality of intermediate vectors further comprises: determining a plurality of candidate motion field positions within a first reference image of the plurality of reference images; scanning the plurality of candidate motion field positions to determine a plurality of candidate motion field positions from the plurality of candidate motion field positions to a second reference image of the plurality of reference images; and determining a first intermediate motion field among the plurality of intermediate motion field positions as a first candidate motion field among the plurality of candidate motion field positions, wherein the first candidate motion field is an unscaled motion field.
[0151] (18) The method according to feature (17), wherein the plurality of candidate motion field locations include: the center position of the reference block in the first reference image of the plurality of reference images indicated by the initial vector, and the four corners of the reference block in the first reference image of the plurality of reference images.
[0152] (19) The method according to any one of features (12) to (18) further includes: constructing the MVP candidate list based on a plurality of MVP candidates, the plurality of MVP candidates including unscaled MVPs from spatially adjacent coding blocks of the current block, candidate MVPs following the unscaled MVPs, and scaled MVPs following the candidate MVPs, wherein the POC associated with the scaled MVP is not equal to the POC of the corresponding reference image among the plurality of reference images associated with the candidate MVP; and reordering the plurality of MVP candidates based on the template cost of the plurality of MVP candidates.
[0153] (20) A method for processing visual media data, the method comprising: processing a bitstream of the visual media data according to a format rule. Wherein: the bitstream includes encoding information of a current block in a current image and encoding information of a plurality of reference images in a reference list; the format rule specifies: determining a plurality of intermediate vectors associated with one of the plurality of reference images, the plurality of intermediate vectors including an initial vector and a plurality of intermediate motion vectors (MVs), the initial vector being associated with the current image, each of the plurality of intermediate MVs being defined between two corresponding reference images in the plurality of reference images; determining a candidate motion vector prediction value (MVP) for the current block based on the sum of the plurality of intermediate vectors; and processing the current block based on an MVP candidate list including the candidate MVPs.
[0154] (21) An apparatus for video decoding, comprising processing circuitry configured to perform the method of any one of features (1) to (11).
[0155] (22) An apparatus for video encoding, comprising processing circuitry configured to perform the method of any one of features (12) to (19).
[0156] (23) A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of features (1) to (20).
Claims
1. A video decoding method, the method comprising: Receive a video stream, the video stream including the encoding information of the current block in the current image and the encoding information of multiple reference images in the reference list; Determine a plurality of intermediate vectors associated with one of the plurality of reference images, the plurality of intermediate vectors including an initial vector and a plurality of intermediate motion vectors (MVs), the initial vector being associated with the current image, and each of the plurality of intermediate MVs being defined between two corresponding reference images in the plurality of reference images; The candidate motion vector prediction (MVP) for the current block is determined based on the sum of the multiple intermediate vectors. and The current block is reconstructed based on the MVP candidate list that includes the candidate MVP.
2. The method according to claim 1, wherein, Determining the plurality of intermediate vectors includes: The initial vector is determined as the first MV from the current block to the first reference block within the first reference image among the plurality of reference images, and The first intermediate MV among the plurality of intermediate MVs is determined as the second MV from the first reference block in the first reference image among the plurality of reference images to the second reference block in the second reference image among the plurality of reference images.
3. The method according to claim 2, wherein, Determining the plurality of intermediate vectors includes: The plurality of intermediate vectors are defined as block vectors (BVs) from the second reference block within the second reference image of the plurality of reference images to the third reference block within the second reference image of the plurality of reference images.
4. The method according to any one of claims 1 to 3, wherein, Determining the plurality of intermediate vectors includes: The plurality of intermediate vectors are determined to include a block vector (BV) from the current block to a first reference block in the current image; and The first intermediate MV among the plurality of intermediate MVs is determined as the MV from the first reference block in the current image to the second reference block within the first reference image among the plurality of reference images.
5. The method according to any one of claims 1 to 4, wherein: The first intermediate MV in the plurality of intermediate MVs is defined as the distance from a first reference block within a first reference image in the plurality of reference images to a second reference block within a second reference image in the plurality of reference images, and The second reference block is the prediction block of the first reference block.
6. The method according to any one of claims 1 to 5, wherein, Determining the plurality of intermediate vectors also includes: Determine the locations of multiple candidate sports fields within the first reference image among the multiple reference images; Scanning the plurality of candidate motion field locations to determine a plurality of candidate motion field locations from the plurality of candidate motion field locations to a second reference image among the plurality of reference images; and The first intermediate MV among the plurality of intermediate MVs is determined as the first candidate MV among the plurality of candidate MVs, wherein the first candidate MV is an unscaled MV.
7. The method according to claim 6, wherein, The multiple candidate sports field locations include: The center position of the reference block within the first reference image of the plurality of reference images, indicated by the initial vector, and The four corners of the reference block within the first reference image of the plurality of reference images.
8. The method according to any one of claims 1 to 7, further comprising: The MVP candidate list is constructed based on multiple MVP candidates; The plurality of MVP candidates includes: an unscaled MVP from a spatially adjacent coded block of the current block, the candidate MVP following the unscaled MVP, and a scaled MVP following the candidate MVP; the picture order count (POC) associated with the scaled MVP is not equal to the POC of the corresponding reference image of the plurality of reference images associated with the candidate MVP; and The MVP candidates are reordered based on their template costs.
9. The method according to any one of claims 1 to 8, wherein, Determining the plurality of intermediate vectors also includes: Determine the MV from the first reference image to the third reference image among the plurality of reference images; The scaled MV is determined by scaling the MV using a time scaling factor; and The MV is derived from the first reference image among the plurality of reference images to the second reference image among the plurality of reference images based on the scaled MV.
10. The method according to any one of claims 1 to 9, wherein, The reference list is one of the forward reference list and the backward reference list relative to the current image.
11. The method according to any one of claims 1 to 10, wherein, The total number of intermediate vectors defined between the current image and the corresponding reference images in the plurality of reference images associated with the plurality of intermediate vectors is determined based on the maximum tracking depth.
12. A video encoding method, the method comprising: Determine multiple intermediate vectors for the current block in the current image, the multiple intermediate vectors being associated with one of multiple reference images in a reference list, the multiple intermediate vectors including an initial vector and multiple intermediate motion vectors (MVs), the initial vector being associated with the current image, and each of the multiple intermediate MVs being defined between two corresponding reference images in the multiple reference images; The candidate motion vector prediction (MVP) for the current block is determined based on the sum of the multiple intermediate vectors. and The current block is encoded based on an MVP candidate list that includes the candidate MVP.
13. The method according to claim 12, wherein, Determining the plurality of intermediate vectors includes: The initial vector is determined as the first MV from the current block to the first reference block within the first reference image among the plurality of reference images, and The first intermediate MV among the plurality of intermediate MVs is determined as the second MV from the first reference block in the first reference image among the plurality of reference images to the second reference block in the second reference image among the plurality of reference images.
14. The method according to claim 13, wherein, Determining the plurality of intermediate vectors includes: The plurality of intermediate vectors are defined as block vectors (BVs) from the second reference block within the second reference image of the plurality of reference images to the third reference block within the second reference image of the plurality of reference images.
15. A method for processing visual media data, the method comprising: The bitstream of the visual media data is processed according to format rules, wherein: The bitstream includes the encoding information of the current block in the current image and the encoding information of multiple reference images in the reference list; and The formatting rules specify: Determine a plurality of intermediate vectors associated with one of the plurality of reference images, the plurality of intermediate vectors including an initial vector and a plurality of intermediate motion vectors (MVs), the initial vector being associated with the current image, and each of the plurality of intermediate MVs being defined between two corresponding reference images in the plurality of reference images; The candidate motion vector prediction (MVP) for the current block is determined based on the sum of the plurality of intermediate vectors; and The current block is processed based on an MVP candidate list that includes the candidate MVP.