System and method for cross-view motion vector prediction

By adopting a parallax compensation prediction method in multi-view video encoding, the MVP of the first view is derived using the motion vector set of the second view, which solves the problem of redundancy between views and improves the encoding efficiency and decoding quality.

CN120359741APending Publication Date: 2025-07-22TENCENT AMERICA LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202480005458.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2024-05-16
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing video encoding technology fails to effectively utilize statistical redundancy between views in multi-view video encoding, resulting in low encoding efficiency.

Method used

Through the parallax compensation prediction method, the motion vector predictor (MVP) in the multi-view video includes pictures of other views in the same time example, reducing statistical redundancy between views, and using the set of motion vectors with the second view to derive the MVP of the first view for decoding.

Benefits of technology

It improves encoding efficiency, achieves a bit rate saving of about 70%, and improves video decoding accuracy and reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359741A_ABST
    Figure CN120359741A_ABST
Patent Text Reader

Abstract

Various implementations described herein include methods and systems of video encoding. In an aspect, a method includes receiving a multi-view video bitstream, the multi-view video bitstream including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view. The method identifies a first set of reference frames in a first view for a first block. The method acquires, for a second block, a motion vector corresponding to a second set of reference frames in a second view. In accordance with a determination that the first set of reference frames shares a display time with the second set of reference frames, the method derives a motion vector predictor (MVP) of a first block corresponding to the first view using a set of motion vectors corresponding to the second view, and decodes the first block using the derived MVP.
Need to check novelty before this filing date? Find Prior Art

Description

Related Applications

[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 548,368, filed on November 13, 2023, entitled "Cross-View Motion Vector Prediction", which is a continuation-in-part of and claims the benefit of priority of U.S. Patent Application No. 18 / 657,708, filed on May 7, 2024, entitled "Systems and Methods for Cross-View Motion Vector Prediction". Technical Field

[0002] The embodiments disclosed herein generally relate to video coding, including but not limited to systems and methods for motion vector prediction for multiview video (MVV) coding. Background Art

[0003] Digital video is supported by various electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data over a communication network and / or store the digital video data on a storage device. Since the bandwidth capacity of the communication network and the storage resources of the storage device are limited, video coding can be used to compress the video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by hardware and / or software on an electronic device / client device or by a server providing cloud services.

[0004] Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that can exploit the inherent redundancy in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard that is a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (Version 1) and 2022 (Version 2), respectively. AOMedia Video 1 (AV1) is an open video coding format designed to replace HEVC. A verified version 1.0.0 with Errata 1 was released on January 8, 2019. Summary of the Invention

[0005] The present disclosure describes a set of methods for video (image) compression, and more particularly, relates to motion vector prediction when encoding multiple views of a scene. In some embodiments, rather than encoding each view independently and sending a bitstream from each view (simulcast coding), a disparity compensation prediction method is implemented such that pictures of other views are included in a reference picture list at the same time instance. This method, also referred to as disparity compensation prediction, can improve the coding efficiency by reducing the statistical redundancy that exists between different views. In certain instances, the methods disclosed herein can achieve a bitrate savings of approximately 70% compared to simulcast coding.

[0006] According to some embodiments, a method of video decoding includes: (i) receiving a multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) identifying a first set of reference frames for the first block in the first view; (iii) obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; (iv) determining that the first set of reference frames and the second set of reference frames share a display time, and using the set of motion vectors corresponding to the second view to derive a motion vector predictor (MVP) for the first block corresponding to the first view; and (v) decoding the first block using the derived MVP.

[0007] According to some embodiments, a method for video encoding includes: (i) receiving video data including a plurality of blocks, where the plurality of blocks includes a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) identifying a first set of reference frames for the first block in the first view; (iii) obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; (iv) selecting a motion vector for the first block from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames share a display time; and (v) encoding the first block using the selected motion vector.

[0008] According to some embodiments, a method for bitstream conversion includes: (i) obtaining a source video sequence corresponding to a set of views; and (ii) performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, where the multi-view video bitstream includes (a) a first plurality of encoded pictures corresponding to a first view; (b) a second plurality of encoded pictures corresponding to a second view; and (c) an indicator for indicating whether motion vector information from the second plurality of encoded pictures is used to encode one or more blocks from the first plurality of encoded pictures.

[0009] According to some embodiments, a computing system is provided, such as a streaming system, a server system, a personal computer system, or other electronic devices. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).

[0010] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.

[0011] Accordingly, methods, devices, and systems for video encoding and video decoding are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for encoding / decoding video. The features and advantages described in this specification are not necessarily an exhaustive listing. Specifically, according to the accompanying drawings, specification, and claims provided by the present disclosure, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes and is not necessarily for detailed description or limitation of the subject matter described herein. Description of the Drawings

[0012] To understand the present disclosure in more detail, a more specific description can be made with reference to the features of various embodiments, some of which are shown in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and thus are not necessarily considered restrictive, because those skilled in the art will understand after reading the present disclosure that the description may include other effective features.

[0013] Figure 1 is a block diagram showing an example communication system according to some embodiments.

[0014] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.

[0015] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.

[0016] Figure 3 is a block diagram showing an example server system according to some embodiments.

[0017] Figure 4 shows an example MVV according to some embodiments.

[0018] Figures 5 to 7 shows an example operation in the MVV according to some embodiments.

[0019] Figures 8A to 8C shows example prediction blocks, residual blocks, and reconstruction blocks according to some embodiments.

[0020] Figure 9 shows an example in-loop filtering stage according to some embodiments.

[0021] Figure 10A shows an example video decoding process according to some embodiments.

[0022] Figure 10B shows an example video encoding process according to some embodiments.

[0023] By convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and throughout the specification and the drawings, the same reference numerals may be used to denote the same features. Detailed Description

[0024] The present disclosure describes video / image compression techniques, including loop filtering techniques for MVV encoding. The disclosed techniques include using a set of motion vectors corresponding to a second view to derive a motion vector predictor (MVP) for a first block in a first view when it is determined that a first set of reference frames in the first view for the first block shares a display time with a second set of reference frames in a second view for a second block, and using the derived MVP to decode the first block. By using the motion vectors for the second view to derive the MVP for the first view, inter-view redundancy is reduced, and encoding / decoding efficiency can be improved. In addition, using the motion vectors from the second view to derive the MVP for the first view can improve the accuracy of the MVP, thereby improving video decoding accuracy (e.g., improving reconstruction quality). Example systems and devices

[0025] Figure 1 FIG. 6 is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), which are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for supporting applications for video, such as video conferencing applications, digital television applications, and media storage and / or distribution applications.

[0026] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder assembly 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder assembly 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 can be of high data volume compared to the encoded video bitstreams 108 generated by the encoder assembly 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder assembly 106 (e.g., configured to transmit uncompressed video to one (or more) networks 110).

[0027] One or more networks 110 represent any number of networks for communicating information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wireline or wired networks and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0028] One or more networks 110 include the server system 112 (e.g., a distributed computing system / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 may be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0029] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be displayed on a display or other type of display device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or including a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0030] The source device and / or the plurality of electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of: server systems, personal computers, portable devices (e.g., smartphones, tablets, or laptop computers), wearable devices, video conferencing devices, and / or other types of electronic devices.

[0031] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may use the encoder component 114 to decode and / or encode the encoded video bitstream 108. For example, the server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0032] Figure 2A is a block diagram showing example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device for storing previously acquired / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that, when viewed in sequence, are given motion. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can easily understand the relationship between pixels and samples.

[0033] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to the encoder component 106 optimized for a certain system design.

[0034] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simple example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and one (or more) reference pictures) and a (local) decoder 210. The decoder 210 (when the compression between the symbols and the encoded video bitstream is lossless) reconstructs the symbols in a manner similar to a (remote) decoder to create sample data. The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the reference picture samples with the same values as the samples that the decoder will interpret when using the prediction during decoding.

[0035] The operation of the decoder 210 may be the same as that of a remote decoder such as the decoder component 122 described in detail below in conjunction with Figure 2B However, briefly referring to Figure 2B , when the symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols into the encoded video sequence losslessly, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.

[0036] Except for parsing / entropy decoding, the decoder techniques described herein may exist in a corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques may be simplified as encoder techniques may be inverse to decoder techniques.

[0037] As part of the operation of the source encoder 202, the source encoder 202 may perform motion-compensated predictive coding, which predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames in a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of one (or more) reference frames, and the one (or more) reference frames may be selected as one (or more) predictive references for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.

[0038] The decoder 210 decodes the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A (not shown)), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 duplicates the decoding process that may be performed by a remote video decoder on a reference frame and enables the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame, which has the same content as the reconstructed reference frame that will be obtained by the remote video decoder (in the absence of transmission errors).

[0039] The predictor 206 may perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor 206 may operate on the sample blocks on a pixel block-by-pixel basis to find an appropriate prediction reference. As determined from the search results obtained by the predictor 206, the input picture may have prediction references obtained from multiple reference pictures stored in the reference picture memory 208.

[0040] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by various functional units according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding), thereby transforming these symbols into an encoded video sequence.

[0041] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer one (or more) encoded video sequences created by the entropy encoder 214 to prepare for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (source not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal-to-noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0042] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, but this may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Those of ordinary skill in the art are familiar with the variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0043] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., each sample block includes 4×4, 8×8, 4×8, or 16×16 samples), and are encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which is determined by the coding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or the block can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0044] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0045] The encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard such as those described herein. In the operation of the encoder component 106, the encoder component 106 can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0046] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0047] In some embodiments, decoder component 122 includes a receiver that is coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of the other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. This additional data may be included as part of one (or more) of the encoded video sequences. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0048] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.

[0049] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 within decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to prevent network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer 252 may not be required, or a small buffer may be used. For use on a best-effort packet network such as the Internet, buffer memory 252 may be necessary, may be relatively large and / or have an adaptive size, and may be implemented at least partially in an operating system or a similar element external to decoder component 122.

[0050] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a display device (e.g., the display screen 124). The control information for a display device(s) may be in the form of, for example, Supplemental Enhancement Information (SEI) messages or Video Usability Information (VUI) parameter set fragments (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser 254 may extract subgroup parameter sets for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to a group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), and so on. The parser 254 may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0051] Depending on the type of the encoded video picture or a portion of the encoded video picture (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols 270 may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For clarity, such subgroup control information flows between the parser 254 and the multiple units below are not depicted.

[0052] The decoder component 122 may be conceptually subdivided into several functional units, and in some implementations, these units interact closely with each other and may be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is retained herein.

[0053] The Scaler / Inverse Transform Unit 258 receives, from the Parser 254, the quantized transform coefficients as symbols 270, and control information such as which transform mode, block size, quantization factor, and / or quantization scaling matrix to use. The Scaler / Inverse Transform Unit 258 may output a block including sample values, which may be input into the Aggregator 268. In some cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to intra-coded blocks; that is, blocks that do not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the Intra Picture Prediction Unit 262. The Intra Picture Prediction Unit 262 may generate a block having the same size and shape as the block being reconstructed, using surrounding reconstructed information extracted from the current (partially reconstructed) picture in the Current Picture Memory 264. The Aggregator 268 may add, on a per-sample basis, the prediction information generated by the Intra Picture Prediction Unit 262 to the output sample information provided by the Scaler / Inverse Transform Unit 258.

[0054] In other cases, the output samples of the Scaler / Inverse Transform Unit 258 belong to inter-coded and potentially motion-compensated blocks. In such cases, the Motion Compensation Prediction Unit 260 may access the Reference Picture Memory 266 to extract samples for prediction. After motion-compensating the extracted samples according to the symbols 270 belonging to the block, these samples may be added by the Aggregator 268 to the output of the Scaler / Inverse Transform Unit 258 (referred to as residual samples or a residual signal in this case), thereby generating output sample information. The address within the Reference Picture Memory 266 from which the Motion Compensation Prediction Unit 260 extracts prediction samples may be controlled by a motion vector. The motion vector may be in the form of the symbols 270 available to the Motion Compensation Prediction Unit 260, which may have, for example, X, Y, and reference picture components. Motion compensation may also include, for example, interpolation of sample values extracted from the Reference Picture Memory 266, a motion vector prediction mechanism when using sub-sampled accurate motion vectors.

[0055] The output samples of the Aggregator 268 may be used in the Loop Filter Unit 256 for various loop filtering techniques. The video compression technique may include in-loop filter techniques that are controlled by parameters included in the encoded video bitstream, and the parameters may be available to the Loop Filter Unit 256 as symbols 270 from the Parser 254, but may also be in response to meta-information obtained during decoding of a previously (in decoding order) portion of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values. The output of the Loop Filter Unit 256 may be a sample stream, which may be output to a display device such as the Display 124, and stored in the Reference Picture Memory 266 for future inter-picture prediction.

[0056] Once reconstructed, certain coded pictures can be used as reference pictures for future prediction. Once a coded picture has been reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 254), the current reference picture can become part of the reference picture memory 266 and a new current picture memory can be reallocated before starting to reconstruct subsequent coded pictures.

[0057] The decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be documented in any standard such as those described herein. The coded video sequence can conform to the syntax of the video compression technique or standard specified in the video compression technique documentation or standard (in particular, the profile documentation therein). In addition, in order to conform to some video compression techniques or standards, the complexity of the coded video sequence can be within the range defined by the levels of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured, for example, in millions of samples per second), the maximum reference picture size, etc. In some cases, the level-based limitations can be further restricted by assuming a Hypothetical Reference Decoder (HRD) specification and HRD buffer management metadata signaled in the coded video sequence.

[0058] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one (or more) field programmable gate arrays (FPGA), hardware accelerators, and / or one (or more) integrated circuits (e.g., application specific integrated circuit).

[0059] One or more network interfaces 304 may be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication networks may be local, wide area, metropolitan area, in-vehicle and industrial, real-time, fault-tolerant, etc. Examples of communication networks include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicle and industrial networks (including CANBus), etc. Such communication may be unidirectional, receive-only (e.g., broadcast television), send-only unidirectional (e.g., CANBus to certain CANbus devices), or bidirectional (e.g., other computer systems using local area networks or wide area digital networks). Such communication may include communication with one or more cloud computing networks.

[0060] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. One or more input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. One or more output devices 308 may include one or more of the following: audio output devices (e.g., speakers), visual output devices (e.g., display screens or monitors), etc.

[0061] The memory 314 may include high-speed random access memory (such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDRRAM), and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disc storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuit 302. The memory 314 (or alternatively, one or more non-volatile solid-state storage devices within the memory 314) includes non-transitory computer-readable storage media. In some embodiments, the memory 314 or the non-transitory computer-readable storage media of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: · An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks. · A network communication module 318 for connecting the server system 112 to other computing devices via one or more network interfaces 304 (e.g., via a wired connection and / or a wireless connection). · A decoding module 320 for performing various functions related to encoding and / or decoding data, such as video data. In some embodiments, the decoding module 320 is an instance of the encoder component 114. The decoding module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322 for performing various functions related to decoding encoded data, such as those functions previously described with respect to the decoder component 122. ○ An encoding module 340 for performing various functions related to encoding data, such as those functions previously described with respect to the encoder component 106. · A picture memory 352 for storing pictures and picture data for use by, for example, the decoding module 320. In some embodiments, the picture memory 352 includes one or more of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0062] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transform module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).

[0063] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described with respect to the source encoder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 includes Figure 3 a subset of the illustrated modules. For example, both the decoding module 322 and the encoding module 340 use a shared prediction module.

[0064] Each of the above-identified modules stored in the memory 314 corresponds to an instruction set for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the decoding module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0065] Although Figure 3 server system 112 is shown in accordance with some embodiments, Figure 3 it is more a functional description of the various features that may be present in one or more server systems than a structural schematic of the embodiments described herein. In practice, items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation and, optionally, in part on the data traffic processed by the server system during peak and average usage periods. Example encoding techniques

[0066] The encoding processes and techniques described below may be performed on the above-described devices and systems (e.g., source device 102, server system 112, and / or electronic device 120).

[0067] In the following, one or more associated motion vectors (MVs) can be used to encode an inter-coded block. The MV is predicted using a dedicated motion vector predictor (MVP), and the difference between the current motion vector and its corresponding predictor can be written into the bitstream. The MVP is identified by an index corresponding to an entry in the constructed motion vector prediction list. This list is constructed based on motion vectors from spatial neighbors or temporal neighbors. Spatial neighbors can include blocks that are directly adjacent to the current block, as well as blocks that are close but not directly adjacent to the current block. Temporal motion vector predictors (TMVPs) can be derived using co-located blocks in the reference frames. One way to generate the temporal MVP is to store the MVs of the reference frames along with the reference indices associated with the respective reference frames, then identify the MVs of the reference frames whose trajectories pass through each 8×8 block of the current frame, and store these MVs along with the reference frame indices in a temporal MV buffer. Thereafter, given predefined block coordinates, the associated MVs stored in the temporal MV buffer are identified and projected onto the current block to derive the temporal MV predictor from the current block to its reference frame.

[0068] In the following, the term "mode 1" refers to an encoding mode that inherits the motion vectors of adjacent blocks. The term "mode 2" refers to an encoding mode that signals the motion vector difference relative to a motion vector predictor selected from spatial neighboring blocks or temporal neighboring blocks or a given derived motion vector (such as a global motion vector).

[0069] In the following, a motion vector library is a set of motion vectors derived from encoded blocks. The motion vector library can be used as the MVP for the current block.

[0070] In MMV, different views can have strong correlations, and reducing the statistical redundancy present between different views helps to improve the coding efficiency. As described in detail below, some MMV techniques utilize the block partitioning information of a first block in a first picture of a first view to perform the block partitioning and encoding of a second block in a second picture of a second view. The first block and the second block can be located at the same coordinates in the first view and the second view, respectively, and the first picture and the second picture can be associated with the same display time.

[0071] Figure 4 Shows two views (e.g., view 0 and view Figure 1) Example MVV500. Each view can be associated with a different viewport or camera. For applications such as stereoscopic video viewing, video of more than one view can be encoded. In some embodiments, the MVV 500 corresponds to a three-dimensional (3D) scene captured by two or more cameras. In some cases, optional processing such as view correction and color correction is performed on the transmitter side. After encoding the MVV sequence, the bitstream is transmitted to the receiver side, where the views are decoded and presented on a suitable 3D display. In some embodiments, the MVV includes more than two views.

[0072] Figure 5 Shows an example prediction structure 600 of an MVV according to some embodiments. Structure 600 uses temporal reference pictures (represented by horizontal and curved arrows) and inter-view reference pictures (represented by vertical arrows) for motion compensation and disparity compensation prediction. Figure 6 Shows two picture sequences corresponding to the left view and the right view, where due to the strong correlation between the two views, the pictures in the left view are used to predict the pictures in the right view, thereby improving the encoding efficiency. The picture 602 in the right view (POC 0) is a P picture / frame encoded (e.g., predicted) using the picture 604 in the left view as a reference picture. The picture 604 is an I picture. The next frame encoded in the right view is the picture 606 (POC 8), corresponding to the last P frame in the sequence. Then, the pictures 602 and 606 are used to derive the picture 610 (POC 4) in the right view. The picture 610 is a bi-directional B picture / frame. Then, the POC 0, POC 8, and POC 4 obtained in the right view are used to derive POC2 and POC 6 in the right view. In some embodiments, the pictures in the right view are multi-layered. For example, the odd POCs (represented by the lowercase letter "b") in the right view are in different layers from the even POCs.

[0073] Figure 6 Shows video data 700A and 700B, each of which includes multiple views (e.g., view 0 to view Figure 5 ) that are spatially stitched together to form a two-dimensional image. Although Figure 7 depicts six views in the video data, it should be understood that any number of views can be stitched together. There are multiple ways to stitch views spatially. For example, for six views, one, two, or three views can be stitched in each row, which may result in a 1×6 stitch, a 2×3 stitch, and a 3×2 stitch, respectively. In some embodiments, the stitching is designed such that the resulting super-large-sized picture has a desired picture size. For example, the super-large-sized picture can be close to square or a rectangle with an aspect ratio of 4:3 or 16:9.

[0074] For P slices or B slices in ultra-large-sized pictures, the motion vectors from previously encoded views in the same picture are highly correlated with the motion vectors in the currently encoded current view. Therefore, the motion vectors from previously encoded views in the same picture are very suitable for use in motion vector prediction or as a starting point for motion estimation of the current view. Additionally, in some embodiments, a motion vector predictor (MVP) can be calculated through a perspective transformation between two views (e.g., from a reference view to the current view). In some embodiments, an MVP candidate is derived for a current block 702 in the current view (Vcur) of the current picture (Pcur) (e.g., view 2). In some embodiments, the block 702 is mapped to a block 704 in the reference view (Vref) of Pcur. If the motion vector of block 704 is (Mx, My) and its reference block 706 is in the same view (Vref) of the reference picture (Pref), then it can be obtained that X2 = X3 + M x , Y2 = Y3 + My, where (X2, Y2) are the coordinates of the reference block 706, and (X3, Y3) are the coordinates of the block 708, and the block 708 is the collocated block of the block 704. Both the block 708 and the block 704 can be in the reference view Vref (e.g., view 0), and the block 708 can reside in Pref.

[0075] In Figure 6 , the block 710 is the reference block of the block 702, the block 712 is the collocated block of the block 702, and the block 712 can reside in Vcur of Pref. In some embodiments, the coordinates of the block 710 are derived based on the coordinates of the block 712 and the calculated (e.g., derived) MVP candidate for the block 702.

[0076] In some embodiments, assuming that the samples in the block share the same disparity, a disparity vector DV (Dx, Dy) is used to find the collocated block of the current block in the same picture in the reference view. In some embodiments, a position offset is established between the current view and its reference view (e.g., the offset can be twice the view width in the x direction and 0 in the y direction). The disparity vector can be added to the view offset to find the collocated block of the current block in the reference view. Since the collocated block indicated by the disparity vector in the reference view (e.g., view 0) may have been encoded, its motion vector (if it exists) can point to the reference block in view 0 of the reference picture. In some embodiments, the reference block is used as the reference block for the current block (in view 2) or as a starting point for motion estimation.

[0077] Figure 7 Shows video data 800 with multiple views according to some embodiments. The video data 800 includes several views (e.g., views 0 to views) spatially stitched together to form a two-dimensional imageFigure 5 )。A block vector (BV) can point to a reference block in a previously encoded view. In some embodiments, prior to block matching, a perspective transformation is estimated between two views (e.g., view 0 and view Figure 5 ). The perspective transformation can establish a bijective mapping between the two views. The perspective transformation can be applied to a reference view (e.g., view 0) that maps coordinates from the reference view to the current view being encoded (e.g., view Figure 5 ). The "transformed" reference view can be used as a reference for performing block matching.

[0078] For a block in the current view (e.g., block 808A), a co-located block (e.g., block 804) of the block can be found at the same location (e.g., the same coordinates) in the "transformed" reference view, and block matching can start with the co-located block as a starting point (e.g., with block 802B). Since the perspective transformation is very close to the transition between the two views, block matching can be limited to a small neighborhood of the co-located block in the "transformed" reference view. In some embodiments, the current block and its co-located block have the same position offset relative to the upper left position of their respective views. By analyzing the reference view and the current view, the perspective transformation can be estimated. The perspective transformation can also be derived using techniques from computational photography (e.g., keypoint detection and matching), or calculated directly based on camera parameters and depth map data.

[0079] In some embodiments, assuming that the samples in a block share the same disparity, a disparity vector DV (Dx, Dy) is used to indicate the disparity between a co-located block in the reference view and a reference block in the reference view. It should be noted that for each view pair, the block vector (BV) can be different for blocks at different positions in the view. The block vector pointing from the current block to the reference block in its reference view consists of two parts: the view position offset plus the disparity vector.

[0080] Figures 8A - 8C An example encoding and subsequent decoding process of the current block is shown. Figure 8A The calculation of a predicted block according to some embodiments is shown. In Figure 8A the example, intra prediction is performed on the current block 902 to generate a predicted block 904. In some embodiments, inter prediction is performed to generate a predicted block. The current block 902 includes a set of samples (e.g., a pixel block), and the predicted block 904 includes a corresponding set of predictions for the set of samples. Figure 8B The calculation of a residual block according to some embodiments is shown. As Figure 8B shown, the predicted block 904 is subtracted from the current block 902 to generate a residual block 906 that includes a set of residuals. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 8CShows the calculation of a reconstructed block according to some embodiments. As Figure 8C shown, the residual block 906 undergoes one or more transforms and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 908. The reconstructed residual block 908 is combined with the prediction block 904 (e.g., adding the reconstructed residual of the reconstructed residual block 908 to the prediction of the prediction block 904) to generate a reconstructed block 910 corresponding to the current block 902.

[0081] Figure 9 Shows an example in-loop filtering stage according to some embodiments. In Figure 9 it, the in-loop filtering stage applied to the decoded frame 922 includes a deblocking filter 924, a constrained direction enhancement filter (CDEF) 926, and a loop restoration filter 930. In some embodiments, the filtered output frame is used as a reference frame for subsequent frames (e.g., stored in the reference frame buffer 928). The loop filtering method can include any filtering process applied to the reconstructed samples (e.g., after adding the residual to the prediction), including Wiener loop filtering, cross-component filtering, and CDEF. In some cases, the reconstructed samples filtered via loop filtering can be used as reference samples for performing prediction within a picture. In some cases, the reconstructed samples filtered via loop filtering cannot be used as reference samples.

[0082] The cross-component filtering method can use the co-located reconstructed samples and adjacent reconstructed samples from a first color component as inputs to perform filtering on the current reconstructed sample of a second color component. The cross-component offset filtering method can use the co-located reconstructed samples and their adjacent reconstructed samples from a first color component as inputs to derive an offset value, which is added to the current sample of the second color component to adjust its reconstructed value. The first color component can refer to the luminance color component, and the second color component can refer to the chrominance color component. The first color component and the second color component can be the same color component (e.g., the luminance component).

[0083] The deblocking filter 924 can be applied across transform block boundaries to remove block artifacts caused by quantization errors. In some embodiments, the filter length is determined based on the minimum transform block size on both sides. In some embodiments, the deblocking filter 924 uses a finite impulse response (FIR) filter (e.g., a low-pass filter). Edge detection can be used to disable the deblocking filter at transitions containing high variance signals (e.g., to avoid blurring actual edges in the original image). In this way, the deblocking filtering method can be applied to the reconstructed samples located near the block boundaries. The block boundaries can include transform block boundaries, motion compensation block boundaries, coding block boundaries, and / or fixed block size boundaries.

[0084] The CDEF 926 applies a non-linear de-ringing filter along a specific (e.g., tilted) direction. The CDEF 926 can operate on the output of the deblocking filter 924. The CDEF 926 can operate in 8x8 units. In some embodiments, eight preset directions are defined by rotating and reflecting a template in a preset direction. The decoder can use the reconstructed pixels to select a dominant direction index. The main filter can be applied along the selected direction, and the secondary filter can be applied along an offset direction (e.g., 45° off the main direction). In some embodiments, up to eight sets of filter parameters are signaled (e.g., in the frame header). The filter parameter sets can include main filter strength indices and secondary filter strength indices for the luminance and chrominance components. The CDEF 926 can apply filtering to the reconstructed samples by identifying the direction of each block, and then perform highly controlled adaptive filtering on the filtering intensity along and across that direction.

[0085] In some embodiments, the loop restoration filter 930 is applied to the reconstructed pixels after any previous in-loop filtering stage (e.g., the deblocking filter 924 and / or the CDEF 926). The loop restoration filter 930 can be applied to a loop restoration unit (LRU), e.g., 64x64, 128x128, and / or 256x256 pixel blocks. Bypass filtering, a Wiener filter (e.g., the Wiener loop filtering method), and / or a self-guiding filter can be independently applied to each LRU. The Wiener loop filtering method can use a linear weighted sum of the current reconstructed sample and multiple spatially adjacent reconstructed samples as input to derive a modified value of the current reconstructed sample as output.

[0086] Figure 10AFIG. is a flowchart showing a method 1000 for decoding video according to some embodiments. Method 1000 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 1000 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0087] The system receives (1002) a multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view. The system identifies (1004) a first set of reference frames for the first block in the first view. The system obtains (1006) a set of motion vectors corresponding to a second set of reference frames in the second view for the second block. Based on determining that the first set of reference frames and the second set of reference frames share a display time, the system uses the set of motion vectors corresponding to the second view to derive (1008) a motion vector prediction (MVP) for the first block corresponding to the first view. The system decodes (1010) the first block using the derived MVP. For example, one (or more) motion vectors derived from encoded information associated with another view are used to derive the MVP of a block in the current view. In some embodiments, a motion vector associated with a first encoded block of a first picture in a first view, i.e., a view-based motion vector predictor (VMVP), is used as the MVP of a second encoded block of a second picture in a second view. In some embodiments, a disparity vector is used to extract the first encoded block, the disparity vector being derived from an adjacent block encoded using the second view as a reference frame. In one example, the disparity vector is (0,0), which means that the first encoded block and the second encoded block are located at the same coordinates in the first view and the second view, respectively. In another example, the disparity vector is derived from a global disparity vector associated with a combination between two frames from different views that are displayed simultaneously.

[0088] In some embodiments, one (or more) reference frames belonging to the first view are used to encode the first encoded block. The display time of one (or more) reference frames of the first view is compared with the display time of the reference frames of the second block in the second view. If the display times are the same, the motion vector associated with the first encoded block is used as the motion vector predictor for the second encoded block. In some embodiments, one (or more) reference frames belonging to the first view are used to encode the first encoded block. The display time of the reference frames is different from the display time of the reference frames of the second block in the second view. The motion vector associated with the first encoded block is scaled and used as the motion vector predictor for the second encoded block.

[0089] In some cases, the scaling factor is proportional to the ratio of the temporal distance between the first picture in the first view and its reference frame and the temporal distance between the second picture in the second view and its reference frame.

[0090] In some embodiments, the MVP of the current block is constructed. The checking order of the TMVP candidate or the VMVP candidate relative to other motion vector predictor candidates (e.g., spatial motion vector predictor candidates, global motion vector predictor candidates, motion vector predictor candidates in the motion vector library) is different. In the following embodiments, both the TMVP and the VMVP are in the motion vector list. Each is associated with a different index in the motion vector list. In some embodiments, the motion vector list has a predefined fixed (maximum) number of TMVP and / or VMVP. The TMVP / VMCP candidates in the list depend on the checking order and availability of the TMVP and VMVP candidates. In one example, two lists are constructed, namely the MVP list from the multi-view and the list from the current frame. A high-level flag or a block-level flag is signaled to indicate which list to use.

[0091] In some embodiments, more than one motion vector extracted from a plurality of coded blocks in the first picture of the first view is used as an MVP candidate for the second block in the second picture of the second view. The extracted motion vectors can be added to the MVP list and used as the MVP of the second block.

[0092] In some embodiments, the coded blocks are from coded blocks that are not adjacent to the first coded block with respect to the first picture in the first view. In some embodiments, a disparity vector is used to determine the first coded block, and the disparity vector is derived from adjacent blocks of the second block encoded using the second view as a reference frame. In some embodiments, the disparity vector is (0, 0), for example, the first coded block and the second coded block are located at the same coordinates in the first view and the second view, respectively.

[0093] In some embodiments, the disparity vector is derived from a global disparity vector associated with the combination between two frames from different views that are displayed simultaneously.

[0094] In some embodiments, if the motion vectors of one (or more) spatial (or temporal) adjacent blocks point to the first view, the motion vectors of the adjacent blocks are also inserted into the MVP list of the current block. In some embodiments, for the current block, an additional motion vector is signaled explicitly or derived implicitly, and the additional motion vector is used to indicate the position displacement in the reference picture of the second view. In some embodiments, the coordinates of the coded blocks are predefined or derived implicitly using encoded information such as block shape, quantization parameter, and / or temporal layer.

[0095] In some embodiments, a motion vector library associated with a block in a first picture of a first view is used to derive a motion vector predictor for a block in a second picture of a second view. In some embodiments, a motion vector library associated with a block in a first picture of a first view is merged with a motion vector library associated with a block in a second picture of a second view.

[0096] In some embodiments, the use or enabling of the above embodiments is controlled by high-level syntax, including but not limited to sequence-level flags, picture-level flags, sub-picture-level flags, slice-level flags, tile-level flags, and maximum coded block-level flags.

[0097] In some embodiments, for a plurality of blocks belonging to different views, the block partitioning mode is jointly signaled. In some embodiments, these blocks are co-located blocks (e.g., blocks located at the same coordinates in different views). In some embodiments, the position of the block depends on the disparity between the views. In certain cases, the disparity between the views is quantized to certain predefined values. These values can be powers of 2 (or 4), such as 4, 8, 16.

[0098] In some embodiments, a flag is signaled for blocks from different views to indicate whether the partitioning mode of the blocks is jointly signaled or separately signaled. In some embodiments, the flag is signaled conditionally. For example, when the block is larger (or smaller) than a predefined threshold (e.g., 128×128, 64×64, 32×32, 16×16, or 8×8), the flag is signaled. In some embodiments, when the partitioning mode of the blocks is jointly signaled, the amount of signaling syntax related to the partitioning mode is less than the number of blocks. For example, if N blocks (where N>1) share the same partitioning mode, the partitioning mode is signaled once instead of N times (for the N blocks). In some embodiments, when the partitioning mode of the blocks is jointly signaled, the partitioning manner of the blocks is shared for a given first depth value, and beyond the first depth value, the partitioning modes of these blocks are signaled separately.

[0099] In some embodiments, when the block partitioning mode is signaled for a plurality of blocks from different views, the prediction mode of the blocks (e.g., intra prediction or inter prediction and its direction) is also jointly signaled. For example, the partitioning mode can specify how to partition the block into smaller block sizes. The partitioning mode can also specify whether different color components share the same partitioning.

[0100] In some embodiments, the block partitioning pattern of the first block in the first view is used to derive the context for signaling the partitioning pattern of the second block from the second view. The context provides a better estimate of the probability of a bit having a certain value, which in turn improves the coding efficiency. In some embodiments, the same context is shared to encode the same type of syntax for different views. For example, when encoding the block partitioning pattern, the context associated with signaling the block partitioning pattern of the first view can be further updated by signaling the block partitioning pattern of the second view, and vice versa.

[0101] In some embodiments, the signaling of the syntax is interleaved instead of encoding different pictures from different views separately. For example, the syntax belonging to the second view can be written between multiple syntaxes written for the first view. In some embodiments, the signaling of the largest coding unit (LCU) for different views is interleaved. An example coding order can be LCU0 of view 0, LCU0 of Figure 1 view Figure 1 and LCU1 of view 0 and LCU1 of

[0102] view. The advantage of interleaving the signaling of the largest coding unit (LCU) is that context update can be performed more efficiently.

[0103] In some embodiments, for multiple pictures from different views, loop filter parameters associated with one or more loop filter methods are jointly signaled. Example loop filter methods include cross-component filtering, cross-component offset filtering, Wiener loop filtering, deblocking filtering, and CDEF methods. In some embodiments, high-level syntax (HLS) is signaled to indicate whether the loop filter parameters are jointly signaled for multiple pictures from different views or are signaled separately for multiple pictures from different views. In some embodiments, high-level syntax is jointly signaled for multiple views, and the high-level syntax is used to signal the enabling of one (or more) loop filter methods.

[0104] In some embodiments, picture-level on / off flags for one (or more) loop filter methods are jointly signaled for multiple views. In some embodiments, depending on the picture-level on / off flag of a loop filter method of one view, the picture-level on / off flag of a loop filter method in another view is signaled. In one example, for a first picture in a first view, three flags for different color components may be signaled to specify whether the first loop filter method is applied to the three color components separately. Then, for a second view, one flag is signaled to indicate whether to inherit the values of the three flags for a second picture in the second view. In another example, for a first picture in a first view, multiple flags are separately signaled to specify whether each of multiple loop filter methods is enabled. Then, for a second picture in a second view, one flag is signaled to indicate whether the enabling of the multiple loop filter methods is inherited from the first picture in the first view or is signaled separately for the second picture in the second view.

[0105] In some embodiments, parameters used in one loop filter method are jointly signaled for multiple views. In some embodiments, parameters of the cross-component offset filtering method are jointly signaled for multiple views. The parameters may include: an offset look-up table, a selection between offset-only, edge-only offset, and edge-combined offset, a filtering shape, a quantization step size, a downsampling filter, and a filter unit size. In some embodiments, loop filter parameters of the Wiener loop filtering method are jointly signaled for multiple views. The parameters may include a selection of filtering parameters, filtering shape, and / or size, and / or a filter unit size. In some embodiments, loop filter parameters of CDEF are jointly signaled for multiple views.

[0106] In some embodiments, a signaling flag is used to indicate whether one or more loop filter parameters are shared from a first picture in a first view to a second picture in a second view. In some embodiments, a signaling flag is used to indicate whether one or more loop filter parameters are predicted from a first picture in a first view to a second picture in a second view. In some embodiments, a prediction residual on one or more loop filter parameters is signaled for a second picture in a second view.

[0107] In some embodiments, some of the parameters of a loop filter are shared between two views, while the remaining parameters are signaled jointly, dependently, or independently for the two views. For example, the on / off flags for the two views are signaled separately, while the parameter set for controlling the loop filter is signaled independently for the two views.

[0108] Some embodiments apply loop filtering to the reconstruction of a first picture from a first view by using the reconstructed picture of a second picture from a second view as one of the inputs to the loop filtering process (e.g., cross-view loop filtering). In some embodiments, the input to the loop filter includes the reconstruction of the first picture in the first view, and the output includes an offset to be added on top of the reconstruction of the second picture in the second view. In some embodiments, the inputs of the original loop filtering method used in a single view are also used together as the inputs to the loop filtering method.

[0109] In some embodiments, a first component of the reconstruction of a first picture in a first view is used as the input for loop filtering a second component of the reconstruction of a second picture in a second view. The first component and the second component may be different components. In some embodiments, the input from the second view is the samples around the co-located sample of the sample to be filtered in the first view. In some embodiments, the input from the second view is the sample located by a disparity vector.

[0110] In some embodiments, the disparity vector is derived from the disparity vectors associated with adjacent blocks that use different views as reference pictures, such as inter-view prediction. In some embodiments, the precision of the disparity vector includes a predefined precision, such as integer pixels, half pixels, or quarter pixels.

[0111] In some embodiments, the filtering order of pictures from different views is different from the coding order of pictures from different views. For example, for residual coding, the first picture from the first view is encoded first, and then the second picture from the second view is residual-encoded. Before loop filtering the reconstruction of the first picture, loop filtering is performed on the reconstruction of the second picture.

[0112] In some embodiments, the reconstructed samples from two views are jointly used as the input to the loop filter for each view. For example, the reconstructed samples from two views are used to derive the filter parameters for each individual view.

[0113] Figure 10B FIG. 1050 is a flow chart illustrating a method 1050 for encoding video according to some embodiments. Method 1050 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions executable by the control circuit. In some embodiments, method 650 is performed by executing instructions stored in a memory (e.g., memory 314) of the computing system.

[0114] The system receives (1052) video data including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view. The system identifies (1054) a first set of reference frames for the first block in the first view. The system obtains (1056) a set of motion vectors corresponding to a second set of reference frames in the second view for the second block. The system selects (1058) a motion vector for the first block from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames share a display time. The system encodes (1060) the first block using the selected motion vector. As previously described, the encoding process may reflect the decoding process described herein (e.g., the derivation and identification of motion vectors). For the sake of brevity, these details are not repeated herein.

[0115] Although Figure 10A and Figure 10B the multiple logical stages are shown in a particular order, stages that are not order-dependent may be reordered, and other stages may be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the order and grouping presented herein are not exhaustive. Additionally, it should be recognized that these stages may be implemented in hardware, firmware, software, or any combination thereof.

[0116] Now turning to some example embodiments:

[0117] (A1)In one aspect, some embodiments include a method for video decoding (e.g., method 1000). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at an encoding module (e.g., decoding module 320). In some embodiments, the method is performed at a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) identifying a first set of reference frames for the first block in the first view; (iii) obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; (iv) based on determining that the first set of reference frames and the second set of reference frames share a display time, using the set of motion vectors corresponding to the second view to derive a motion vector predictor (MVP) for the first block corresponding to the first view; and (v) decoding the first block using the derived MVP. For example, one (or more) motion vectors derived from encoded information associated with another view are used to derive the MVP of a block in the current view. As an example, a motion vector associated with a first encoded block of a first picture in a first view, i.e., a view-based motion vector predictor (VMVP), is used as the MVP of a second encoded block of a second picture in a second view. As an example, if a first encoded block is encoded using reference frames belonging to a first view, the display times of these reference frames are compared with the display times of the reference frames of the second block in the second view, and if the display times are the same, the motion vector associated with the first encoded block can be used as the MVP of the second encoded block. In some embodiments, based on determining that the first set of reference frames and the second set of reference frames do not share a display time, the MVP of the first block is derived without using the set of motion vectors corresponding to the second view.

[0118] (A2)In some embodiments of A1, a disparity vector is used to identify the second block, where the disparity vector is derived from a set of neighboring blocks that are encoded using the first view as a reference frame. For example, a disparity vector is used to determine (e.g., extract) a first encoded block, which is derived from neighboring blocks encoded using the second view as a reference frame. As an example, the disparity vector can always be (0,0), i.e., the first encoded block and the second encoded block are located at the same coordinates in the first view and the second view, respectively. In another example, the disparity vector is derived from a global disparity vector that is associated with a combination between two frames from different views that are displayed simultaneously. In some embodiments, the global disparity vector represents the viewing angle difference between the first view and the second view. In some embodiments, the global disparity vector is derived based on the relative positioning of a first camera corresponding to the first view and a second camera corresponding to the second view.

[0119] (A3)In some embodiments of A1 or A2, the method includes: when the first set of reference frames and the second set of reference frames do not share a display time, using scaling information from a set of motion vectors corresponding to the second view to derive the MVP of a first block corresponding to the first view. For example, if a first encoded block is encoded using reference frames belonging to the first view and the display times of these reference frames are different from the reference frames of a second block in the second view, the motion vectors associated with the first encoded block are scaled and used as the MVP of the second encoded block.

[0120] (A4)In some embodiments of A3, the set of motion vectors is scaled according to a scaling factor proportional to the ratio of the temporal distances. For example, the scaling factor is proportional to the ratio of the temporal distance between a first picture and its reference frame in the first view to the temporal distance between a second picture and its reference frame in the second view.

[0121] (A5)In some embodiments of any one of A1 to A4, the method further includes: constructing an MVP list corresponding to the first block, where the MVP list includes TMVP candidates and / or VMVP candidates. For example, when constructing the MVP list of the current block, the checking order of TMVP or VMVP candidates relative to other MVP candidates (e.g., spatial MVP candidates, global MVP candidates, MVP candidates in the motion vector library) is different.

[0122] (A6)In some embodiments of A5, the MVP list includes a TMVP candidate at a first index and a VMVP candidate at a second index. For example, both TMVP and VMVP can remain in the motion vector list, and each of the two is associated with a different index in the motion vector list.

[0123] (A7)In some embodiments of A5 or A6, the MVP list is restricted to not having more than a predefined number of TMVP and / or VMVP candidates. For example, only a predefined fixed (maximum) number of TMVP and / or VMVP can remain in the motion vector list, and which TMVP / VMCP candidates can remain in the motion vector list depends on the checking order and availability of the TMVP and VMCP candidates.

[0124] (A8)In some embodiments of any one of A5 to A7, the MVP list includes one or more motion vectors corresponding to a second view. For example, more than one motion vector extracted from a plurality of coded blocks in a first picture of a first view can be used as an MVP candidate for a second block in a second picture of the second view, and the extracted motion vectors can be added to the MVP list and used as the MVP of the second block.

[0125] (A9)In some embodiments of any one of A1 to A8, the method includes: constructing a first MVP list corresponding to a plurality of views and constructing a second MVP list corresponding to a current frame, wherein, according to an indicator written in the multi-view video bitstream, the first MVP list or the second MVP list is used to derive the MVP of a first block. For example, construct an MVP list from multiple views and construct a list from the current frame, and signal a high-level flag or a block-level flag to indicate which list to use.

[0126] (A10)In some embodiments of any one of A1 to A9, the motion vectors corresponding to a plurality of blocks in a second frame corresponding to a second view are used to derive the MVP of a first block. For example, the plurality of blocks (e.g., a plurality of coded blocks) can be from coded blocks that are not adjacent to a first coded block in a first picture in a first view.

[0127] (A11)In some embodiments of A10, the coordinates of the plurality of blocks in the second frame are predefined or derived by a computing system. For example, the coordinates of a plurality of coded blocks can be predefined, or these coordinates can be implicitly derived using encoded information (including but not limited to block shape, quantization parameter, temporal layer).

[0128] (A12)In some embodiments of any one of A1 to A11, the method further includes: obtaining an additional motion vector of a first block, the additional motion vector being used to indicate a position displacement of a second set of reference frames. For example, for a current block, the additional motion vector can be explicitly signaled or implicitly derived. The motion vector is used to indicate a position displacement in a reference picture of a second view.

[0129] (A13)In some embodiments of any one of A1 to A12, the MVP of the first block is derived based on the motion vector library associated with the second block. For example, the motion vector library associated with a block in the first picture of the first view can be used to derive the MVP of a block in the second picture of the second view. In some embodiments, the motion vector library associated with a block in the first picture of the first view and the motion vector library associated with a block in the second picture of the second view are combined.

[0130] (A14)In some embodiments of any one of A1 to A13, according to a first indicator in the multi-view video bitstream, the MVP of the first block corresponding to the first view is derived using a set of motion vectors corresponding to the second view, and the first indicator is used to indicate that motion vectors from different views are to be used for the first block. For example, the first indicator is written in the high-level syntax.

[0131] (B1)In another aspect, some embodiments include a method of video encoding (e.g., method 1050). In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) receiving video data including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) identifying a first set of reference frames for the first block in the first view; (iii) obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; (iv) selecting a motion vector for the first block from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames share a display time; and (v) encoding the first block using the selected motion vector.

[0132] (B2)In some embodiments of B1, the method further includes: identifying the second block using a disparity vector, wherein the set of motion vectors is obtained according to the identification of the second block.

[0133] (B3)In some embodiments of B1 or B2, the method further includes: deriving a motion vector for the first block using scaling information from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames do not share a display time.

[0134] (B4)In some embodiments of any one of B1 to B3, the set of motion vectors corresponds to a plurality of blocks in the second view, the plurality of blocks including the second block.

[0135] (C1)In another aspect, some embodiments include a method for processing visual media data. In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) obtaining a source video sequence corresponding to a set of views; and (ii) performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, wherein the multi-view video bitstream includes: (a) a first plurality of encoded pictures corresponding to a first view; (b) a second plurality of encoded pictures corresponding to a second view; and (c) an indicator for indicating whether motion vector information from the second plurality of encoded pictures is used to encode one or more blocks from the first plurality of encoded pictures.

[0136] (C2)In some embodiments of C1, the indicator is written in the high-level syntax of the multi-view video bitstream.

[0137] (D1)In another aspect, some embodiments include a method for video decoding. In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) receiving a multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) obtaining a motion vector predictor (MVP) list for the first block, the MVP list including at least one MVP corresponding to the second view; (iii) identifying the MVP for the first block from the MVP list; and (iv) decoding the first block using the identified MVP. In some embodiments, the method further includes any of the previously described motion vector identification and / or derivation techniques.

[0138] (E1)In another aspect, some embodiments include a method for video decoding. In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) receiving a multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; (ii) identifying a motion vector library corresponding to the second block; (iii) deriving an MVP for the first block corresponding to the first view using the motion vector library corresponding to the second block; and (iv) decoding the first block using the identified MVP. In some embodiments, the method further includes any of the previously described motion vector identification and / or derivation techniques.

[0139] In another aspect, some embodiments include a computing system (e.g., server system 112) that includes control circuitry (e.g., control circuitry 302) and a memory (e.g., memory 314) coupled to the control circuitry, the memory storing one or more instruction sets configured to be executed by the control circuitry, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 - A14, B1 - B4, C1 - C2, D1, and E1 above).

[0140] In yet another aspect, some embodiments include a non - transitory computer - readable storage medium storing one or more instruction sets that are executed by control circuitry of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 - A14, B1 - B4, C1 - C2, D1, and E1 above).

[0141] Unless otherwise specified, any syntactic element described herein can be high - level syntax (HLS). As used herein, HLS is written at a level higher than the block level. For example, HLS can correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements can be written in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, picture header, tile header, and / or CTU header.

[0142] It should be understood that although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The proper nouns used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that when used in this specification, the terms “comprises” and / or “comprising” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0143] As used herein, the term "when" can be interpreted as "if" or "upon" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the prerequisite is true, depending on the context. Similarly, the phrase "if it is determined [that the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting" that the prerequisite is true, depending on the context. As used herein, N refers to a variable number. Unless explicitly stated, different instances of N can refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.

[0144] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.

Claims

1. A method for video decoding, the method being executed on a computing system having a memory and one or more processors, the method comprising: Receiving a multi-view video bitstream, the multi-view video bitstream including a plurality of blocks, wherein the plurality of blocks includes a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; Identifying a first set of reference frames for the first block in the first view; Obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; Based on determining that the first set of reference frames and the second set of reference frames share a display time, using the set of motion vectors corresponding to the second view to derive a motion vector predictor (MVP) for the first block corresponding to the first view; and Decoding the first block using the derived MVP.

2. The method according to claim 1, wherein, Identifying the second block using a disparity vector, wherein the disparity vector is derived from a set of adjacent blocks that are encoded using the first view as a reference frame.

3. The method according to claim 1, further comprising, based on determining that the first set of reference frames and the second set of reference frames do not share the display time, using scaling information from the set of motion vectors corresponding to the second view to derive the MVP for the first block corresponding to the first view.

4. The method according to claim 3, wherein, Scaling the set of motion vectors according to a scaling factor proportional to a ratio of time distances.

5. The method according to claim 1 further comprises: Constructing an MVP list corresponding to the first block, wherein the MVP list includes a temporal motion vector predictor (TMVP) candidate and / or a view-based motion vector predictor (VMVP) candidate.

6. The method according to claim 5, wherein The MVP list includes the TMVP candidate at a first index and the VMVP candidate at a second index.

7. The method according to claim 5, wherein The MVP list is limited to not more than a predefined number of TMVP and / or VMVP candidates.

8. The method according to claim 5, wherein The MVP list includes one or more motion vectors corresponding to the second view.

9. The method according to claim 1, further comprising: Constructing a first MVP list corresponding to a plurality of views; And Constructing a second MVP list corresponding to a current frame, wherein, according to an indicator written in the multi-view video bitstream, using the first MVP list or the second MVP list to derive the MVP for the first block.

10. The method according to claim 1, wherein The MVP for the first block is derived using motion vectors corresponding to a plurality of blocks in the second frame corresponding to the second view.

11. The method according to claim 10, wherein, The coordinates of the plurality of blocks in the second frame are predefined or derived by the computing system.

12. The method according to claim 1 further comprises: Obtaining an additional motion vector for the first block, the additional motion vector being used to indicate a position displacement of the second set of reference frames.

13. The method according to claim 1, wherein, The MVP for the first block is derived based on a motion vector library associated with the second block.

14. The method according to claim 1, wherein, Derive the MVP of the first block corresponding to the first view using the set of motion vectors corresponding to the second view according to the first indicator in the multi-view video bitstream, where the first indicator is used to indicate that motion vectors from different views are to be used for the first block.

15. A computing system, comprising: A control circuit; A memory; And One or more instruction sets stored in the memory and configured to be executed by the control circuit, the one or more instruction sets including instructions for the following operations: Receiving video data, the video data including a plurality of blocks, where the plurality of blocks includes a first block in a first frame corresponding to a first view and a second block in a second frame corresponding to a second view; Identifying a first set of reference frames for the first block in the first view; Obtaining a set of motion vectors corresponding to a second set of reference frames in the second view for the second block; Selecting a motion vector for the first block from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames share a display time; and Encoding the first block using the selected motion vector.

16. The computing system according to claim 15, further comprising: Identifying the second block using a disparity vector, where the set of motion vectors is obtained according to the identification of the second block.

17. The computing system according to claim 15, further comprising: Deriving the motion vector of the first block using scaling information from the set of motion vectors corresponding to the second view based on determining that the first set of reference frames and the second set of reference frames do not share the display time.

18. The computing system according to claim 15, wherein, The set of motion vectors corresponds to a plurality of blocks in the second view, the plurality of blocks including the second block.

19. A non-transitory computer-readable storage medium storing one or more instruction sets configured to be executed by a computing device including a control circuit and a memory, the one or more instruction sets including instructions for the following operations: Obtain a source video sequence corresponding to a view set; And Performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, where The multi-view video bitstream includes: A first plurality of encoded pictures corresponding to a first view; A second plurality of encoded pictures corresponding to a second view; An indicator for indicating whether to use motion vector information from the second plurality of encoded pictures to encode one or more blocks from the first plurality of encoded pictures.

20. The non-transitory computer-readable storage medium according to claim 19, wherein, The indicator is written in the high-level syntax of the multi-view video bitstream.

Citation Information

Cited By

  • Video coding pre-analysis method and device

    CN121567873A