Systems and methods for coding order for block partitioning and interleaving in multi-view video coding

By jointly signaling block partition modes across views in MVV coding, the method addresses inefficiencies in existing MVV technologies, achieving a 70% reduction in bit rate and enhancing compression efficiency.

CN120323010APending Publication Date: 2025-07-15TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004607.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-07
Filing Date
2024-05-16
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing multi-view video encoding and decoding technology has high redundancy when encoding blocks of different views, resulting in low encoding efficiency.

Method used

By combining signaling modes and parameters in multi-view video encoding, the redundancy between different views is reduced, and the parallax compensation prediction method is adopted to encode blocks of multiple views simultaneously using the same block partition mode to improve the encoding and decoding efficiency.

Benefits of technology

Significantly reduces statistical redundancy between different views, improves encoding and decoding efficiency, and saves about 70% of the bit rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323010A_ABST
    Figure CN120323010A_ABST
Patent Text Reader

Abstract

An example method includes receiving a multi-view video bitstream, the multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view, the second frame corresponding to a second view; determining, based on a first indicator in the multi-view video bitstream, whether to jointly signal, or to respectively signal, a plurality of block partition modes for the first block corresponding to the first view and the second block corresponding to the second view; and, when the first indicator indicates that the plurality of block partition modes are jointly signaled, applying a same block partition mode to the first block and the second block according to a second indicator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by Reference

[0002] This application claims priority to U.S. Provisional Application No. 63 / 599,274, filed on Nov. 15, 2023, with the title "Coding and Decoding Order of Block Partitioning and Interleaving in Multi-View Video Coding and Decoding", and to U.S. Application No. 18 / 657,709, filed on May 7, 2024, with the title "Systems and Methods for Coding and Decoding Order of Block Partitioning and Interleaving in Multi-View Video Coding and Decoding", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to video coding and decoding technologies, and particularly to a system and method for coding and decoding block partitioning and interleaving in multi-view video (MVV) coding and decoding. Background Art

[0004] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. These electronic devices transmit and receive or otherwise communicate digital video data over communication networks and / or store digital video data on storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding and decoding can be used to compress video data according to at least one video coding and decoding standard before communicating or storing the video data. Video coding and decoding can be performed by hardware and / or software on an electronic / client device or a server providing cloud services.

[0005] Video codecs typically utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, or the like) that exploit the redundancy inherent in video data. Video codecs aim to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. A variety of video codec standards have been developed. For example, High Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. ITU-T and ISO / IEC released the HEVC / H.265 standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Versatile Video Coding (VVC / H.266) is a video compression standard intended to be a successor to HEVC. ITU-T and ISO / IEC released the VVC / H.266 standard in 2020 (version 1) and 2022 (version 2). AOMedia Video 1 (AV1) is an open video codec format designed as an alternative to HEVC. On January 8, 2019, the validated specification version 1.0.0 with Errata 1 was released. Summary of the invention

[0006] The present application describes a set of methods for video (image) compression, and more specifically, relates to deriving motion vector predictions when encoding multiple views of a scene. In some embodiments, instead of encoding each view and sending a bitstream from each view independently (simulcast coding), a disparity compensation prediction method is implemented to include pictures of other views in the reference picture list at the same time. This method, also known as disparity compensation prediction, can improve codec efficiency by reducing the statistical redundancy that exists between different views. In some cases, the method disclosed in the present application can save about 70% of the bit rate compared to simulcast coding.

[0007] According to an embodiment of the present application, a video decoding method is provided, including:

[0008] Receive a multi-view video stream, the multi-view video stream comprising a plurality of blocks, the plurality of blocks comprising a first block in a first frame and a second block in a second frame, the first frame corresponds to a first view, and the second frame corresponds to a second view;

[0009] determining, based on a first indicator in the multi-view video stream, whether to jointly signal or separately signal a plurality of block partition modes for the first block corresponding to the first view and the second block corresponding to the second view; and

[0010] When the first indicator indicates that the plurality of block partition modes are jointly signaled, the same block partition mode is applied to the first block and the second block according to a second indicator.

[0011] According to an embodiment of the present application, a video encoding method is provided, including:

[0012] Receiving video data, where the video data includes a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view;

[0013] Determining whether to signal jointly or separately for a plurality of block partitioning patterns for the first block and the second block; and,

[0014] When it is determined to signal jointly for the plurality of block partitioning patterns, signaling a first indicator in the video bitstream, the first indicator indicating that the plurality of block partitioning patterns are signaled jointly.

[0015] According to an embodiment of the present application, a method for bitstream conversion is provided, including:

[0016] Obtaining a source video sequence, the source video sequence including a set of views;

[0017] Performing conversion between the source video sequence and a multi-view video bitstream of visual media data, where the multi-view video bitstream includes:

[0018] (i) A first plurality of encoded pictures corresponding to a first view;

[0019] (ii) A second plurality of encoded pictures corresponding to a second view;

[0020] (iii) A first indicator indicating whether to signal jointly or separately for block partitioning patterns for the first block and the second block;

[0021] (iv) When the first indicator indicates signaling jointly for the block partitioning patterns, a second indicator indicating the block partitioning patterns;

[0022] (v) When the first indicator indicates signaling separately for the block partitioning patterns, a set of indicators indicating the block partitioning patterns of the first block and the second block respectively.

[0023] According to an embodiment of the present application, a computing system, such as a streaming media system, a server system, a personal computer system, or other electronic devices, is provided. The computing system includes a control circuit and a memory storing one or more sets of instructions. The one or more sets of instructions include instructions for performing any method described in the present application. In some embodiments, the computing system includes an encoder component and a decoder component (such as a code converter).

[0024] According to some embodiments, a non - volatile computer - readable storage medium is provided. The non - volatile computer - readable storage medium stores one or more sets of instructions for execution by a computing system. The one or more sets of instructions include instructions for performing any of the methods described in this application.

[0025] Accordingly, an apparatus and a system with a video encoding and decoding method are disclosed. Such methods, apparatuses, and systems may supplement or replace traditional methods, apparatuses, or systems for video encoding and decoding. The features and advantages described in the specification are not necessarily all - encompassing. In particular, in view of the accompanying drawings, specification, and claims provided in this application, some additional features and advantages will be apparent to those of ordinary skill in the art. In addition, it should be noted that the language used in the specification is mainly selected for readability and indicative purposes and is not necessarily for describing or limiting the subject matter described in this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] For a more detailed understanding of this application, a more specific description can be made with reference to the features of various embodiments, some of which are shown in the accompanying drawings. However, the drawings only show the relevant features of this application and are not necessarily considered limiting, as those skilled in the art will understand after reading this application that the description may admit other valid features.

[0027] Figure 1 is a schematic diagram of an exemplary block diagram of a communication system according to an embodiment of this application;

[0028] Figure 2A is a schematic diagram of an exemplary block diagram of an encoder component according to an embodiment of this application;

[0029] Figure 2B is a schematic diagram of an exemplary block diagram of a decoder component according to an embodiment of this application;

[0030] Figure 3 is a schematic diagram of an exemplary block diagram of a server system according to an embodiment of this application;

[0031] Figures 4A - 4D shows an example coding tree structure according to an embodiment of this application;

[0032] Figure 5 shows an example MVV according to an embodiment of this application;

[0033] Figures 6 - 8 shows an example operation in the MVV according to an embodiment of this application;

[0034] Figures 9A - 9C shows an example prediction block, residual block, and reconstruction block according to an embodiment of this application;

[0035] Figure 10AShows an example video decoding process according to an embodiment of the present application;

[0036] Figure 10B Shows an example video encoding process according to an embodiment of the present application.

[0037] By convention, the various features shown in the drawings are not necessarily drawn to scale, and throughout the specification and drawings, the same reference numerals may be used to denote the same features. Detailed implementation

[0038] The present application describes video / image compression techniques, including block segmentation and interleaved coding for MVV coding. The disclosed techniques include joint signaling modes and / or parameters for blocks belonging to different views of MVV. For example, depending on whether the block segmentation modes of the first block and the second block are signaled jointly or separately, it is determined whether the same block segmentation method is applied to the first block in the first frame corresponding to the first view (e.g., from the first camera in a stereoscopic vision system) and the second block in the second frame corresponding to the second view (e.g., from the second camera in the system). By reducing the redundancy between views through joint signaling modes and / or parameters, the encoding and decoding efficiency is improved.

[0039] Example systems and devices

[0040] Figure 1 Is a block diagram illustrating a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and at least two electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m) communicatively coupled to each other via at least one network. In some embodiments, the communication system 100 is a streaming system, e.g., for applications where video is available, such as video conferencing applications, digital TV applications, and media storage and / or distribution applications.

[0041] The source device 102 includes a video source 104 (e.g., a camera component or media storage) and an encoder component 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder component 106 generates at least one encoded video bitstream from the video stream. Compared with the video stream from the video source 104, the video stream from the video source 104 can be of high data volume. Since the encoded video bitstream 108 has a lower data volume (less data) compared with the video stream from the video source, the encoded video bitstream 108 requires less bandwidth for transmission and less storage space for storage compared with the video stream from the video source 104. In some embodiments, the source device 102 does not include an encoder component 106 (e.g., configured to transmit uncompressed video to at least one network 110).

[0042] At least one network 110 represents any number of networks for transferring information between the source device 102, the server system 112, and / or the electronic device 120, including, for example, wired (wired) and / or wireless communication networks. The at least one network 110 can exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0043] The at least one network 110 includes the server system 112 (e.g., a distributed / cloud computing system). In some embodiments, the server system 112 is or includes a streaming media server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes codec components 114 (e.g., configured to encode and / or decode video data). In some embodiments, the codec components 114 include an encoder component and / or a decoder component. In various embodiments, the codec components 114 are instantiated as hardware, software, or a combination thereof. In some embodiments, the codec components 114 are configured to decode the encoded video bitstream 108 and re-encode the video data using different coding standards and / or methods to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate at least two video formats and / or encodings from the encoded video bitstream 108. In some embodiments, the server system 112 serves as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for at least one of the electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.

[0044] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an outgoing video stream that can be rendered on a display or other type of rendering device. In some embodiments, at least one of the electronic devices 120 does not include a display component (e.g., is communicatively coupled to an external display device and / or includes media storage). In some embodiments, the electronic device 120 is a streaming media client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.

[0045] The source device and / or at least two electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, at least one of the source device 102 and / or the electronic devices 120 is an instance of a server system, a personal computer, a portable device (e.g., a smart phone, a tablet computer, or a laptop computer), a wearable device, a video conferencing device, and / or other types of electronic devices.

[0046] In an example operation of the communication system 100, the source device 102 transmits an encoded video bitstream 108 to the server system 112. For example, the source device 102 may encode a picture stream captured by the source device. The server system 112 receives the encoded video bitstream 108 and may decode and / or encode the encoded video bitstream 108 using codec components 114. For example, the server system 112 may apply an encoding that is more optimal for network transmission and / or storage to the video data. The server system 112 may transmit the encoded video data 116 (e.g., at least one encoded video bitstream) to at least one of the electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.

[0047] Figure 2A is a block diagram illustrating example elements of an encoder component 106 according to some embodiments. The encoder component 106 receives video data (e.g., a source video sequence) from a video source 104. In some embodiments, the encoder component includes a receiver (e.g., a transceiver) component configured to receive the source video sequence. In some embodiments, the encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from the encoder component 106). The video source 104 may provide a source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, the video source 104 is a storage device that stores previously captured / prepared video. In some embodiments, the video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as at least two separate pictures that impart motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include at least one sample, depending on the sampling structure, color space, etc. being used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.

[0048] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by the application. In some embodiments, the encoder component 106 is configured to perform a conversion between the source video sequence and a bitstream of visual media data (e.g., a video bitstream). Implementing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls other functional units as described below and is functionally coupled to other functional units. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skip, quantizer, and / or λ value of rate distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to the encoder component 106 optimized for a particular system design.

[0049] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simplified example, the encoding loop includes a source decoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and at least one reference picture), and a (local) decoder 210. The decoder 210 reconstructs the symbols in a manner similar to a (remote) decoder to create sample data (when the compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input to the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the same sample values as the sample values that the decoder will interpret during decoding when using prediction as reference picture samples.

[0050] The operation of the decoder 210 may be the same as the operation of a remote decoder (such as the decoder component 122), which will be described in detail below in conjunction with Figure 2B However, briefly referring to Figure 2B , since the symbols are available and the entropy encoder 214 and the parser 254 encoding / decoding the symbols into the encoded video sequence can be lossless, the entropy decoding part of the decoder component 122 (including the buffer memory 252 and the parser 254) may not be fully implemented in the local decoder 210.

[0051] Except for parsing / entropy decoding, the decoder techniques described in this application may exist in the corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. Additionally, the description of encoder techniques may be brief as they may be the opposite of decoder techniques.

[0052] As part of its operation, the source decoder 202 may perform motion compensated predictive coding, where the source decoder predictively encodes an input frame by referring to at least one previously encoded frame designated as a reference frame from a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of at least one reference frame that may be selected as at least one predictive reference for the input frame. The controller 204 may manage the coding and decoding operations of the source decoder 202, including, for example, setting parameters and sub-group parameters for encoding video data.

[0053] The decoder 210 decodes the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source decoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by a remote video decoder on a reference frame, and may store the reconstructed reference frame in the reference picture memory 208. In this way, the encoder component 106 locally stores copies of the reconstructed reference frames, which have the same content as the reconstructed reference frames that would be obtained by a remote video decoder (in the absence of transmission errors).

[0054] The predictor 206 may perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new picture. The predictor 206 may operate on a per-pixel-block basis of sample blocks to find an appropriate prediction reference. As determined by the search results obtained by the predictor 206, the input picture may have prediction references extracted from at least two reference pictures stored in the reference picture memory 208.

[0055] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 converts the symbols generated by the various functional units into an encoded video sequence by losslessly compressing the symbols according to techniques known to those of ordinary skill in the art (e.g., Huffman coding, variable length coding, and / or arithmetic coding).

[0056] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer at least one encoded video sequence created by the entropy encoder 214 to prepare them for transmission via a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source decoder 202 with other data to be transmitted (e.g., encoded audio data and / or auxiliary data streams (not shown in the source)). In some embodiments, the transmitter may transmit additional data along with the encoded video. The source decoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data (such as redundant pictures and slices), supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.

[0057] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a specific encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). Intra pictures can be encoded and decoded without using any other frames in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including for example instant decoder refresh (IDR) pictures. Those of ordinary skill in the art are aware of those variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. Predictive pictures can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and reference index to predict the sample values of each block. Bi-predictive pictures can be encoded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indexes to predict the sample values of each block. Similarly, at least two predictive pictures can be used to reconstruct a single block using more than two reference pictures and associated metadata.

[0058] Source pictures can typically be spatially subdivided into at least two sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and encoded on a block-by-block basis. A block can be predictively encoded with reference to other (already encoded) blocks determined by an encoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture can be non-predictively encoded, or they can be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture can be non-predictively encoded via spatial prediction or via temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be non-predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures.

[0059] Video can be captured as at least two source pictures (video pictures) in a time series. Intra-picture prediction (commonly abbreviated as intra prediction) exploits spatial correlations within a given picture, and inter-picture prediction exploits (temporal or other) correlations between pictures. In an example, a particular picture (which is referred to as the current picture) in encoding / decoding is partitioned into blocks. When a block in the current picture resembles a reference block in a previously encoded and still buffered reference picture in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture and can have a third dimension identifying the reference picture in cases where at least two reference pictures are used.

[0060] The encoder component 106 can perform encoding operations according to a predetermined video coding technique or standard (such as any technique or standard described in this application). In its operation, the encoder component 106 can perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.

[0061] Figure 2B is a block diagram illustrating example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter coupled to the loop filter 256 and configured to transmit data to the display 124 (e.g., via a wired or wireless connection).

[0062] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data (e.g., via a wired or wireless connection) from channel 218. The receiver may be configured to receive at least one encoded video sequence to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device that stores the encoded video data. The receiver may receive the encoded video data as well as other data, such as encoded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data with the encoded video. The additional data may be included as part of at least one encoded video sequence. Decoder component 122 may use the additional data to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0063] According to some embodiments, decoder component 122 includes buffer memory 252, parser 254 (sometimes also referred to as an entropy decoder), scaler / inverse transform unit 258, intra picture prediction unit 262, motion compensation prediction unit 260, aggregator 268, loop filter unit 256, reference picture memory 266, and current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.

[0064] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to counter network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 inside decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided outside decoder component 122 (e.g., to counter network jitter). When receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer memory 252 may not be needed, or buffer memory 252 may be small. For use on a best-effort packet network such as the Internet, buffer memory 252 may be needed, buffer memory 252 may be relatively large and / or have an adaptive size, and may be implemented at least partially in an operating system or similar element outside decoder component 122.

[0065] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. The symbols can include, for example, information for managing the operation of the decoder component 122 and / or information for controlling a rendering device such as the display 124. The control information for at least one rendering device can be in the form of, for example, a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence can be according to a video coding technique or standard and can follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 254 can extract a set of subgroup parameters of the group based on at least one parameter corresponding to at least one of the pixel subgroups in the video decoder. The subgroups can include groups of pictures (GOPs), pictures, tiles, strips, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), and so on. The parser 254 can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0066] The reconstruction of the symbols 270 can involve at least two different units, depending on the type of the encoded video picture or its part (such as: inter and intra pictures, inter and intra blocks) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser 254 from the encoded video sequence. For clarity, the flow of such subgroup control information between the parser 254 and the at least two units below is not depicted.

[0067] The decoder component 122 can be conceptually subdivided into at least two functional units, and in some embodiments, these units interact closely with each other and can be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is maintained here.

[0068] The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform to use, block size, quantization factor, and / or quantization scaling matrix) as at least one symbol 270 from the parser 254. The scaler / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268. In some cases, the output samples of the scaler / inverse transform unit 258 relate to intra-coded blocks; that is, blocks that do not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 262. The intra-picture prediction unit 262 may use the surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture memory 264 to generate a block having the same size and shape as the block being reconstructed. The aggregator 268 may add the prediction information generated by the intra-picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 on a per-sample basis.

[0069] In other cases, the output samples of the scaler / inverse transform unit 258 relate to inter-coded and possibly motion-compensated blocks. In such cases, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After the extracted samples are motion-compensated according to the symbol 270 associated with the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or residual signals in this case) to generate output sample information. The address within the reference picture memory 266 from which the motion compensation prediction unit 260 extracts prediction samples may be controlled by a motion vector. The motion vector may be available to the motion compensation prediction unit 260 in the form of a symbol 270, which may have, for example, X, Y, and reference picture components. Motion compensation may also include, for example, interpolation of sample values extracted from the reference picture memory 266 when using sub-sample accurate motion vectors, a motion vector prediction mechanism.

[0070] The output samples of the aggregator 268 may undergo various loop filtering techniques in the loop filter unit 256. Video compression techniques may include in-loop filter techniques, which are controlled by parameters included in the encoded video bitstream and provided to the loop filter unit 256 as a symbol 270 from the parser 254, but may also respond to meta-information obtained during the decoding of a previously (in decoding order) portion of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream, which may be output to a rendering device such as a display 124 and stored in the reference picture memory 266 for use in future inter-picture prediction.

[0071] Once certain coded pictures are reconstructed, they can be used as reference pictures for future prediction. Once a coded picture is reconstructed and the coded picture is identified as a reference picture (e.g., by parser 254), the current reference picture can become part of the reference picture memory 266, and a new current picture memory can be reallocated before starting to reconstruct the next coded picture.

[0072] The decoder component 122 can perform decoding operations according to predetermined video compression techniques that can be specified in a standard (such as any standard described in this application). The coded video sequence can conform to the syntax specified by the video compression technique or standard being used, in the sense that it follows the syntax of the video compression technique or standard as specified in the video compression technique document or standard and particularly in the profile document therein. And, to conform to some video compression techniques or standards, the complexity of the coded video sequence can be within the bounds defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specification and metadata signaled in the coded video sequence for HRD buffer management.

[0073] Figure 3 is a block diagram illustrating a server system 112 according to some embodiments. The server system 112 includes control circuitry 302, at least one network interface 304, a memory 314, a user interface 306, and at least one communication bus 312 for interconnecting these components. In some embodiments, the control circuitry 302 includes at least one processor (e.g., a CPU, GPU, and / or DPU). In some embodiments, the control circuitry includes at least one field programmable gate array, a hardware accelerator, and / or at least one integrated circuit (e.g., an application specific integrated circuit).

[0074] At least one network interface 304 may be configured to interface with at least one communication network (e.g., wireless, wired, and / or optical networks). The communication network can be local, wide area, metro area, vehicular and industrial, real-time, delay tolerant, etc. Examples of communication networks include local area networks (such as Ethernet, wireless LAN), cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial (including CANBus), etc. Such communication can be one-way, receive-only (e.g., broadcast TV), one-way send-only (e.g., CANbus to certain CANbus devices), or two-way (e.g., to other computer systems using local or wide area digital networks). Such communication can include communication with at least one cloud computing network.

[0075] The user interface 306 includes at least one output device 308 and / or at least one input device 310. The at least one input device 310 may include at least one of the following: keyboard, mouse, touchpad, touchscreen, data glove, joystick, microphone, scanner, camera, etc. The at least one output device 308 may include at least one of the following: audio output device (e.g., speaker), visual output device (e.g., display or detector), etc.

[0076] The memory 314 may include high-speed random access memory (such as DRAM, SRAM, DDR RAM, and / or other random access solid-state memory devices) and / or non-volatile memory (such as at least one disk storage device, optical disk storage device, flash memory device, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes at least one storage device located remotely from the control circuit 302. The memory 314 or alternatively at least one non-volatile solid-state memory device within the memory 314 includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures or subsets or supersets thereof:

[0077] · An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks;

[0078] · A network communication module 318, which is used to connect the server system 112 to other computing devices via at least one network interface 304 (e.g., via wired and / or wireless connections);

[0079] · A codec module 320 that is configured to perform various functions related to encoding and / or decoding data such as video data. In some embodiments, the codec module 320 is an example of the codec component 114. The codec module 320 includes at least one of the following:

[0080] o A decoding module 322 that is configured to perform various functions related to decoding encoded data, such as those functions previously described with respect to the decoder component 122; and

[0081] o An encoding module 340 that is configured to perform various functions related to encoding data, such as those functions previously described with respect to the encoder component 106; and

[0082] · A picture memory 352 that is configured to store pictures and picture data, e.g., for use with the codec module 320. In some embodiments, the picture memory 352 includes at least one of the following: a reference picture memory 208, a buffer memory 252, a current picture memory 264, and a reference picture memory 266.

[0083] In some embodiments, the decoding module 322 includes a parsing module 324 (e.g., configured to perform various functions previously described with respect to the parser 254), a transformation module 326 (e.g., configured to perform various functions previously described with respect to the scaler / inverse transform unit 258), a prediction module 328 (e.g., configured to perform various functions previously described with respect to the motion compensation prediction unit 260 and / or the intra picture prediction unit 262), and a filter module 330 (e.g., configured to perform various functions previously described with respect to the loop filter 256).

[0084] In some embodiments, the encoding module 340 includes a code module 342 (e.g., configured to perform various functions previously described with respect to the source coder 202 and / or the encoding engine 212) and a prediction module 344 (e.g., configured to perform various functions previously described with respect to the predictor 206). In some embodiments, the decoding module 322 and / or the encoding module 340 include Figure 3 a subset of the modules shown. For example, a shared prediction module is used by both the decoding module 322 and the encoding module 340.

[0085] Each of the above-identified modules stored in the memory 314 corresponds to a set of instructions for performing the functions described in this application. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, the codec module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, the memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, the memory 314 stores additional modules and data structures not described above.

[0086] Although Figure 3 FIG. illustrates a server system 112 according to some embodiments, Figure 3 it is more intended as a functional description of the various features that may exist in at least one server system rather than a structural schematic of the embodiments described in this application. In practice, the items shown separately may be combined, and some items may be separated. For example, Figure 3 some of the items shown separately in may be implemented on a single server, and a single item may be implemented by at least one server. The actual number of servers used to implement the server system 112 and how the features are allocated among them will vary depending on the implementation, and optionally partially depend on the data traffic processed by the server system during peak usage periods as well as during average usage periods.

[0087] Example encoding techniques

[0088] The encoding processes and techniques described below may be performed at the devices and systems described above (e.g., the source device 102, the server system 112, and / or the electronic device 120). A set of encoding processes / techniques involves block partitioning. In some encoding schemes, when encoding a picture, the entire picture is first divided into a plurality of non-overlapping blocks, which may be referred to as maximum coding blocks / units (or coding tree blocks or super blocks), and each maximum coding block may be further partitioned into smaller coding blocks according to a specific design of a block partitioning pattern, and each coding block is further encoded using a specific prediction pattern, transform coding, quantization, and / or entropy coding. Examples of the maximum coding block size include but are not limited to 64×64, 128×128, 256×256, and 512×512.

[0089] Hereinafter, a block (or sub-block) may refer to a coding block having a maximum coding block size (such as a super block, or a maximum coding unit, or a coding tree block), or a coding block, a prediction block, a transform block, a filtering unit, or a predefined fixed block size. As an example, a sub-block of block A is a block whose area is completely covered by block A.

[0090] Figures 4A through 4D illustrates an example coding tree structure according to some embodiments. As Figure 4A shown in the first coding tree structure (400) in, some coding methods use a 4-way partition tree starting from the 64′64 level and going down to the 4′4 level. For example, there are some additional restrictions for blocks of 8′8. In Figure 4A , the partition designated as "R" is recursive because the same partition tree repeats at a lower ratio until the lowest level is reached. As Figure 4B shown in the example coding tree structure (402) in, some coding methods extend the partition tree to a 10-way structure and increase the maximum size (e.g., sometimes referred to as a super block) to start from 128′128. The second coding tree structure includes 4:1 / 1:4 rectangular partitions not in the first coding tree structure. Figure 4B The partition type with 3 sub-partitions in the second row of is called a T-type partition. In addition to the coding block size, the coding tree depth can be defined to indicate the splitting depth from the root node.

[0091] As an example, the coding tree unit (CTU) can be split into coding units (CUs) by using a quadtree structure represented as a coding tree to adapt to various local characteristics. In some embodiments, a decision is made at the CU level on whether to use inter-frame (temporal) or intra-frame (spatial) prediction to encode a picture region. Each CU can be further split into one, two, or four prediction units (PUs) according to the PU split type. Inside the PU, the same prediction process can be applied, and relevant information can be transmitted to the decoder based on the PU. After obtaining the residual block by applying the prediction process based on the PU split type, the CU can be partitioned into transform units (TUs) according to another quadtree structure (such as the coding tree of the CU).

[0092] A quadtree with a nested multi-type tree using binary and ternary split structures can be used to replace the concept of multiple partition unit types. In the coding tree structure, the CU can have a square or rectangular shape. The CTU is first partitioned by a quadtree structure. The quadtree leaf nodes can be further partitioned by a multi-type tree structure. As Figure 4C shown in the third coding tree structure (404) in, the multi-type tree structure includes four split types. The multi-type tree leaf nodes are called CUs, and unless the CU is too large for the maximum transform length. This means that the CU, PU, and TU can have the same block size in a quadtree with a nested multi-type tree coding block structure. In Figure 4D an example of the block partition of a CTU (406) is shown, which illustrates an example quadtree.

[0093] Another set of encoding processes / techniques relates to motion estimation. As described above, motion estimation involves determining motion vectors that describe the transformation from one image (picture) to another image (picture) (e.g., using reference images (or blocks) from neighboring frames in a video sequence). Motion vectors can relate to the entire image (global motion estimation) or to specific blocks. Additionally, motion vectors can correspond to translational models or warping models that approximate motion (e.g., rotation, translation, and scaling in three dimensions). In some cases (e.g., for more complex video objects), motion estimation can be improved by further partitioning the blocks.

[0094] In MMV, different views can have strong correlations, and reducing the statistical redundancy present between different views helps improve the coding and decoding efficiency. As described in detail below, some MMV techniques utilize the block partition information of a first block in a first picture of a first view to perform block partition and encoding of a second block in a second picture of a second view. The first block and the second block can be located at the same coordinates in the first view and the second view, respectively, and the first picture and the second picture can be associated with the same display time.

[0095] Figure 5 Illustrated is an example MVV500 having two views (e.g., view 0 and view Figure 1 ) according to some embodiments. Each view can be associated with a different viewport or camera. For applications such as stereoscopic video viewing, video of more than one view can be encoded. In some embodiments, MVV 500 corresponds to a 3D scene captured by two or more cameras. In some cases, optional processing, such as view correction and color correction, is performed on the transmitter side. After encoding the MVV sequence, the bitstream is transmitted to the receiver side, where the views are decoded and presented on a suitable 3D display. In some embodiments, MVV includes more than two views.

[0096] Figure 6 Illustrated is an example prediction structure 600 for MVV according to some embodiments. Structure 600 uses temporal reference pictures (represented by horizontal and curved arrows) and inter-view reference pictures (represented by vertical arrows) for motion compensation prediction and disparity compensation prediction. Figure 6Two picture sequences are shown, corresponding to a left view and a right view, where due to the strong correlation between the two views, the pictures in the left view are used to predict the pictures in the right view, thereby improving the encoding and decoding efficiency. Picture 602 in the right view (POC 0) is a P picture / frame encoded (e.g., predicted) using picture 604 in the left view as a reference picture. Picture 604 is an I picture. The next frame encoded in the right view is picture 606 (POC 8), corresponding to the last P frame in the sequence. Then picture 610 (POC 4) in the right view is derived using picture 602 and picture 606. Picture 610 is a bi-directional B picture / frame. Then POC 0, POC 8, and POC 4 in the obtained right view are used to derive POC 2 and POC 6 in the right view. In some embodiments, the pictures in the right view are in multiple layers. For example, the odd POCs (denoted by the lowercase letter "b") in the right view are in different layers from the even POCs.

[0097] Figure 7 Video data 700A and 700B are illustrated, each including a plurality of views (e.g., views 0 to view Figure 5 ) that are spatially stitched together to form a two-dimensional image. Although Figure 7 six views in the video data are depicted, it should be understood that any number of views can be stitched together. There can be several ways to stitch the views spatially. For example, for six views, one, two, or three views can be stitched in each row, which can result in a 1×6 stitch, a 2×3 stitch, and a 3×2 stitch, respectively. In some embodiments, the stitching is designed such that the resulting super-sized picture has a desired picture size. For example, the super-sized picture can be close to a square shape or a rectangular shape with a 4:3 aspect ratio or a 16:9 aspect ratio.

[0098] For P stripes or B stripes in the super-sized picture, the motion vectors from the previously encoded views in the same picture are highly correlated with the motion vectors in the current view being encoded. Therefore, the motion vectors from the previously encoded views in the same picture are very suitable for motion vector prediction or as a starting point in the motion estimation of the current view. Additionally, in some embodiments, using the perspective transformation between two views (e.g., from a reference view to the current view), a motion vector predictor (MVP) can be calculated. In some embodiments, an MVP candidate is derived for the current block 702 in the current view (e.g., view 2) of the current picture (Pcur). In some embodiments, block 702 is mapped to block 704 in the reference view (Vref) of Pcur. If the motion vector of block 704 is (Mx, My), where its reference block 706 resides in the same view (Vref) of the reference picture (Pref), then X2 = X3 + M xand Y2 = Y3 + My, where (X2, Y2) are the coordinates of reference block 706, (X3, Y3) are the coordinates of block 708, where block 708 is the corresponding block of block 704. Both block 708 and block 704 can be in the reference view Vref (e.g., view 0), and block 708 can reside in Pref.

[0099] In Figure 7 , block 710 is the reference block of block 702, block 712 is the corresponding block of block 702, and block 712 can reside in Vcur of Pref. In some embodiments, the coordinates of block 710 are derived based on the coordinates of block 712 and the MVP candidate for the calculation (e.g., derivation) for block 702.

[0100] In some embodiments, assuming that the samples in a block share the same disparity, a disparity vector DV (Dx, Dy) is used to find the corresponding block of the current block in the same picture in the reference view. In some embodiments, a position offset is established between the current view and its reference view (e.g., the offset may be twice the view width in the x - direction and 0 in the y - direction). The disparity vector can be added to the view offset to find the corresponding block of the current block in the reference view. Since the corresponding block indicated by the disparity vector in the reference view (e.g., view 0) may already be encoded, its motion vector (if any) can point to the reference block in view 0 of the reference picture. In some embodiments, the reference block is used as the reference block for the current block (in view 2), or as a starting point in motion estimation.

[0101] Figure 8 FIG. illustrates video data 800 with multiple views according to some embodiments. The video data 800 includes a number of views (e.g., view 0 to view Figure 5 ) that are spatially stitched together to form a two - dimensional image. A block vector (BV) can point to a reference block in a previously encoded view. In some embodiments, a perspective transformation between two views (e.g., view 0 and view Figure 5 ) is estimated before block matching. The perspective transformation can establish a bijective mapping between the two views. The perspective transformation can be applied to the reference view (e.g., view 0), which maps the coordinates from the reference view to the current view being encoded (e.g., view Figure 5 ). The "transformed" reference view can be used as a reference for performing block matching.

[0102] For a block in the current view (e.g., block 808A), its corresponding block (e.g., block 804) can be found at the same position (e.g., the same coordinates) in the "transformed" reference view, and the block matching (e.g., with block 802B) starts with the corresponding block as the starting point. Since the perspective transformation is very close to the transition between the two views, the block matching can be restricted to a small neighborhood of the corresponding block in the "transformed" reference view. In some embodiments, the current block and its corresponding block have the same position offset relative to the upper left position of their corresponding views. The perspective transformation can be estimated by analyzing the reference view and the current view. The perspective transformation can also be derived using techniques from computational photography (such as keypoint detection and matching), or directly calculated based on camera parameters and depth map data.

[0103] In some embodiments, assuming that the samples in a block share the same disparity, a disparity vector DV (Dx, Dy) is used to indicate the disparity between the corresponding block in the reference view and the reference block in the reference view. Note that for each view pair, the block vector (BV) can be different for blocks at different positions in the view. In the reference view, the block vector pointing from the current block to its reference block consists of two parts: the view position offset plus the disparity vector.

[0104] Figure 9A Illustrates the calculation of a predicted block according to some embodiments. In Figure 9A the example, intra-frame prediction is performed on the current block 902 to generate a predicted block 904. In some embodiments, inter-frame prediction is performed to generate a predicted block. The current block 902 includes a set of samples (e.g., a pixel block), and the predicted block 904 includes a set of predictions corresponding to the set of samples. Figure 9B Illustrates the calculation of a residual block according to some embodiments. As Figure 9B shown, the predicted block 904 is subtracted from the current block 902 to generate a residual block 906 including a set of residuals. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 9C Illustrates the calculation of a reconstructed block according to some embodiments. As Figure 9C shown, the residual block 906 undergoes one or more transforms and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 908. The reconstructed residual block 908 is combined with the predicted block 904 (e.g., adding the reconstructed residuals of the reconstructed residual block 908 to the predictions of the predicted block 904) to generate a reconstructed block 910 corresponding to the current block 902.

[0105] Figure 10AFIG. is a flowchart of a method 1000 for decoding video according to some embodiments. Method 1000 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, method 1000 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0106] The system receives (1002) a multi-view video bitstream that includes a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view. The system determines (1004), based on a first indicator in the multi-view video bitstream, whether to signal jointly or separately a plurality of block partitioning patterns for the first block corresponding to the first view and the second block corresponding to the second view. When the first indicator indicates signaling the plurality of block partitioning patterns jointly, the same block partitioning pattern is applied (1006) to the first block and the second block according to a second indicator.

[0107] In some embodiments, the block partitioning patterns are signaled jointly for a plurality of blocks belonging to different views. In some embodiments, the blocks are co-located blocks (e.g., blocks located at the same coordinates in different views). In some embodiments, the position of a block depends on the disparity of the view. In some cases, the disparity between views is quantized to certain predefined values. These values may be powers of 2 (or 4), such as 4, 8, 16.

[0108] In some embodiments, a flag is signaled for blocks from different views to indicate whether to signal the block partitioning patterns jointly or separately. In some embodiments, the flag is signaled conditionally. For example, the flag is signaled when a block is larger (or smaller) than a predefined threshold (e.g., 128×128, 64×64, 32×32, 16×16, or 8×8). In some embodiments, when the block partitioning patterns are signaled jointly, the number of signaling syntaxes related to the partitioning pattern is less than the number of blocks. For example, if N blocks (where N>1) share the same partitioning pattern, the partitioning pattern is signaled once (for the N blocks) instead of N times. In some embodiments, when the block partitioning patterns are signaled jointly, for a given first depth value, the block partitioning pattern is shared, and beyond the first depth value, the block partitioning patterns are signaled separately.

[0109] In some embodiments, when signaling block partition modes for multiple blocks from different views, the prediction modes of the blocks (e.g., intra prediction or inter prediction and their directions) are also signaled jointly. In some embodiments, the partition mode specifies how to partition a block into smaller block sizes. The partition mode may also specify whether different color components share the same partition.

[0110] In some embodiments, a context for signaling the block partition mode of a second block from a second view is derived using the block partition mode of a first block from a first view. This context provides a better estimate of the bit probabilities having a certain value, thereby improving the coding and decoding efficiency. In some embodiments, the same context is shared for encoding the same type of syntax for different views. For example, when encoding the block partition mode, the context associated with signaling the block partition mode for a first view can be further updated by signaling the block partition mode for a second view, and vice versa.

[0111] In some embodiments, instead of encoding different pictures from different views separately, the signaling of the syntax is interleaved. For example, the syntax belonging to a second view can be signaled between multiple syntaxes signaled for a first view. In some embodiments, the signaling of the largest coding units (LCUs) for different views is interleaved. An example coding order can be LCU0 of view 0, LCU0 of view Figure 1 of view 0, LCU1 of view 0 and LCU1 of view Figure 1 of view 0. The advantage of interleaving the signaling of the largest coding units (LCUs) is that context updates are potentially more efficient.

[0112] In some embodiments, for different views, only part of the syntax is signaled in an interleaved manner; for different views, part of the syntax is still signaled separately without interleaving. In some embodiments, the syntax related to block partitioning is signaled separately for different views; for different views, other syntax is signaled in an interleaved manner. In some embodiments, the syntax related to loop filtering is signaled separately for different views; for different views, other syntax is signaled in an interleaved manner. In some embodiments, the interleaving of the syntax signaling is performed in different units, including, for example, tiles, strips, largest coding units, largest coding unit rows, coding blocks, transform blocks, prediction blocks, or predefined block sizes. In some embodiments, the interleaving of the syntax signaling is performed conditionally (e.g., when the first picture and the second picture are of the same type (e.g., intra or inter)).

[0113] In some embodiments, for multiple pictures from different views, loop filter parameters associated with one or more loop filter methods are signaled jointly. Example loop filter methods include cross-component filtering, cross-component offset filtering, Wiener loop filtering, deblocking filtering, and Constrained Direction Enhancement Filter (CDEF) method. In some embodiments, high-level syntax is signaled to indicate whether loop filter parameters are signaled jointly for multiple pictures from different views or signaled separately for multiple pictures from different views. In some embodiments, high-level syntax for signaling the enabling of one or more loop filter methods is signaled jointly for multiple views.

[0114] In some embodiments, picture-level on / off flags for one or more loop filter methods are signaled jointly for multiple views. In some embodiments, the picture-level on / off flag for a loop filter method in one view depends on the picture-level on / off flag for the loop filter method in another view. In one example, for a first picture in a first view, three flags can be signaled for different color components to specify whether a first loop filter method is applied to the three color components separately. Then, for a second view, a flag is signaled to indicate whether the values of the three flags are inherited for a second picture in the second view. In another example, for a first picture in a first view, multiple flags are signaled separately to specify whether each of multiple loop filter methods is enabled. Then, for a second picture in a second view, a single flag is signaled regardless of whether the enabling of the multiple loop filter methods is inherited from the first picture in the first view or signaled separately for the second picture in the second view.

[0115] In some embodiments, parameters used in a loop filter method are signaled jointly for multiple views. In some embodiments, parameters for the cross-component offset filtering method are signaled jointly for multiple views. The parameters can include a selection between an offset lookup table, offset-only, edge-only offset, and band-edge combined offset, filtering shape, quantization step, downsampling filter, and filtering unit size. In some embodiments, loop filter parameters for the Wiener loop filtering method are signaled jointly for multiple views. The parameters can include a selection of filtering parameters, filtering shape, and / or size and / or filtering unit size. In some embodiments, loop filter parameters for the CDEF method are signaled jointly for multiple views.

[0116] In some embodiments, a signaling flag is used to indicate whether one or more loop filter parameters are shared from a first picture in a first view to a second picture in a second view. In some embodiments, a signaling flag is used to indicate whether one or more loop filter parameters are predicted from a first picture in a first view to a second picture in a second view. In some embodiments, for a second picture in a second view, a prediction residual of one or more loop filter parameters is signaled.

[0117] In some embodiments, partial parameters of a loop filter are shared between two views, while the remaining parameters are signaled jointly, dependently, or independently for the two views. For example, the on / off flag for the two views is signaled separately, and the set of parameters for controlling the loop filter for the two views is signaled independently.

[0118] Some embodiments apply loop filtering to the reconstruction of a first picture from a first view, using the reconstructed picture of a second picture from a second view as one of the inputs in the loop filtering process (e.g., cross-view loop filtering). In some embodiments, the input to the loop filter includes the reconstruction of the first picture in the first view, and the output includes an offset to be added on top of the reconstruction of the second picture in the second view. In some embodiments, the input of the original loop filtering method used in a single view is also used together as the input to the loop filtering method.

[0119] In some embodiments, the first component of the reconstruction of a first picture in a first view is used as the input for loop filtering the second component of the reconstruction of a second picture in a second view. The first component and the second component may be different components. In some embodiments, the input from the second view is the samples surrounding the co-located sample of the samples in the first view to be filtered. In some embodiments, the input from the second view is the samples located by a disparity vector.

[0120] In some embodiments, the disparity vector is derived from disparity vectors associated with adjacent blocks that use different views as reference pictures, e.g., inter-view prediction. In some embodiments, the accuracy of the disparity vector includes a predefined accuracy, e.g., integer pixel, half pixel, or quarter pixel.

[0121] In some embodiments, the filtering order of pictures from different views is different from the coding order of pictures from different views. For example, for residual coding, the first picture from the first view is encoded first, and then the second picture from the second view is residual-encoded. The reconstruction of the second picture is loop-filtered before the reconstruction of the first picture is loop-filtered.

[0122] In some embodiments, the reconstructed samples from two views are jointly used as the input to the loop filter for each view. For example, the reconstructed samples from two views are used to derive the filter parameters for each individual view.

[0123] In some embodiments, one or more motion vectors derived from the encoded information associated with another view are used to derive the MVP of a block in the current view. In some embodiments, the motion vector associated with the first encoded block of the first picture in the first view, i.e., the view-based motion vector predictor (VMVP), is used as the MVP of the second encoded block of the second picture in the second view. In some embodiments, a disparity vector is used to obtain the first encoded block, and the disparity vector is derived from adjacent blocks encoded using the second view as a reference frame. In one example, the disparity vector is (0, 0), meaning that the first encoded block and the second encoded block are located at the same coordinates in the first view and the second view, respectively. In another example, the disparity vector is derived from a global disparity vector associated with the combination between two frames simultaneously displayed from different views.

[0124] In some embodiments, one or more reference frames belonging to the first view are used to encode the first encoded block. The display times of one or more reference frames of the first view are compared with the display times of the reference frames of the second block in the second view. If the display times are the same, the motion vector associated with the first encoded block is used as the motion vector predictor for the second encoded block. In some embodiments, one or more reference frames belonging to the first view are used to encode the first encoded block. The display time of the reference frame is different from the display time of the reference frame of the second block in the second view. The motion vector associated with the first encoded block is scaled and used as the motion vector predictor for the second encoded block.

[0125] In some cases, the scaling factor is proportional to the ratio between the temporal distance between the first picture and its reference frame in the first view and the temporal distance between the second picture and its reference frame in the second view.

[0126] In some embodiments, the MVP of the current block is constructed. The checking order of the Temporal MVP (TMVP) or VMVP candidate relative to other motion vector predictor candidates (e.g., spatial motion vector predictor candidate, global motion vector predictor candidate, motion vector predictor candidate in the motion vector library) is different. In some embodiments, both the TMVP and the VMVP are in the motion vector list. Each is associated with a different index in the motion vector list. In some embodiments, the motion vector list has a predefined fixed (maximum) number of TMVP and / or VMVP. The TMVP / VMCP candidates in the list depend on the checking order and availability of the TMVP and VMVP candidates. In an example, two lists are constructed, namely the MVP list from the multi-view and the list from the current frame. A high-level flag or a block-level flag is signaled to indicate which list to use.

[0127] In some embodiments, more than one motion vector obtained from a plurality of coded blocks in the first picture of the first view is used as the MVP candidate for the second block in the second picture of the second view. The obtained motion vectors can be added to the MVP list and used as the MVP for the second block.

[0128] In some embodiments, the coded blocks are non-adjacent coded blocks relative to the first coded block in the first picture in the first view. In some embodiments, a disparity vector is used to determine the first coded block, and the disparity vector is derived from adjacent blocks of the second block encoded using the second view as a reference frame. In some embodiments, the disparity vector is (0,0), for example, the first coded block and the second coded block are located at the same coordinates in the first view and the second view, respectively.

[0129] In some embodiments, the disparity vector is derived from a global disparity vector associated with the combination between two frames simultaneously displayed from different views.

[0130] In some embodiments, if the motion vectors of one or more spatial (or temporal) adjacent blocks point to the first view, the motion vectors of the adjacent blocks are also inserted into the MVP list of the current block. In some embodiments, for the current block, additional motion vectors are signaled explicitly or derived implicitly. The motion vector is used to indicate the position displacement in the reference picture of the second view. In some embodiments, the coordinates of the coded block are predefined or derived implicitly using the coded information such as block shape, quantization parameter, and / or temporal layer.

[0131] In some embodiments, a motion vector library associated with the blocks in the first picture of the first view is used to derive the motion vector predictor for the blocks in the second picture of the second view. In some embodiments, the motion vector library associated with the blocks in the first picture of the first view is merged with the motion vector library associated with the blocks in the second picture of the second view.

[0132] In some embodiments, the use or enabling of the above embodiments is controlled by high-level syntax, including but not limited to sequence-level, picture-level, sub-picture-level, slice-level, tile-level, and maximum coded block-level flags.

[0133] Figure 10B FIG. is a flowchart illustrating a method 1050 for encoding video according to some embodiments. Method 1050 may be performed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions for execution by the control circuit. In some embodiments, method 1050 is performed by executing instructions stored in the memory (e.g., memory 314) of the computing system.

[0134] The system receives (1052) video data including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view. The system determines (1054) whether to signal jointly or separately for a plurality of block partitioning patterns for the first block and the second block. When it is determined to signal jointly for the plurality of block partitioning patterns, a first indicator is signaled (1056) in the video bitstream, the first indicator indicating that the plurality of block partitioning patterns are signaled jointly. As previously mentioned, the encoding process may mirror the decoding process described in this application. For the sake of brevity, these details are not repeated here.

[0135] Although Figure 10A and Figure 10B FIG. illustrates a number of logical stages in a particular order, stages that are not order-dependent may be reordered, and other stages may be combined or decomposed. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, and thus the orderings and groupings presented in this application are not exhaustive. Additionally, it should be recognized that these stages may be implemented in hardware, firmware, software, or any combination thereof.

[0136] Now turning to some example embodiments:

[0137] (A1)In one aspect, some embodiments include a method for video decoding (e.g., method 1000). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at a codec module (e.g., codec module 320). In some embodiments, the method is performed at a source coding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a multi-view video bitstream, the multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view; (ii) based on a first indicator in the multi-view video bitstream, determining whether to signal jointly or separately a plurality of block partitioning patterns for the first block corresponding to the first view and the second block corresponding to the second view; and (iii) when the first indicator indicates signaling the plurality of block partitioning patterns jointly, applying the same block partitioning pattern to the first block and the second block according to a second indicator. For example, block partitioning patterns are signaled jointly for a plurality of blocks belonging to different views. In some embodiments, the multi-view video bitstream includes video from two or more cameras (e.g., offset relative to each other). In some embodiments, a flag is signaled for a plurality of blocks from different views to indicate whether to signal these block partitioning patterns jointly or separately. In some embodiments, the first frame and the second frame are temporally synchronized (e.g., captured and / or streamed simultaneously). In some embodiments, when the first indicator indicates signaling the block partitioning patterns separately, applying a first block partitioning pattern to the first block and a second block partitioning pattern to the second block, where the first block partitioning pattern is different from the second block partitioning pattern.

[0138] (A2)In some embodiments of A1, the first block is located at a set of coordinates in the first view, and the second block is located at the same set of coordinates in the second view. For example, the plurality of blocks are corresponding blocks (e.g., blocks located at the same coordinates in different views).

[0139] (A3)In some embodiments of A1 or A2, the first block is located at a first set of coordinates in the first view, and the second block is located at a second set of coordinates in the second view, the second set of coordinates being offset relative to the first set of coordinates. For example, the positions of the plurality of blocks may depend on the disparity of the plurality of views.

[0140] (A4)In some embodiments of A3, the offset between the first set of coordinates and the second set of coordinates is determined according to a predefined quantization disparity value. For example, the disparity between multiple views is quantized to certain predefined values, such as powers of 2 (or 4), for example, 4, 8, 16.

[0141] (A5)In some embodiments of any one of A1 to A4, the first frame includes a first plurality of blocks, and the second frame includes a second plurality of blocks; for the first plurality of blocks and the second plurality of blocks, a first block partitioning mode is signaled jointly. For example, when signaling the block partitioning mode jointly, the number of syntaxes signaled related to the partitioning mode is less than the number of blocks, that is, N (N>1) blocks share the same partitioning mode, and the partitioning mode is signaled once instead of signaling the partitioning mode for these N blocks separately N times.

[0142] (A6)In some embodiments of any one of A1 to A5, when the first indicator indicates signaling the plurality of block partitioning modes jointly, (i) a shared partitioning mode of a first depth is applied to the first block and the second block; and (ii) when exceeding the first depth, different partitioning modes are applied to the first block and the second block. For example, when signaling the block partitioning mode jointly, for a given first depth value, these block partitioning modes are shared, and when exceeding the first depth value, these block partitioning modes are signaled separately. In some embodiments, according to the indicator indicating signaling the block partitioning mode jointly, a shared partitioning mode of the first depth is applied to the first block and the second block, and separate partitioning modes are applied to the first block and the second block exceeding the first depth.

[0143] (A7)In some embodiments of any one of A1 to A6, the first indicator is signaled conditionally. For example, this flag is signaled only conditionally. Examples of conditions include whether the block is greater than (or less than) a predefined threshold, such as 128×128, 64×64, 32×32, 16×16, or 8×8.

[0144] (A8)In some embodiments of any one of A1 to A7, the method further includes determining whether to signal jointly or separately the prediction mode for the first block and the second block based on a third indicator in the video bitstream. For example, when signaling the block partitioning mode for multiple blocks from different views, the prediction mode can also be signaled jointly (for example, whether these multiple blocks are encoded by intra prediction or inter prediction, the intra prediction direction, and / or the inter prediction mode).

[0145] (A9)In some embodiments of any one of A1 to A8, the method further includes determining a context for signaling a first block partition mode for a first block based on a second block partition mode for a second block. For example, the block partition mode of the first block from the first view can be used to derive the context for signaling the second block partition mode from the second view. In some embodiments, the first block partition mode is entropy coded according to a first context, and the second block partition mode is entropy coded according to a second context.

[0146] (A10)In some embodiments of any one of A1 to A9, the same block partition mode includes instructions for indicating partitioning a block into a plurality of sub-blocks, and / or instructions for indicating whether different color components of the block share the same partition. For example, the partition mode can refer to how to partition a block into smaller block sizes, and the partition mode can also refer to whether different color components share the same partition.

[0147] (A11)In some embodiments of any one of A1 to A10, in the multi-view video bitstream, a first set of syntax elements corresponding to the first view and a second set of syntax elements corresponding to the second view are interleaved. For example, instead of encoding different pictures from different views separately, the signaling of the syntax is interleaved. That is, the syntax belonging to the second view can be signaled between multiple syntaxes signaled for the first view. Interleaving the syntax elements from different views allows for more efficient context updates (e.g., for contexts that depend on information from another view). In some embodiments, the first set of syntax elements and the second set of syntax elements are conditionally interleaved. For example, the interleaving of the syntax signaling can be performed conditionally. The conditions can include whether the first picture and the second picture are of the same type (e.g., only intra or inter) and / or share the same pattern.

[0148] (A12)In some embodiments of A11, at the largest coding unit (LCU) level, the first set of syntax elements and the second set of syntax elements are signaled. For example, the signaling of the largest coding units (LCUs) for different views is interleaved. That is, the coding order of the largest coding units is as follows: LCU0 of view 0, LCU0 of view Figure 1 view, LCU1 of view 0, LCU1 of view Figure 1 view. In this way, more efficient context updates can be achieved. In some embodiments, the syntax elements corresponding to a specific level / unit are interleaved. For example, the interleaving of the syntax signaling can be performed in different units, such as tiles, strips, largest coding units, largest coding unit rows, coding blocks, transform blocks, prediction blocks, or predefined block sizes.

[0149] (A13)In some embodiments of A11 or A12, a first syntax element in the first set of syntax elements shares the same context with a second syntax element in the second set of syntax elements. For example, the same context can be shared for encoding the same type of syntax for different views. As an example, when encoding a block partitioning pattern, the context associated with signaling the block partitioning pattern for a first view can be further updated by signaling the block partitioning pattern for a second view, or vice versa.

[0150] (A14)In some embodiments of any one of A11 to A13, a first subset of the first set of syntax elements is interleaved with the second set of syntax elements, and a second subset of the first set of syntax elements is not interleaved with the second set of syntax elements. For example, for different views, only a part of the syntax is signaled in an interleaved manner, and for different views, a part of the syntax is still signaled separately without any interleaving. In some embodiments, the syntax elements related to a first mode / operation are interleaved, while the syntax elements related to different modes / operations are not interleaved. For example, the syntax related to block partitioning is signaled separately for different views, and other syntax is signaled in an interleaved manner for different views. As another example, the syntax related to loop filtering is signaled separately for different views, and other syntax is signaled in an interleaved manner for different views.

[0151] (B1)On the other hand, some embodiments include a method for video encoding (e.g., method 1050). In some embodiments, the method is executed at a computing system having a memory and one or more processors. The method includes: (i) receiving video data, the video data including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view; (ii) determining whether to signal jointly or separately a plurality of block partitioning patterns for the first block and the second block; and (iii) when it is determined to signal jointly the plurality of block partitioning patterns, signaling, in a video bitstream, a first indicator that indicates signaling jointly the plurality of block partitioning patterns.

[0152] In some embodiments, the first block and the second block are partitioned using the same partitioning pattern, and the same partitioning pattern is signaled jointly for the first block and the second block. In some embodiments, a second indicator is signaled in the video bitstream according to a determination to jointly signal the block partitioning pattern, the second indicator indicating the block partitioning pattern for the first block and the second block. In some embodiments, a second indicator is signaled in the video bitstream according to a determination to separately signal the block partitioning pattern, the second indicator indicating the block partitioning pattern for the first block, and a third indicator is signaled in the video bitstream, the third indicator indicating the block partitioning pattern for the second block.

[0153] (B2) In some embodiments of B1, the method includes (i) determining whether to jointly signal or separately signal a plurality of prediction patterns for the first block and the second block; and (ii) when determining to jointly signal the plurality of prediction patterns, signaling in the video bitstream a second indicator indicating that the plurality of prediction patterns are jointly signaled.

[0154] (B3) In some embodiments of B1 or B2, the method includes signaling in the video bitstream a first set of syntax elements corresponding to the first view, the first set of syntax elements being interleaved with a second set of syntax elements corresponding to the second view.

[0155] (C1) In another aspect, some embodiments include a method for visual media data processing. In some embodiments, the method is performed at a computing system having a memory and one or more processors. The method includes: (i) obtaining a source video sequence, the source video sequence including a set of views; and (ii) performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, wherein the multi-view video bitstream includes: (a) a first plurality of encoded pictures corresponding to a first view; (b) a second plurality of encoded pictures corresponding to a second view; (c) a first indicator indicating whether to jointly signal or separately signal a block partitioning pattern for the first block and the second block; (d) when the first indicator indicates jointly signaling the block partitioning pattern, a second indicator indicating the block partitioning pattern; (e) when the first indicator indicates separately signaling the block partitioning pattern, a set of indicators indicating the respective block partitioning patterns of the first block and the second block.

[0156] (C2) In some embodiments of C1, the multi-view video bitstream further includes a third indicator indicating whether to jointly signal a prediction pattern for the first block and the second block.

[0157] (C3) In some embodiments of C1 or C2, in the multi-view video bitstream, a first set of syntax elements corresponding to the first view and a second set of syntax elements corresponding to the second view are interleaved.

[0158] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including control circuitry (e.g., control circuitry 302) and a memory coupled to the control circuitry (e.g., memory 314), the memory storing one or more sets of instructions configured to be executed by the control circuitry, the one or more sets of instructions including instructions for performing any of the methods described in this application (e.g., A1 - A1 - A14, B1 - B3, and C1 - C3 above).

[0159] In another aspect, some embodiments include a non - volatile computer - readable storage medium storing one or more sets of instructions, the instructions being executed by control circuitry of a computing system, the instructions including instructions for performing any of the methods described in this application (e.g., A1 - A14, B1 - B3, and C1 - C3 above).

[0160] Unless otherwise specified, any syntax element described in this application may be high - level syntax (HLS). As used in this application, the signal level of HLS is higher than the block level. For example, HLS may correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements may be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), slice header, picture header, tile header, and / or CTU header.

[0161] The proposed methods can be used alone or in any order of combination. In addition, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). For example, one or more processors execute a program stored in a non - transient computer - readable medium. Hereinafter, the term block may be interpreted as a prediction block, coding block, or coding unit, i.e., a CU.

[0162] It should be understood that although the terms "first", "second", etc. may be used in this application to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used in this application refers to and encompasses any and all possible combinations of at least one of the associated listed items. It will be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of at least one other feature, integer, step, operation, element, component, and / or group thereof.

[0163] As used in this application, depending on the context, the term "if" can be interpreted to mean "when", "at the time of", "in response to determining", "in accordance with determining", or "in response to detecting" that the stated prerequisite is true. Similarly, the phrases "if it is determined [that the stated prerequisite is true]" or "if [the stated prerequisite is true]" or "when [the stated prerequisite is true]" can be interpreted to mean "at the time of determining", "in response to determining", "in accordance with determining", "at the time of detecting", or "in response to detecting" that the stated prerequisite is true, depending on the context.

[0164] For purposes of explanation, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and practical applications, thereby enabling others skilled in the art to implement.

Claims

1. A video decoding method, characterized in that, Performed by a computer system, the computer system including a memory and one or more processors, the method comprising: Receiving a multi-view video bitstream, the multi-view video bitstream including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view; Based on a first indicator in the multi-view video bitstream, determining whether to signal jointly or separately a plurality of block partitioning patterns for the first block corresponding to the first view and the second block corresponding to the second view; and, When the first indicator indicates signaling the plurality of block partitioning patterns jointly, applying the same block partitioning pattern to the first block and the second block according to a second indicator.

2. The method according to claim 1, characterized in that, The first block is located at a set of coordinates in the first view, which is the same as a set of coordinates where the second block is located in the second view.

3. The method according to claim 1, characterized in that, The first block is located at a first set of coordinates in the first view, and the second block is located at a second set of coordinates in the second view.

4. The method according to claim 3, wherein The offset between the first set of coordinates and the second set of coordinates is determined according to a predefined quantization disparity value.

5. The method according to claim 1, wherein The first frame includes a first plurality of blocks, and the second frame includes a second plurality of blocks; a first block splitting pattern is signaled jointly for the first plurality of blocks and the second plurality of blocks.

6. The method according to claim 1, characterized in that When the first indicator indicates signaling the plurality of block partitioning patterns jointly, Applying a shared partitioning pattern of a first depth to the first block and the second block; When exceeding the first depth, applying different partitioning patterns to the first block and the second block.

7. The method according to claim 1, wherein Signaling the first indicator conditionally.

8. The method according to claim 1, further comprising: Based on a third indicator in the video bitstream, determining whether to signal jointly or separately a plurality of prediction patterns for the first block and the second block.

9. The method according to claim 1, further comprising: Based on a second block partitioning pattern of the second block, determining a context for signaling a first block partitioning pattern of the first block.

10. The method according to claim 1, characterized in that The same block partitioning pattern includes instructions for indicating dividing the block into a plurality of sub-blocks, and / or instructions for indicating whether different color components of the block share the same partitioning.

11. The method according to claim 1, wherein In the multi-view video bitstream, a first set of syntax elements corresponding to the first view and a second set of syntax elements corresponding to the second view are interleaved.

12. The method according to claim 11, wherein At the largest coding unit (LCU) level, signaling the first set of syntax elements and the second set of syntax elements.

13. The method according to claim 11, wherein A first syntax element in the first set of syntax elements and a second syntax element in the second set of syntax elements share the same context.

14. The method according to claim 11, wherein A first subset of the first set of syntax elements is interleaved with the second set of syntax elements, and a second subset of the first set of syntax elements is not interleaved with the second set of syntax elements.

15. A computing system, characterized in that, Comprising: A control circuit; A memory; One or more sets of instructions stored in the memory and configured to be executed by the control circuit, the one or more sets of instructions including instructions for: Receiving video data, the video data including a plurality of blocks, the plurality of blocks including a first block of a first frame and a second block of a second frame, the first frame corresponding to a first view and the second frame corresponding to a second view; Determining whether to signal jointly or separately a plurality of block partitioning patterns for the first block and the second block; And, When it is determined to signal jointly the plurality of block partitioning patterns, signaling, in a video bitstream, a first indicator that indicates signaling jointly the plurality of block partitioning patterns.

16. The computing system according to claim 15, further comprising: Determining whether to signal jointly or separately a plurality of prediction patterns for the first block and the second block; When it is determined to signal jointly the plurality of prediction patterns, signaling, in the video bitstream, a second indicator that indicates signaling jointly the plurality of prediction patterns.

17. The computing system according to claim 15, further comprising: Signaling, in the video bitstream, a first set of syntax elements corresponding to the first view, the first set of syntax elements being interleaved with a second set of syntax elements corresponding to the second view.

18. A non-volatile computer-readable storage medium, characterized in that, Storing one or more sets of instructions, the one or more sets of instructions being configured to be executed by a computing device having a control circuit and a memory, the one or more sets of instructions including instructions for: Obtaining a source video sequence, the source video sequence including a set of views; Performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, wherein the multi-view video bitstream includes: A first plurality of encoded pictures corresponding to a first view; A second plurality of encoded pictures corresponding to a second view; A first indicator that indicates whether to signal jointly or separately a block partitioning pattern for the first block and the second block; When the first indicator indicates signaling jointly the block partitioning pattern, a second indicator that indicates the block partitioning pattern; When the first indicator indicates signaling separately the block partitioning pattern, a set of indicators that indicate the block partitioning patterns of the first block and the second block, respectively.

19. The non-volatile computer-readable storage medium according to claim 18, wherein The multi-view video bitstream further includes a third indicator that indicates whether to signal jointly a prediction pattern for the first block and the second block.

20. The non-volatile computer-readable storage medium according to claim 18, wherein In the multi-view video bitstream, a first set of syntax elements corresponding to the first view and a second set of syntax elements corresponding to the second view are interleaved.