Systems and methods for loop filtering for multi-view coding
By adopting the parallax compensation prediction method in multi-view video encoding and jointly sending loop filtering parameters, the problem of unused redundancy between views in the prior art is solved, and more efficient coding efficiency and quality are achieved.
Patent Information
- Application Number
- CN202480005405.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-07
- Filing Date
- 2024-05-16
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing multi-view video encoding technology, the independent transmission of loop filter parameters of each view leads to low encoding efficiency and fails to make full use of statistical redundancy between views.
The parallax compensation prediction method is adopted to jointly transmit loop filtering parameters, and the pictures in multi-view videos are subjected to joint loop filtering to reduce statistical redundancy between views.
Improve coding efficiency, achieve a bit rate saving of about 70%, and improve coding quality.
Smart Images

Figure CN120359743A_ABST
Abstract
Description
Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 601,177, filed on November 20, 2023, entitled "Loop Filtering for Multi-View Coding", which is a continuation-in-part of and claims the benefit of priority of U.S. Patent Application No. 18 / 657,710, filed on May 7, 2024, entitled "Systems and Methods for Loop Filtering for Multi-View Coding". Technical Field
[0002] Embodiments disclosed herein generally relate to video coding and decoding, including but not limited to systems and methods for loop filter design for multiview video (MVV) coding and decoding. Background Art
[0003] Digital video is supported by various electronic devices, such as digital televisions, laptop or desktop computers, tablets, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices send and receive or otherwise transmit digital video data over communication networks and / or store digital video data on storage devices. Since the bandwidth capacity of communication networks and the storage resources of storage devices are limited, video coding can be used to compress video data according to one or more video coding standards before transmitting or storing the video data. Video coding can be performed by hardware and / or software on an electronic device / client device or by a server providing cloud services.
[0004] Video coding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that can exploit the inherent redundancy in video data. Video coding aims to compress video data into a form that uses a lower bitrate while avoiding or minimizing a degradation in video quality. A variety of video codec standards have been developed. For example, High-Efficiency Video Coding (HEVC / H.265) is a video compression standard designed as part of the MPEG-H project. The HEVC / H.265 standard was released by ITU-T and ISO / IEC in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4), respectively. Versatile Video Coding (VVC / H.266) is a video compression standard that is a successor to HEVC. The VVC / H.266 standard was released by ITU-T and ISO / IEC in 2020 (Version 1) and 2022 (Version 2), respectively. AOMedia Video 1 (AV1) is an open video coding format designed to replace HEVC. A verified version 1.0.0 with Errata 1 was released on January 8, 2019. Summary of the Invention
[0005] The present disclosure describes a set of methods for video (image) compression, and more particularly, relates to signaling loop filter parameters when encoding multiple views of a scene. In some embodiments, rather than encoding each view independently and transmitting a bitstream from each view (simulcast coding), a disparity compensation prediction method is implemented such that pictures from other views are included in the reference picture list at the same time instance. Disparity compensation prediction can improve coding efficiency by reducing the statistical redundancy that exists between different views. In certain instances, the methods disclosed herein can achieve a bitrate savings of approximately 70% compared to simulcast coding.
[0006] According to some embodiments, a method of video decoding includes: (i) receiving a multi-view video bitstream including a plurality of pictures, the plurality of pictures including a first picture corresponding to a first view and a second picture corresponding to a second view; (ii) based on a first indicator in the multi-view video bitstream, determining whether the loop filter parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are signaled jointly or separately; and (iii) in accordance with the first indicator indicating that the loop filter parameters are signaled jointly, performing a first loop filter process on the first picture and a second loop filter process on the second picture using a shared set of loop filter parameters (e.g., a set of one or more parameters).
[0007] According to some embodiments, a method of video encoding includes: (i) receiving video data including a plurality of pictures, the plurality of pictures including a first picture corresponding to a first view and a second picture corresponding to a second view; (ii) determining whether loop filter parameters corresponding to the first picture of the first view and the second picture of the second view are to be signaled jointly or separately; and (iii) writing a first indicator into a multi-view video bitstream according to determining that the loop filter parameters are to be signaled jointly, the first indicator being used to indicate that the loop filter parameters are signaled jointly for the first picture and the second picture.
[0008] According to some embodiments, a method of bitstream conversion includes: (i) obtaining a source video sequence corresponding to a set of views; and (ii) performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, where the multi-view video bitstream includes (a) a first plurality of encoded pictures corresponding to a first view, the first plurality of encoded pictures including a first picture; (b) a second plurality of encoded pictures corresponding to a second view, the second plurality of encoded pictures including a second picture; and (c) a first indicator, the first indicator being used to indicate whether loop filter parameters are signaled jointly for the first picture and the second picture.
[0009] According to some embodiments, a computing system such as a streaming system, a server system, a personal computer system, or other electronic devices is provided. The computing system includes a control circuit and a memory storing one or more instruction sets. The one or more instruction sets include instructions for performing any of the methods described herein. In some embodiments, the computing system includes an encoder component and a decoder component (e.g., a transcoder).
[0010] According to some embodiments, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores one or more instruction sets for execution by a computing system. The one or more instruction sets include instructions for performing any of the methods described herein.
[0011] Therefore, methods, devices, and systems for video encoding and video decoding are disclosed. Such methods, devices, and systems may supplement or replace conventional methods, devices, and systems for encoding / decoding video. The features and advantages described in this specification are not necessarily an exhaustive list. Specifically, according to the drawings, the specification, and the claims provided by the present disclosure, some additional features and advantages are obvious to those of ordinary skill in the art. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes and is not necessarily for detailed description or limitation of the subject matter described herein. Description of the Drawings
[0012] For a more detailed understanding of the present disclosure, a more specific description may be made with reference to the features of various embodiments, some of which are illustrated in the accompanying drawings. However, the accompanying drawings only show the relevant features of the present disclosure and thus are not necessarily considered restrictive, as those skilled in the art will understand after reading the present disclosure that the description may include other effective features.
[0013] Figure 1 is a block diagram showing an example communication system according to some embodiments.
[0014] Figure 2A is a block diagram showing example elements of an encoder component according to some embodiments.
[0015] Figure 2B is a block diagram showing example elements of a decoder component according to some embodiments.
[0016] Figure 3 is a block diagram showing an example server system according to some embodiments.
[0017] Figure 4 shows an example MVV according to some embodiments.
[0018] Figures 5 to 7 shows example operations in an MVV according to some embodiments.
[0019] Figures 8A to 8C shows example prediction blocks, residual blocks, and reconstruction blocks according to some embodiments.
[0020] Figure 9 shows an example in-loop filtering stage according to some embodiments.
[0021] Figure 10A shows an example video decoding process according to some embodiments.
[0022] Figure 10B shows an example video encoding process according to some embodiments.
[0023] By convention, the various features shown in the accompanying drawings are not necessarily drawn to scale, and throughout the specification and the drawings, the same reference numerals may be used to denote the same features. Detailed Description
[0024] The present disclosure describes video / image compression techniques, including loop filtering techniques for MVV encoding. The disclosed techniques include signaling loop filtering parameters jointly for pictures of different views belonging to an MVV. For example, in the disclosed techniques, based on whether the loop filtering parameters of a first picture and a second picture are signaled jointly or separately, it is determined whether to perform a first loop filtering process on a first picture corresponding to a first view and a second loop filtering process on a second picture corresponding to a second view using a shared set of loop filtering parameters. Inter-view redundancy is reduced through joint signaling modes and / or parameters, thereby improving the encoding / decoding efficiency. Example systems and devices
[0025] Figure 1 FIG. 6 is a block diagram showing a communication system 100 according to some embodiments. The communication system 100 includes a source device 102 and a plurality of electronic devices 120 (e.g., electronic devices 120-1 to electronic devices 120-m), which are communicatively coupled to each other via one or more networks. In some embodiments, the communication system 100 is a streaming system, e.g., for supporting applications for video, such as video conferencing applications, digital television applications, and media storage and / or distribution applications.
[0026] The source device 102 includes a video source 104 (e.g., a camera assembly or a media memory) and an encoder assembly 106. In some embodiments, the video source 104 is a digital camera (e.g., configured to create an uncompressed video sample stream). The encoder assembly 106 generates one or more encoded video bitstreams from the video stream. The video stream from the video source 104 may be of high data volume compared to the encoded video bitstreams 108 generated by the encoder assembly 106. Since the encoded video bitstreams 108 have a lower data volume (less data) compared to the video stream from the video source, the encoded video bitstreams 108 require less bandwidth for transmission and less storage space for storage compared to the video stream from the video source 104. In some embodiments, the source device 102 does not include the encoder assembly 106 (e.g., configured to transmit uncompressed video to one (or more) networks 110).
[0027] One or more networks 110 represent any number of networks that transfer information between the source device 102, the server system 112, and / or the electronic devices 120, including, for example, wireline or wired networks and / or wireless communication networks. One or more networks 110 may exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0028] One or more networks 110 include a server system 112 (e.g., a distributed computing system / cloud computing system). In some embodiments, the server system 112 is or includes a streaming server (e.g., configured to store and / or distribute video content, such as an encoded video stream from the source device 102). The server system 112 includes an encoder component 114 (e.g., configured to encode and / or decode video data). In some embodiments, the encoder component 114 includes an encoder component and / or a decoder component. In various embodiments, the encoder component 114 is instantiated as hardware, software, or a combination thereof. In some embodiments, the encoder component 114 is configured to decode the encoded video bitstream 108 and re-encode the video data using a different encoding standard and / or method to generate encoded video data 116. In some embodiments, the server system 112 is configured to generate multiple video formats and / or encodings based on the encoded video bitstream 108. In some embodiments, the server system 112 functions as a Media-Aware Network Element (MANE). For example, the server system 112 can be configured to trim the encoded video bitstream 108 to customize potentially different bitstreams for one or more electronic devices 120. In some embodiments, the MANE is provided separately from the server system 112.
[0029] The electronic device 120-1 includes a decoder component 122 and a display 124. In some embodiments, the decoder component 122 is configured to decode the encoded video data 116 to generate an output video stream that can be displayed on a display or other type of display device. In some embodiments, one or more of the electronic devices 120 do not include a display component (e.g., communicatively coupled to an external display device and / or including a media memory). In some embodiments, the electronic device 120 is a streaming client. In some embodiments, the electronic device 120 is configured to access the server system 112 to obtain the encoded video data 116.
[0030] The source device and / or the multiple electronic devices 120 are sometimes referred to as "terminal devices" or "user devices". In some embodiments, the source device 102 and / or one or more of the electronic devices 120 are examples of: a server system, a personal computer, a portable device (e.g., a smartphone, a tablet, or a laptop), a wearable device, a video conferencing device, and / or other types of electronic devices.
[0031] In an example operation of communication system 100, source device 102 transmits an encoded video bitstream 108 to server system 112. For example, source device 102 may encode a picture stream captured by the source device. Server system 112 receives encoded video bitstream 108 and may use encoder component 114 to decode and / or encode encoded video bitstream 108. For example, server system 112 may apply an encoding that is more suitable for network transmission and / or storage to the video data. Server system 112 may transmit encoded video data 116 (e.g., one or more encoded video bitstreams) to one or more electronic devices 120. Each electronic device 120 may decode the encoded video data 116 and optionally display the video pictures.
[0032] Figure 2A is a block diagram showing example elements of encoder component 106 according to some embodiments. Encoder component 106 receives video data (e.g., a source video sequence) from video source 104. In some embodiments, the encoder component includes a receiver (e.g., transceiver) component configured to receive the source video sequence. In some embodiments, encoder component 106 receives a video sequence from a remote video source (e.g., a video source that is a component of a device different from encoder component 106). Video source 104 may provide the source video sequence in the form of a digital video sample stream, which may have any suitable bit depth (e.g., 8 bits, 10 bits, or 12 bits), any color space (e.g., BT.601 Y CrCb or RGB), and any suitable sampling structure (e.g., Y CrCb 4:2:0 or Y CrCb 4:4:4). In some embodiments, video source 104 is a storage device for storing previously captured / prepared video. In some embodiments, video source 104 is a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. Those of ordinary skill in the art can readily understand the relationship between pixels and samples.
[0033] The encoder component 106 is configured to encode and / or compress pictures of a source video sequence into an encoded video sequence 216 in real time or under other time constraints required by an application. In some embodiments, the encoder component 106 is configured to perform a conversion between a source video sequence and a bitstream of visual media data (e.g., a video bitstream). Enforcing an appropriate encoding speed is a function of the controller 204. In some embodiments, the controller 204 controls and is functionally coupled to other functional units as described below. Parameters set by the controller 204 may include rate control related parameters (e.g., picture skipping, quantizer, and / or λ value of rate-distortion optimization techniques), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those of ordinary skill in the art can readily identify other functions of the controller 204 as they may relate to the encoder component 106 optimized for a particular system design.
[0034] In some embodiments, the encoder component 106 is configured to operate in an encoding loop. In a simple example, the encoding loop includes a source encoder 202 (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and one (or more) reference pictures) and a (local) decoder 210. The decoder 210 (when the compression between the symbols and the encoded video bitstream is lossless) reconstructs the symbols in a manner similar to a (remote) decoder to create sample data. The reconstructed sample stream (sample data) is input into the reference picture memory 208. Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory 208 is also bit-exact between the local encoder and the remote encoder. In this way, the prediction part of the encoder interprets the reference picture samples with the same values as the samples that the decoder will interpret during decoding when using prediction.
[0035] The operation of the decoder 210 may be the same as that of a remote decoder such as the decoder component 122 described in detail below in connection with Figure 2B However, briefly referring to Figure 2B , when symbols are available and the entropy encoder 214 and the parser 254 can encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the decoder component 122, including the buffer memory 252 and the parser 254, may not be fully implemented in the local decoder 210.
[0036] Except for parsing / entropy decoding, the decoder techniques described herein may exist in a corresponding encoder in a substantially identical functional form. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder techniques can be simplified as encoder techniques can be inverse to decoder techniques.
[0037] As part of the operation of the source encoder 202, the source encoder 202 may perform motion-compensated predictive coding, which predictively encodes an input frame by referring to one or more previously encoded frames designated as reference frames in a video sequence. In this way, the encoding engine 212 encodes the difference between a pixel block of the input frame and a pixel block of one (or more) reference frames, which one (or more) reference frames may be selected as one (or more) prediction references for the input frame. The controller 204 may manage the encoding operations of the source encoder 202, including, for example, setting parameters and subgroup parameters for encoding video data.
[0038] The decoder 210 decodes the encoded video data of a frame that may be designated as a reference frame based on the symbols created by the source encoder 202. The operation of the encoding engine 212 may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 2A not shown), the reconstructed video sequence may be a copy of the source video sequence with some errors. The decoder 210 replicates the decoding process that may be performed by a remote video decoder on a reference frame and may cause the reconstructed reference frame to be stored in the reference picture memory 208. In this way, the encoder component 106 locally stores a copy of the reconstructed reference frame, which has the same content (in the absence of transmission errors) as the reconstructed reference frame that will be obtained by the remote video decoder.
[0039] The predictor 206 may perform a prediction search on the encoding engine 212. That is, for a new frame to be encoded, the predictor 206 may search the reference picture memory 208 for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new picture. The predictor 206 may operate on sample blocks on a per-pixel-block basis to find an appropriate prediction reference. As determined according to the search results obtained by the predictor 206, the input picture may have prediction references obtained from a plurality of reference pictures stored in the reference picture memory 208.
[0040] The outputs of all the above functional units may be entropy encoded in the entropy encoder 214. The entropy encoder 214 performs lossless compression on the symbols generated by the various functional units according to techniques known to those of ordinary skill in the art (such as Huffman coding, variable length coding, and / or arithmetic coding), thereby transforming these symbols into an encoded video sequence.
[0041] In some embodiments, the output of the entropy encoder 214 is coupled to a transmitter. The transmitter may be configured to buffer the encoded video sequence(s) created by the entropy encoder 214 in preparation for transmission over a communication channel 218, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter may be configured to combine the encoded video data from the source encoder 202 with other data to be transmitted, such other data being, for example, encoded audio data and / or auxiliary data streams (sources not shown). In some embodiments, the transmitter may transmit additional data when transmitting the encoded video. The source encoder 202 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / signal-to-noise ratio (SNR) enhancement layers, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual usability information (VUI) parameter set segments, etc.
[0042] The controller 204 may manage the operation of the encoder components 106. During encoding, the controller 204 may assign a certain encoded picture type to each encoded picture, which may affect the encoding technique applied to the corresponding picture. For example, a picture may be assigned as an intra picture (I picture), a predictive picture (P picture), or a bi-predictive picture (B picture). An intra picture may be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh (IDR) pictures. Persons of ordinary skill in the art are aware of the variants of I pictures and their corresponding applications and characteristics, and thus will not be repeated here. A predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict the sample values of each block. A bi-predictive picture may be encoded and decoded using intra prediction or inter prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0043] The source pictures can typically be spatially subdivided into multiple sample blocks (e.g., each sample block includes 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which is determined by the coding assignment for the corresponding picture applied to the block. For example, the blocks of an I picture can be non-predictively encoded, or these blocks can be predictively encoded with reference to the already encoded blocks of the same picture (spatial prediction or intra-frame prediction). The pixel blocks of a P picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. The blocks of a B picture can be non-predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.
[0044] The captured video can be multiple source pictures (video pictures) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In one example, the particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.
[0045] The encoder component 106 can perform encoding operations according to any predetermined video coding technique or standard such as those described herein. In the operation of the encoder component 106, the encoder component 106 can perform various compression operations, including predictive coding operations that utilize the temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard being used.
[0046] Figure 2B is a block diagram showing example elements of a decoder component 122 according to some embodiments. Figure 2B The decoder component 122 in is coupled to the channel 218 and the display 124. In some embodiments, the decoder component 122 includes a transmitter that is coupled to the loop filter 256 and is configured to transmit data to the display 124 (e.g., via a wired or wireless connection).
[0047] In some embodiments, decoder component 122 includes a receiver coupled to channel 218 and configured to receive data from channel 218 (e.g., via a wired or wireless connection). The receiver may be configured to receive one or more encoded video sequences (e.g., a video bitstream) to be decoded by decoder component 122. In some embodiments, the decoding of each encoded video sequence is independent of other encoded video sequences. Each encoded video sequence may be received from channel 218, which may be a hardware / software link to a storage device storing the encoded video data. The receiver may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver may separate the encoded video sequences from the other data. In some embodiments, the receiver receives additional (redundant) data along with the encoded video. This additional data may be included as part of one (or more) of the encoded video sequences. The additional data may be used by decoder component 122 to decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0048] According to some embodiments, decoder component 122 includes a buffer memory 252, a parser 254 (sometimes also referred to as an entropy decoder), a scaler / inverse transform unit 258, an intra picture prediction unit 262, a motion compensation prediction unit 260, an aggregator 268, a loop filter unit 256, a reference picture memory 266, and a current picture memory 264. In some embodiments, decoder component 122 is implemented as an integrated circuit, a series of integrated circuits, and / or other electronic circuits. Decoder component 122 may be implemented at least partially in software.
[0049] Buffer memory 252 is coupled between channel 218 and parser 254 (e.g., to prevent network jitter). In some embodiments, buffer memory 252 is separate from decoder component 122. In some embodiments, a separate buffer memory is provided between the output of channel 218 and decoder component 122. In some embodiments, in addition to buffer memory 252 within decoder component 122 (e.g., which is configured to handle playout timing), a separate buffer memory is provided external to decoder component 122 (e.g., to prevent network jitter). When receiving data from a store-and-forward device with sufficient bandwidth and controllability or from an isochronous network, buffer 252 may not be needed, or a small buffer may be used. For use on a best-effort packet network such as the Internet, buffer memory 252 may be necessary, may be relatively large and / or have an adaptive size, and may be implemented at least partially in an operating system or a similar element external to decoder component 122.
[0050] The parser 254 is configured to reconstruct symbols 270 from the encoded video sequence. These symbols may include, for example, information for managing the operation of the decoder component 122, and / or information for controlling a display device (e.g., the display screen 124). The control information for the display device may be in the form of, for example, a Supplementary Enhancement Information (SEI) message or a fragment of a Video Usability Information (VUI) parameter set (not depicted). The parser 254 parses (entropy decodes) the encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technique or standard and may follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser 254 may extract subgroup parameter sets for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and the like. The parser 254 may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0051] Depending on the type of the encoded video picture or a portion of the encoded video picture (such as: inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol 270 may involve multiple different units. Which units are involved and the manner of involvement may be controlled by the subgroup control information parsed by the parser 254 from the encoded video sequence. For clarity, such subgroup control information flows between the parser 254 and the multiple units below are not depicted.
[0052] The decoder component 122 may be conceptually subdivided into several functional units, and in some implementations, these units interact closely with each other and may be at least partially integrated with each other. However, for clarity, the conceptual subdivision of the functional units is retained herein. The scaler / inverse transform unit 258 receives the quantized transform coefficients and control information (such as which transform mode to use, block size, quantization factor, and / or quantization scaling matrix) as symbols 270 from the parser 254. The scaler / inverse transform unit 258 may output a block including sample values, which may be input into the aggregator 268.
[0053] In some cases, the output samples of the scaler / inverse transform unit 258 belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra picture prediction unit 262. The intra picture prediction unit 262 may generate a block having the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current (partially reconstructed) picture in the current picture memory 264. The aggregator 268 may add the prediction information generated by the intra picture prediction unit 262 to the output sample information provided by the scaler / inverse transform unit 258 based on each sample.
[0054] In other cases, the output samples of the scaler / inverse transform unit 258 belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit 260 may access the reference picture memory 266 to extract samples for prediction. After motion compensating the extracted samples according to the symbol 270 belonging to the block, these samples may be added by the aggregator 268 to the output of the scaler / inverse transform unit 258 (referred to as residual samples or a residual signal in this case) to generate output sample information. The address in the reference picture memory 266 from which the motion compensation prediction unit 260 extracts the prediction samples may be controlled by a motion vector. The motion vector may be in the form of the symbol 270 for use by the motion compensation prediction unit 260, and the symbol 270 may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory 266, a motion vector prediction mechanism, etc. when using sub-sampled accurate motion vectors.
[0055] The output samples of the aggregator 268 may be used in the loop filter unit 256 for various loop filtering techniques. The video compression technique may include an in-loop filter technique that is controlled by a parameter included in the encoded video bitstream, and the parameter may be used for the loop filter unit 256 as the symbol 270 from the parser 254, but may also be in response to meta-information obtained during decoding of a previously (in decoding order) portion of the encoded picture or encoded video sequence, and in response to previously reconstructed and loop-filtered sample values. The output of the loop filter unit 256 may be a sample stream that may be output to a display device such as the display 124, and stored in the reference picture memory 266 for future inter-picture prediction.
[0056] Once reconstructed, certain encoded pictures may be used as reference pictures for future prediction. Once an encoded picture has been reconstructed and the encoded picture has been identified (by, for example, the parser 254) as a reference picture, the current reference picture may become part of the reference picture memory 266, and a new current picture memory may be reallocated before starting to reconstruct subsequent encoded pictures.
[0057] The decoder component 122 may perform decoding operations according to predetermined video compression techniques that may be recorded in a standard such as any standard described herein. The encoded video sequence may conform to the syntax of the video compression technique or standard used in the sense that the encoded video sequence follows the video compression technique or standard specified in the video compression technique document or standard (in particular, the profile document thereof). Additionally, to conform to some video compression techniques or standards, the complexity of the encoded video sequence may be within the range defined by the level of the video compression technique or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, millions of samples per second), maximum reference picture size, etc. In certain cases, the level-based restrictions can be further restricted by assuming a Hypothetical Reference Decoder (HRD) specification and HRD buffer management metadata signaled in the encoded video sequence.
[0058] Figure 3 is a block diagram showing a server system 112 according to some embodiments. The server system 112 includes a control circuit 302, one or more network interfaces 304, a memory 314, a user interface 306, and one or more communication buses 312 for interconnecting these components. In some embodiments, the control circuit 302 includes one or more processors (e.g., CPU, GPU, and / or DPU). In some embodiments, the control circuit includes one (or more) field programmable gate arrays (FPGAs), hardware accelerators, and / or one (or more) integrated circuits (e.g., application specific integrated circuits).
[0059] One (or more) network interfaces 304 may be configured to connect to one or more communication networks (e.g., wireless networks, wired networks, and / or optical networks). The communication network may be local, wide area, metropolitan, vehicular and industrial, real-time, fault-tolerant, etc. Examples of communication networks include local area networks (such as Ethernet), wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Such communication may be unidirectional, receive-only (e.g., broadcast television), send-only unidirectional (e.g., CANBus to certain CANbus devices), or bidirectional (e.g., other computer systems using local area networks or wide area digital networks). Such communication may include communication with one or more cloud computing networks.
[0060] The user interface 306 includes one or more output devices 308 and / or one or more input devices 310. One or more input devices 310 may include one or more of the following: keyboard, mouse, touchpad, touch screen, data glove, joystick, microphone, scanner, camera, etc. One or more output devices 308 may include one or more of the following: audio output devices (such as speakers), visual output devices (such as displays or monitors), etc.
[0061] The memory 314 may include high-speed random access memory (such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDRRAM), and / or other random access solid-state memory devices) and / or non-volatile memory (such as one or more disk storage devices, optical disc storage devices, flash memory devices, and / or other non-volatile solid-state storage devices). The memory 314 optionally includes one or more storage devices remote from the control circuit 302. The memory 314 (or alternatively, one or more non-volatile solid-state storage devices within the memory 314) includes a non-transitory computer-readable storage medium. In some embodiments, the memory 314 or the non-transitory computer-readable storage medium of the memory 314 stores the following programs, modules, instructions, and data structures, or subsets or supersets thereof: ● An operating system 316, which includes procedures for handling various basic system services and for performing hardware-related tasks. · A network communication module 318, which is used to connect the server system 112 to other computing devices via one or more network interfaces 304 (such as via a wired connection and / or a wireless connection). · A decoding module 320, which is used to perform various functions related to encoding and / or decoding data (such as video data). In some embodiments, the decoding module 320 is an instance of the encoder component 114. The decoding module 320 includes, but is not limited to, one or more of the following: ○ A decoding module 322, which is used to perform various functions related to decoding encoded data, such as those functions previously described with respect to the decoder component 122. ○ An encoding module 340, which is used to perform various functions related to encoding data, such as those functions previously described with respect to the encoder component 106. ● Picture memory 352 for storing pictures and picture data for use by, for example, decoding module 320. In some embodiments, picture memory 352 includes one or more of the following: reference picture memory 208, buffer memory 252, current picture memory 264, and reference picture memory 266.
[0062] In some embodiments, decoding module 322 includes parsing module 324 (e.g., configured to perform the various functions previously described with respect to parser 254), transformation module 326 (e.g., configured to perform the various functions previously described with respect to scaler / inverse transform unit 258), prediction module 328 (e.g., configured to perform the various functions previously described with respect to motion compensation prediction unit 260 and / or intra picture prediction unit 262), and filter module 330 (e.g., configured to perform the various functions previously described with respect to loop filter 256).
[0063] In some embodiments, encoding module 340 includes code module 342 (e.g., configured to perform the various functions previously described with respect to source encoder 202 and / or encoding engine 212) and prediction module 344 (e.g., configured to perform the various functions previously described with respect to predictor 206). In some embodiments, decoding module 322 and / or encoding module 340 include Figure 3 a subset of the modules shown. For example, both decoding module 322 and encoding module 340 use a shared prediction module.
[0064] Each of the above-identified modules stored in memory 314 corresponds to an instruction set for performing the functions described herein. The above-identified modules (e.g., instruction sets) need not be implemented as separate software programs, procedures, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various embodiments. For example, decoding module 320 optionally does not include separate decoding and encoding modules, but uses the same set of modules to perform both sets of functions. In some embodiments, memory 314 stores a subset of the above-identified modules and data structures. In some embodiments, memory 314 stores additional modules and data structures not described above.
[0065] Although Figure 3 server system 112 is shown in accordance with some embodiments, Figure 3 it is more of a functional description of the various features that may exist in one or more server systems than a structural diagram of the embodiments described herein. In practice, items shown separately may be combined, and some items may be separated. For example, Figure 3Some of the items shown separately in the Chinese can be implemented on a single server, and a single item can be implemented by one or more servers. The actual number of servers used to implement the server system 112 and how the features are distributed among them will vary depending on the implementation, and optionally, partly depend on the data traffic processed by the server system during peak usage periods and average usage periods. Example encoding techniques
[0066] The encoding processes and techniques described below can be performed on the above-mentioned devices and systems (e.g., the source device 102, the server system 112, and / or the electronic device 120). Hereinafter, a block may refer to the largest coding block, coding block / unit, prediction block, transform block, or a predefined fixed block size. The term "block" can also be used to refer to a filtering unit, which is a block unit that performs a loop filtering method.
[0067] In MMV, different views can have strong correlations, and reducing the statistical redundancy existing between different views helps to improve the encoding efficiency. As described in detail below, some MMV techniques utilize the block partitioning information of a first block in a first picture of a first view to perform block partitioning and encoding of a second block in a second picture of a second view. The first block and the second block can be located at the same coordinates in the first view and the second view respectively, and the first picture and the second picture can be associated with the same display time.
[0068] Figure 4 An example MVV 500 having two views (e.g., view 0 and view Figure 1 ) is shown according to some embodiments. Each view can be associated with a different viewport or camera. For applications such as stereoscopic video viewing, videos of more than one view can be encoded. In some embodiments, the MVV 500 corresponds to a three-dimensional (3D) scene captured by two or more cameras. In certain cases, optional processing such as view correction and color correction is performed on the transmitter side. After encoding the MVV sequence, the bitstream is transmitted to the receiver side, where the views are decoded and presented on a suitable 3D display. In some embodiments, the MVV includes more than two views.
[0069] Figure 5 An example prediction structure 600 of an MVV according to some embodiments is shown. The structure 600 uses temporal reference pictures (represented by horizontal arrows and curved arrows) and inter-view reference pictures (represented by vertical arrows) for motion compensation and disparity compensation prediction. Figure 6Two picture sequences corresponding to the left view and the right view are shown, where due to the strong correlation between the two views, the pictures in the left view are used to predict the pictures in the right view, thereby improving the coding efficiency. Picture 602 in the right view (POC 0) is a P picture / frame encoded (e.g., predicted) using picture 604 in the left view as a reference picture. Picture 604 is an I picture. The next frame encoded in the right view is picture 606 (POC 8), corresponding to the last P frame in the sequence. Then, picture 602 and picture 606 are used to derive picture 610 (POC 4) in the right view. Picture 610 is a bi-directional B picture / frame. Then, POC 0, POC 8, and POC 4 obtained in the right view are used to derive POC 2 and POC 6 in the right view. In some embodiments, the pictures in the right view are multi-layered. For example, the odd POCs (denoted by the lowercase letter "b") in the right view are in different layers from the even POCs.
[0070] Figure 6 Video data 700A and 700B are shown, each of which includes a plurality of views (e.g., view 0 to view Figure 5 ) that are spatially stitched together to form a two-dimensional image. Although Figure 7 six views in the video data are depicted, it should be understood that any number of views can be stitched together. There are multiple ways to stitch the views spatially. For example, for six views, one, two, or three views can be stitched in each row, which may result in a 1×6 stitch, a 2×3 stitch, and a 3×2 stitch, respectively. In some embodiments, the stitching is designed such that the resulting super-sized picture has a desired picture size. For example, the super-sized picture can be close to square or a rectangle with an aspect ratio of 4:3 or 16:9.
[0071] For a P slice or a B slice in the super-sized picture, the motion vectors from the previously encoded views in the same picture are highly correlated with the motion vectors in the current view being encoded. Therefore, the motion vectors from the previously encoded views in the same picture are very suitable for use in motion vector prediction or as a starting point for motion estimation of the current view. Additionally, in some embodiments, a motion vector predictor (MVP) can be calculated through a perspective transformation between two views (e.g., from a reference view to the current view). In some embodiments, an MVP candidate is derived for the current block 702 in the current view (e.g., view 2) of the current picture (Pcur). In some embodiments, block 702 is mapped to block 704 in the reference view (Vref) of Pcur. If the motion vector of block 704 is (Mx, My) and its reference block 706 is in the same view (Vref) of the reference picture (Pref), then X2 = X3 + Mx , Y2 = Y3 + My, where (X2, Y2) are the coordinates of reference block 706, (X3, Y3) are the coordinates of block 708, and block 708 is a collocated block of block 704. Both block 708 and block 704 can be located in the reference view Vref (e.g., view 0), and block 708 can reside in Pref.
[0072] In Figure 6 , block 710 is a reference block of block 702, block 712 is a collocated block of block 702, and block 712 can reside in Vcur of Pref. In some embodiments, the coordinates of block 710 are derived based on the coordinates of block 712 and the calculated (e.g., derived) MVP candidate for block 702.
[0073] In some embodiments, assuming that the samples in a block share the same disparity, a disparity vector DV (Dx, Dy) is used to find the collocated block of the current block in the same picture in the reference view. In some embodiments, a position offset is established between the current view and its reference view (e.g., the offset can be twice the view width in the x direction and 0 in the y direction). The disparity vector can be added to the view offset to find the collocated block of the current block in the reference view. Since the collocated block indicated by the disparity vector in the reference view (e.g., view 0) may have been encoded, its motion vector (if any) can point to the reference block in view 0 of the reference picture. In some embodiments, the reference block is used as the reference block for the current block (in view 2) or as the starting point for motion estimation.
[0074] Figure 7 Video data 800 with multiple views is shown according to some embodiments. The video data 800 includes multiple views (e.g., views 0 to view Figure 5 ) that are spatially stitched together to form a two-dimensional image. A block vector (BV) can point to a reference block in a previously encoded view. In some embodiments, a perspective transformation is estimated between two views (e.g., views 0 and view Figure 5 ) before block matching. The perspective transformation can establish a bijective mapping between the two views. The perspective transformation can be applied to the reference view (e.g., view 0), which maps the coordinates from the reference view to the current view being encoded (e.g., view Figure 5 ). The "transformed" reference view can be used as the reference for block matching.
[0075] For a block in the current view (e.g., block 808A), the co-located block (e.g., block 804) of the block can be found at the same location (e.g., the same coordinates) in the "transformed" reference view, and the block matching can start with the co-located block (e.g., with block 802B) as the starting point of the block matching. Since the perspective transformation is very close to the transition between the two views, the block matching can be limited to a small neighborhood of the co-located block in the "transformed" reference view. In some embodiments, the current block and its co-located block have the same position offset relative to the upper left position of their respective views. By analyzing the reference view and the current view, the perspective transformation can be estimated. The perspective transformation can also be derived using techniques from computational photography (e.g., key point detection and matching), or calculated directly based on camera parameters and depth map data.
[0076] In some embodiments, assuming that the samples in the block share the same disparity, the disparity vector DV (Dx, Dy) is used to indicate the disparity between the co-located block in the reference view and the reference block in the reference view. It should be noted that for each view pair, the block vector (BV) can be different for blocks at different positions in the view. The block vector pointing from the current block to the reference block in its reference view consists of two parts: the view position offset plus the disparity vector.
[0077] Figure 8A Shows the calculation of a predicted block according to some embodiments. In Figure 8A the example, intra-frame prediction is performed on the current block 902 to generate a predicted block 904. In some embodiments, inter-frame prediction is performed to generate a predicted block. The current block 902 includes a set of samples (e.g., a pixel block), and the predicted block 904 includes a corresponding set of predictions for the set of samples. Figure 8B Shows the calculation of a residual block according to some embodiments. As Figure 8B shown, the predicted block 904 is subtracted from the current block 902 to generate a residual block 906 including a set of residuals. For example, the corresponding difference between each sample and the corresponding prediction is calculated. Figure 8C Shows the calculation of a reconstructed block according to some embodiments. As Figure 8C shown, the residual block 906 undergoes one or more transforms and quantization to generate a set of residual coefficients. The set of residual coefficients can be transmitted from the encoder component to the decoder component. The set of residual coefficients undergoes inverse quantization and inverse transform to generate a reconstructed residual block 908. The reconstructed residual block 908 is combined with the predicted block 904 (e.g., adding the reconstructed residuals of the reconstructed residual block 908 to the prediction of the predicted block 904) to generate a reconstructed block 910 corresponding to the current block 902.
[0078] Figure 9 Shows an example in-loop filtering stage according to some embodiments. In Figure 9Among them, the in-loop filtering stage applied to the decoded frame 922 includes a deblocking filter 924, a constrained direction enhancement filter (CDEF) 926, and an in-loop restoration filter 930. In some embodiments, the filtered output frame is used as a reference frame for subsequent frames (e.g., stored in the reference frame buffer 928). The in-loop filtering method may include any filtering process applied to the reconstructed samples (e.g., after adding the residual to the prediction), including Wiener loop filtering, cross-component filtering, and CDEF. In some cases, the reconstructed samples filtered via in-loop filtering can be used as reference samples for performing prediction within a picture. In some cases, the reconstructed samples filtered via in-loop filtering cannot be used as reference samples.
[0079] The cross-component filtering method can use the co-located reconstructed samples and adjacent reconstructed samples from the first color component as inputs to perform filtering on the current reconstructed samples of the second color component. The cross-component offset filtering method can use the co-located reconstructed samples and their adjacent reconstructed samples from the first color component as inputs to derive an offset value, which is added to the current samples of the second color component to adjust their reconstructed values. The first color component may refer to the luminance color component, and the second color component may refer to the chrominance color component. The first color component and the second color component may be the same color component (e.g., the luminance component).
[0080] The deblocking filter 924 can be applied across the transform block boundaries to remove the block artifacts caused by quantization errors. In some embodiments, the filter length is determined based on the minimum transform block size on both sides. In some embodiments, the deblocking filter 924 uses a finite impulse response (FIR) filter (e.g., a low-pass filter). Edge detection can be used to disable the deblocking filter at transitions containing high-variance signals (e.g., to avoid blurring the actual edges in the original image). In this way, the deblocking filtering method can be applied to the reconstructed samples located near the block boundaries. The block boundaries may include transform block boundaries, motion compensation block boundaries, coding block boundaries, and / or fixed block size boundaries.
[0081] The CDEF 926 applies a non-linear de-ringing filter along a specific (e.g., tilted) direction. The CDEF 926 can operate on the output of the de-blocking filter 924. The CDEF 926 can operate in 8x8 units. In some embodiments, eight preset directions are defined by rotating and reflecting a template in a preset direction. The decoder can use the reconstructed pixels to select a dominant direction index. The main filter can be applied along the selected direction, and the secondary filter can be applied along an offset direction (e.g., 45° off the main direction). In some embodiments, up to eight sets of filter parameters are signaled (e.g., in the frame header). The sets of filter parameters can include main filter strength indices and secondary filter strength indices for luminance and chrominance components. The CDEF 926 can apply filtering to the reconstructed samples by identifying the direction of each block and then performing highly controlled adaptive filtering on the filtering strength along and across that direction.
[0082] In some embodiments, the loop restoration filter 930 is applied to the reconstructed pixels after any previous in-loop filtering stage (e.g., the de-blocking filter 924 and / or the CDEF 926). The loop restoration filter 930 can be applied to a loop restoration unit (LRU), such as a 64x64, 128x128, and / or 256x256 pixel block. Bypass filtering, a Wiener filter (e.g., the Wiener loop filtering method), and / or a self-guiding filter can be independently applied to each LRU. The Wiener loop filtering method can use a linear weighted sum of the current reconstructed sample and multiple spatially adjacent reconstructed samples as an input to derive a modified value of the current reconstructed sample as an output.
[0083] Figure 10A is a flowchart showing a method 1000 for decoding video according to some embodiments. The method 1000 can be executed at a computing system (e.g., the server system 112, the source device 102, or the electronic device 120) having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, the method 1000 is executed by executing instructions stored in the memory (e.g., the memory 314) of the computing system.
[0084] The system receives (1002) a multi-view video bitstream that includes a plurality of pictures, the plurality of pictures including a first picture corresponding to a first view and a second picture corresponding to a second view. The system determines (1004) whether the loop filter parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are signaled jointly or separately based on a first indicator in the multi-view video bitstream. According to the first indicator indicating that the loop filter parameters are signaled jointly, the system performs a first loop filtering process on the first picture and a second loop filtering process on the second picture using a shared set of loop filter parameters (1006). For example, for a plurality of pictures from different views, loop filter parameters associated with one or more loop filtering methods are signaled jointly. Example loop filtering methods include cross-component filtering, cross-component offset filtering, Wiener loop filtering, deblocking filtering, and CDEF methods. In some embodiments, high-level syntax (HLS) is signaled to indicate whether the loop filter parameters are signaled jointly for a plurality of pictures from different views or separately for a plurality of pictures from different views. In some embodiments, high-level syntax is signaled jointly for a plurality of views, the high-level syntax being used to signal the enabling of one (or more) loop filtering methods.
[0085] In some embodiments, a picture-level on / off flag for one (or more) loop filtering methods is signaled jointly for a plurality of views. In some embodiments, depending on the picture-level on / off flag of a loop filtering method of one view, the picture-level on / off flag of a loop filtering method in another view is signaled. In one example, for a first picture in a first view, three flags for different color components may be signaled to specify whether the first loop filtering method is applied to the three color components separately. Then, for the second view, one flag is signaled to indicate whether to inherit the values of the three flags for the second picture in the second view. In another example, for a first picture in a first view, a plurality of flags are signaled separately to specify whether each of a plurality of loop filtering methods is enabled. Then, for a second picture in a second view, one flag is signaled to indicate whether the enabling of the plurality of loop filtering methods is inherited from the first picture in the first view or is signaled separately for the second picture in the second view.
[0086] In some embodiments, parameters used in a loop filtering method are signaled jointly for multiple views. In some embodiments, parameters of a cross-component offset filtering method are signaled jointly for multiple views. The parameters may include: an offset lookup table, a selection between only-offset, only-edge-offset, and edge-combined-with-offset, a filtering shape, a quantization step size, a downsampling filter, and a filtering unit size. In some embodiments, loop filtering parameters of a Wiener loop filtering method are signaled jointly for multiple views. The parameters may include: a selection of filtering parameters, filtering shape, and / or size, and / or a filtering unit size. In some embodiments, loop filtering parameters of CDEF are signaled jointly for multiple views.
[0087] In some embodiments, a flag is signaled to indicate whether one or more loop filtering parameters are shared from a first picture in a first view to a second picture in a second view. In some embodiments, a flag is signaled to indicate whether one or more loop filtering parameters are predicted from a first picture in a first view to a second picture in a second view. In some embodiments, a prediction residual on one or more loop filtering parameters is signaled for a second picture in a second view.
[0088] In some embodiments, some parameters of a loop filter are shared between two views, while the remaining parameters are signaled jointly, dependently, or independently for the two views. For example, the on / off flags for the two views are signaled separately, while the set of parameters controlling the loop filter is signaled independently for the two views.
[0089] Some embodiments apply loop filtering to the reconstruction of a first picture from a first view by using a reconstructed picture of a second picture from a second view as one of the inputs to the loop filtering process (e.g., cross-view loop filtering). In some embodiments, the input to the loop filter includes the reconstruction of the first picture in the first view, and the output includes an offset to be added on top of the reconstruction of the second picture in the second view. In some embodiments, the inputs of the original loop filtering method used in a single view are also used together as the inputs to the loop filtering method.
[0090] In some embodiments, a first component of the reconstruction of a first picture in a first view is used as an input for loop filtering a second component of the reconstruction of a second picture in a second view. The first component and the second component may be different components. In some embodiments, the input from the second view is samples around a co-located sample of the sample to be filtered in the first view. In some embodiments, the input from the second view is samples located by a disparity vector.
[0091] In some embodiments, the disparity vector is derived from the disparity vectors associated with adjacent blocks that use different views as reference pictures, such as inter-view prediction. In some embodiments, the precision of the disparity vector includes a predefined precision, such as integer pixels, half pixels, or quarter pixels.
[0092] In some embodiments, the filtering order of pictures from different views is different from the encoding order of pictures from different views. For example, for residual coding, the first picture from the first view is encoded first, and then the second picture from the second view is residual-encoded. Loop filtering is performed on the reconstruction of the second picture before loop filtering is performed on the reconstruction of the first picture.
[0093] In some embodiments, the reconstructed samples from two views are jointly used as the input to the loop filters for the respective views. For example, the reconstructed samples from two views are used to derive the filter parameters for each individual view.
[0094] In some embodiments, for multiple blocks belonging to different views, the block partitioning mode is signaled jointly. In some embodiments, these blocks are co-located blocks (e.g., blocks located at the same coordinates in different views). In some embodiments, the position of the blocks depends on the disparity of the views. In some cases, the disparity between views is quantized to certain predefined values. These values can be powers of 2 (or 4), such as 4, 8, 16.
[0095] In some embodiments, a flag is signaled for blocks from different views to indicate whether the partitioning mode of the blocks is signaled jointly or individually. In some embodiments, the flag is signaled conditionally. For example, the flag is signaled when the block is larger (or smaller) than a predefined threshold (e.g., 128×128, 64×64, 32×32, 16×16, or 8×8). In some embodiments, when the partitioning mode of the blocks is signaled jointly, the number of signaling syntaxes related to the partitioning mode is less than the number of blocks. For example, if N blocks (where N>1) share the same partitioning mode, the partitioning mode is signaled once instead of N times (for the N blocks). In some embodiments, when the partitioning mode of the blocks is signaled jointly, for a given first depth value, the partitioning of the shared blocks is the same, and beyond the first depth value, the partitioning modes of these blocks are signaled individually.
[0096] In some embodiments, when the block partitioning mode is signaled for multiple blocks from different views, the prediction mode of the blocks (e.g., intra prediction or inter prediction and its direction) is also signaled jointly. For example, the partitioning mode can specify how to partition the block into smaller block sizes. The partitioning mode can also specify whether different color components share the same partitioning.
[0097] In some embodiments, the block partitioning pattern of the first block in the first view is used to derive the context for signaling the partitioning pattern of the second block from the second view. The context provides a better estimate of the probability of bits having a certain value, which in turn improves the coding efficiency. In some embodiments, the same context is shared to encode the same type of syntax for different views. For example, when encoding the block partitioning pattern, the context associated with signaling the block partitioning pattern of the first view can be further updated by signaling the block partitioning pattern of the second view, and vice versa.
[0098] In some embodiments, the signaling of the syntax is interleaved instead of encoding different pictures from different views separately. For example, the syntax belonging to the second view can be written between multiple syntaxes written for the first view. In some embodiments, the signaling of the largest coding unit (LCU) for different views is interleaved. An example coding order can be LCU0 of view 0, LCU0 of Figure 1 view Figure 1 and LCU1 of view 0 and LCU1 of
[0099] view. The advantage of interleaving the signaling of the largest coding unit (LCU) is that context update can be performed more efficiently.
[0100] In some embodiments, one (or more) motion vectors derived from coded information associated with another view are used to derive the MVP of a block in the current view. In some embodiments, a motion vector associated with a first coded block of a first picture in a first view, i.e., a view-based motion vector predictor (VMVP), is used as the MVP of a second coded block of a second picture in a second view. In some embodiments, a disparity vector is used to extract the first coded block, and the disparity vector is derived from an adjacent block encoded using the second view as a reference frame. In one example, the disparity vector is (0,0), which means that the first coded block and the second coded block are located at the same coordinates in the first view and the second view, respectively. In another example, the disparity vector is derived from a global disparity vector associated with a combination between two frames from different views that are displayed simultaneously.
[0101] In some embodiments, one (or more) reference frames belonging to a first view are used to encode the first coded block. The display time of one (or more) reference frames of the first view is compared with the display time of the reference frame of a second block in the second view. If the display times are the same, the motion vector associated with the first coded block is used as the motion vector predictor of the second coded block. In some embodiments, one (or more) reference frames belonging to a first view are used to encode the first coded block. The display time of the reference frame is different from the display time of the reference frame of a second block in the second view. The motion vector associated with the first coded block is scaled and used as the motion vector predictor of the second coded block.
[0102] In some cases, the scaling factor is proportional to the ratio between the temporal distance between the first picture and its reference frame in the first view and the temporal distance between the second picture and its reference frame in the second view.
[0103] In some embodiments, the MVP of the current block is constructed. The checking order of the temporal MVP (TMVP) candidate or the VMVP candidate is different from that of other motion vector predictor candidates (e.g., spatial motion vector predictor candidate, global motion vector predictor candidate, motion vector predictor candidate in the motion vector library). In some embodiments, both the TMVP and the VMVP are in the motion vector list. Each is associated with a different index in the motion vector list. In some embodiments, the motion vector list has a predefined fixed (maximum) number of TMVP and / or VMVP. The TMVP / VMCP candidates in the list depend on the checking order and availability of the TMVP and VMVP candidates. In one example, two lists are constructed, i.e., the MVP list from multiple views and the list from the current frame. A high-level flag or a block-level flag is signaled to indicate which list is used.
[0104] In some embodiments, more than one motion vector extracted from a plurality of coded blocks in a first picture of a first view is used as an MVP candidate for a second block in a second picture of a second view. The extracted motion vectors may be added to an MVP list and used as the MVPs for the second block.
[0105] In some embodiments, the coded blocks are from coded blocks that are not adjacent to a first coded block with respect to a first picture in the first view. In some embodiments, a disparity vector is used to determine the first coded block, and the disparity vector is derived from adjacent blocks of a second block encoded using the second view as a reference frame. In some embodiments, the disparity vector is (0, 0), for example, the first coded block and the second coded block are located at the same coordinates in the first view and the second view, respectively.
[0106] In some embodiments, the disparity vector is derived from a global disparity vector associated with a combination between two frames from different views that are simultaneously displayed.
[0107] In some embodiments, if the motion vectors of one (or more) spatially (or temporally) adjacent blocks point to the first view, the motion vectors of the adjacent blocks are also inserted into the MVP list of the current block. In some embodiments, for the current block, an additional motion vector is signaled explicitly or derived implicitly. The motion vector is used to indicate a position displacement in a reference picture of the second view. In some embodiments, the coordinates of the coded block are predefined or derived implicitly using coded information such as block shape, quantization parameter, and / or temporal layer.
[0108] In some embodiments, a motion vector library associated with a block in a first picture of a first view is used to derive a motion vector predictor for a block in a second picture of a second view. In some embodiments, the motion vector library associated with a block in a first picture of a first view is merged with the motion vector library associated with a block in a second picture of a second view.
[0109] In some embodiments, the use or enabling of the above embodiments is controlled by an advanced syntax, including but not limited to sequence-level flags, picture-level flags, sub-picture-level flags, slice-level flags, tile-level flags, and maximum coded block-level flags.
[0110] Figure 10B is a flowchart showing a method 1050 for encoding video according to some embodiments. Method 1050 may be executed at a computing system (e.g., server system 112, source device 102, or electronic device 120) having a control circuit and a memory storing instructions executed by the control circuit. In some embodiments, method 1050 is executed by executing instructions stored in a memory (e.g., memory 314) of the computing system.
[0111] The system receives (1052) video data including a plurality of pictures, the plurality of pictures including a first picture corresponding to a first view and a second picture corresponding to a second view. The system determines (1054) whether loop filter parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are to be signaled jointly or separately. Based on determining that the loop filter parameters are to be signaled jointly, the system writes (1056) a first indicator in a multi-view video bitstream. The first indicator is used to indicate that the loop filter parameters are signaled jointly for the first picture and the second picture. As described above, the encoding process may reflect the decoding process described herein. For simplicity, these details are not repeated herein.
[0112] Although Figure 10A and Figure 10B the various logical stages are shown in a particular order, stages that are not order-dependent may be reordered and other stages may be combined or broken apart. Some reorderings or other groupings not specifically mentioned will be apparent to those of ordinary skill in the art, so the order and grouping presented herein are not exhaustive. Additionally, it should be recognized that these stages may be implemented in hardware, firmware, software, or any combination thereof.
[0113] Now turning to some example embodiments:
[0114] (A1)In one aspect, some embodiments include a method for video decoding (e.g., method 1000). In some embodiments, the method is performed at a computing system (e.g., server system 112) having a memory and control circuitry. In some embodiments, the method is performed at an encoding module (e.g., encoding module 320). In some embodiments, the method is performed at a source encoding component (e.g., source encoder 202), an encoding engine (e.g., encoding engine 212), and / or an entropy encoder (e.g., entropy encoder 214). The method includes: (i) receiving a multi-view video bitstream including a plurality of pictures, where the plurality of pictures includes a first picture corresponding to a first view and a second picture corresponding to a second view; (ii) determining, based on a first indicator in the multi-view video bitstream, whether the loop filter parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are signaled jointly or separately; and (iii) in accordance with the first indicator indicating that the loop filter parameters are signaled jointly, performing a first loop filtering process on the first picture and a second loop filtering process on the second picture using a shared set of loop filter parameters. For example, loop filter related parameters for a plurality of pictures from different views can be signaled jointly. In one example, a high-level syntax is signaled to indicate whether loop filter related parameters are signaled jointly for a plurality of pictures from different views or separately for a plurality of pictures from different views. In some embodiments, when the first indicator indicates that the loop filter parameters are signaled separately, performing a first loop filtering process on the first picture using a first set of loop filter parameters and performing a second loop filtering process on the second picture using a second set of loop filter parameters.
[0115] (A2)In some embodiments of A1, the method further includes: determining, based on a second indicator in the multi-view video bitstream, whether a particular loop filter technique is jointly enabled for the first picture and the second picture. For example, a high-level syntax is signaled jointly for a plurality of views, and the high-level syntax is used to signal the enabling of one or more loop filter methods.
[0116] (A3)In some embodiments of A2, the second indicator is a picture-level indicator. For example, the second indicator is signaled in a picture parameter set (PPS). As an example, a picture-level on / off flag for one or more loop filter methods is signaled jointly for a plurality of views.
[0117] (A4)In some embodiments of any one of A1 - A3, the method further includes: determining whether to enable a specific loop filtering technique for a first picture based on a second indicator in a multi-view video bitstream, where the second indicator depends on a corresponding indicator of a second picture. For example, depending on a picture-level on / off flag of a loop filtering method for one view, signaling a picture-level on / off flag of the loop filtering method in another view.
[0118] (A5)In some embodiments of A4, the second indicator indicates whether to inherit the value of the corresponding indicator for the first picture. For example, for a first picture in a first view, signaling three flags for different color components to specify whether to apply a first loop filtering method to the three color components separately, and then for a second view, signaling a flag to indicate whether to inherit the values of the three flags for a second picture in the second view.
[0119] (A6)In some embodiments of any one of A1 - A5, for a second picture, a set of indicators is written in the multi-view video bitstream. The set of indicators indicates whether to apply a corresponding loop filtering technique to the second picture. The second indicator is written in the multi-view video bitstream. The second indicator indicates whether the set of indicators is applicable to the first picture. For example, for a first picture in a first view, signaling multiple flags separately to specify whether to enable each of multiple loop filtering methods, and then for a second picture in a second view, signaling a flag indicating whether the enabling of the multiple loop filtering methods is inherited from the first picture in the first view or signaled separately for the second picture in the second view.
[0120] (A7)In some embodiments of any one of A1 - A6, a shared loop filtering parameter set corresponds to a first loop filtering technique. For example, signaling parameters used in a loop filtering method jointly for multiple views. As an example, for one or more loop filtering parameters, signaling a flag to indicate whether the associated one or more loop filtering parameters are shared from a first picture in a first view to a second picture in a second view.
[0121] (A8)In some embodiments of A7, the first loop filtering technique is one of the following: cross-component offset filtering technique; Wiener loop filtering technique; and constrained direction enhancement filtering technique. For example, loop filtering parameters used in the cross-component offset filtering method are signaled jointly for multiple views, such as the selection between an offset lookup table, band-only offset, edge-only offset, and band-edge combined offset, filtering shape, quantization step size, downsampling filter, and / or filtering unit size. As another example, loop filtering parameters used in the Wiener loop filtering method are signaled jointly for multiple views, such as the selection of filtering parameters, filtering shape and size, and / or filtering unit size. As another example, loop filtering parameters used in the constrained direction enhancement filter method are signaled jointly for multiple views.
[0122] (A9)In some embodiments of any one of A1 - A8, the method further includes: determining, based on a second indicator in the multi-view video bitstream, whether the loop filtering parameters of a first picture corresponding to a first view are to be predicted from corresponding loop filtering parameters of a second picture corresponding to a second view. In some embodiments, the residual of the predicted loop filtering parameters is written in the multi-view video bitstream. For example, for one or more loop filtering parameters, a flag is signaled to indicate whether one or more associated loop filtering parameters in the loop filtering parameters are predicted from a first picture in a first view to a second picture in a second view, and the residual of predicting the loop filtering parameters can also be signaled for the second picture in the second view.
[0123] (A10)In some embodiments of any one of A1 - A9, the first indicator indicates that a shared set of loop filtering parameters is signaled jointly, and a second set of loop filtering parameters is signaled separately for the first picture. For example, some parameters of a loop filter are shared between two views, while the remaining parameters are signaled jointly, dependently, or independently for the two views. For example, the on / off flags for two views are signaled separately, while the set of parameters controlling the loop filter is signaled independently for the two views.
[0124] (A11)In some embodiments of any one of A1 - A10, performing second loop filtering processing on a second picture includes using a reconstructed picture of a first picture as an input of the second loop filtering processing. For example, the loop filtering applied to the reconstruction of the first picture from a first view uses the reconstructed picture of the second picture from a second view as one of the inputs of the loop filtering processing (sometimes referred to as cross - view loop filtering). In some embodiments, the reconstructed picture is used to derive loop filtering parameters for the second loop filtering processing. In some embodiments, the reconstructed samples from two views are jointly used as inputs to the loop filters of each view. For example, the reconstructed samples from two views are used for classification to derive the filter parameters for each individual view.
[0125] (A12)In some embodiments of A11, performing second loop filtering processing on a second picture including using a reconstructed picture of a first picture includes: (i) obtaining a set of offsets based on the reconstructed picture of the first picture; (ii) obtaining the reconstruction of the second picture; and (iii) applying the set of offsets to the reconstruction of the second picture. For example, for a cross - component offset filtering method, and / or a Wiener loop filtering method and / or a constrained - direction enhancement filtering method, the input to the filter includes the reconstruction of the first picture in the first view, and the output is the offset to be added on top of the reconstruction of the second picture in the second view. In some embodiments, the first picture and the second picture are two views of the same picture. In some embodiments, the first picture and the second picture are pictures of two or more views simultaneously captured and / or displayed by different cameras. In some embodiments, the first picture and the second picture are time - synchronized.
[0126] (A13)In some embodiments of A11 or A12, the second picture is also used as an input to the second loop filtering processing. For example, the inputs of the original loop filtering method used in a single view are also used together as inputs to the loop filtering method.
[0127] (A14)In some embodiments of any one of A11 - A13, a first component of the reconstructed picture of the first picture is used as an input for loop - filtering a second component of the second picture, and wherein the first component is a different component from the second component. For example, the first component of the reconstruction of the first picture in the first view can be used as an input for loop - filtering the second component of the reconstruction of the second picture in the second view, and the first component and the second component can be different components (e.g., different color components).
[0128] (A15)In some embodiments of any one of A11 - A14, performing second loop filtering processing on a second picture includes using a reconstructed picture of a first picture, including: using one or more samples from the reconstructed picture that are co - located with a sample to be filtered in the first picture. For example, the input from the second view can be samples around a co - located sample of the sample to be filtered in the first view. In some embodiments, a disparity vector is used to identify one or more samples from the reconstructed picture. For example, the input from the second view can be a sample located by the disparity vector. In some embodiments, the disparity vector is derived from disparity vectors associated with adjacent blocks that use different views as reference pictures (e.g., the adjacent blocks use an inter - view prediction mode). In some embodiments, the precision of the disparity vector is limited to a predefined precision, such as integer pixels, half - pixels, or quarter - pixels.
[0129] (A16)In some embodiments of any one of A1 - A15, multiple pictures have an encoding order and a filtering order, and wherein, the encoding order is different from the filtering order. For example, the filtering order of pictures from different views can be different from the encoding order of pictures from different views. For example, for residual encoding, a first picture from a first view is encoded first, and then a second picture from a second view is residual - encoded. However, before loop filtering the reconstruction of the first picture, one or more loop filtering operations are performed on the reconstruction of the second picture.
[0130] (B1)In another aspect, some embodiments include a method of video coding (e.g., method 1050). In some embodiments, the method is performed on a computing system having a memory and one or more processors. The method includes: (i) receiving video data including multiple pictures, where the multiple pictures include a first picture corresponding to a first view and a second picture corresponding to a second view; (ii) determining whether the loop filtering parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are to be signaled jointly or separately; and (iii) writing a first indicator into the multi - view video bitstream according to determining that the loop filtering parameters are to be signaled jointly, where the first indicator is used to indicate that the loop filtering parameters are signaled jointly for the first picture and the second picture.
[0131] (B2)In some embodiments of B1, the method further includes: writing a second indicator into the multi - view video bitstream, where the second indicator is used to indicate whether a specific loop filtering technique is jointly enabled for the first picture and the second picture.
[0132] (B3)In some embodiments of B1 or B2, the method further includes: writing a second indicator into a multi-view video bitstream, where the second indicator is used to indicate whether one or more loop filter indicators of a first picture are applicable to a second picture.
[0133] (C1)In another aspect, some embodiments include a method for visual media data processing. In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) obtaining a source video sequence corresponding to a set of views; and (ii) performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, where the multi-view video bitstream includes: (a) a first plurality of encoded pictures corresponding to a first view, the first plurality of encoded pictures including a first picture; (b) a second plurality of encoded pictures corresponding to a second view, the second plurality of encoded pictures including a second picture; and (c) a first indicator, where the first indicator is used to indicate whether loop filter parameters are signaled jointly for the first picture and the second picture.
[0134] (D1)In another aspect, some embodiments include a method for video decoding. In some embodiments, the method is executed on a computing system having a memory and one or more processors. The method includes: (i) receiving a multi-view video bitstream including a plurality of pictures, where the plurality of pictures includes a first picture corresponding to a first view and a second picture corresponding to a second view; (ii) obtaining a reconstructed first picture by performing a first loop filter process on the first picture; and (iii) performing a second loop filter process on the second picture using the reconstructed first picture as an input to the second loop filter process. In some embodiments, the first picture and the second picture are two views of the same picture. In some embodiments, the first picture and the second picture are pictures of two or more views simultaneously captured and / or displayed by different cameras. In some embodiments, the first picture and the second picture are temporally synchronized.
[0135] In another aspect, some embodiments include a computing system (e.g., server system 112), the computing system including a control circuit (e.g., control circuit 302) and a memory (e.g., memory 314) coupled to the control circuit, the memory storing one or more instruction sets configured to be executed by the control circuit, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1 - A16, B1 - B3, C1, and D1 above).
[0136] In another aspect, some embodiments include a non-transitory computer-readable storage medium storing one or more instruction sets that are executed by a control circuit of a computing system, the one or more instruction sets including instructions for performing any of the methods described herein (e.g., A1-A16, B1-B3, C1, and D1 above).
[0137] Unless otherwise specified, any syntactic element described herein can be high-level syntax (HLS). As used herein, HLS is signaled at a level higher than the block level. For example, HLS can correspond to the sequence level, frame level, slice level, or tile level. As another example, HLS elements can be signaled in a video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), adaptation parameter set (APS), slice header, picture header, tile header, and / or CTU header.
[0138] It should be understood that although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. The proper nouns used herein are for the purpose of describing particular embodiments only and are not intended to limit the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0139] As used herein, the term "when" can be interpreted as "if" or "upon" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the prerequisite is true, depending on the context. Similarly, the phrase "if it is determined that [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as "upon determining" or "in response to determining" or "in accordance with determining" or "upon detecting" or "in response to detecting" that the prerequisite is true, depending on the context. As used herein, N refers to a variable number. Unless explicitly stated otherwise, different instances of N can refer to the same number (e.g., the same integer value, such as the number 2) or different numbers.
[0140] For purposes of explanation, the above description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the claims to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the operating principles and the practical application, thereby enabling others skilled in the art to understand.
Claims
1. A method for video decoding, the method being executed on a computing system including a memory and one or more processors, the method comprising: Receiving a multi-view video bitstream, the multi-view video bitstream including a plurality of pictures, wherein the plurality of pictures includes a first picture corresponding to a first view and a second picture corresponding to a second view; Based on a first indicator in the multi-view video bitstream, determining whether the loop filter parameters for the first picture corresponding to the first view and the second picture corresponding to the second view are signaled jointly or separately; and According to the first indicator indicating that the loop filter parameters are signaled jointly, performing a first loop filtering process on the first picture and a second loop filtering process on the second picture using a shared loop filter parameter set.
2. The method according to claim 1 further comprises: Based on a second indicator in the multi-view video bitstream, determining whether a specific loop filtering technique is jointly enabled for the first picture and the second picture.
3. The method according to claim 2, wherein The second indicator is a picture-level indicator.
4. The method according to claim 1 further comprises: Based on the second indicator in the multi-view video bitstream, determining whether to enable a specific loop filtering technique for the first picture, wherein the second indicator depends on a corresponding indicator of the second picture.
5. The method according to claim 4, wherein The second indicator indicates whether to inherit the value of the corresponding indicator for the first picture.
6. The method according to claim 1, wherein: For the second picture, an indicator set is written in the multi-view video bitstream, the indicator set indicating whether a corresponding loop filtering technique is to be applied to the second picture; And A second indicator is written in the multi-view video bitstream, the second indicator indicating whether the indicator set is applicable to the first picture.
7. The method according to claim 1, wherein, The shared loop filter parameter set corresponds to a first loop filtering technique.
8. The method according to claim 7, wherein The first loop filtering technique is one of the following: Cross-component offset filtering technique; Wiener loop filtering technique; and Constrained directional enhancement filtering technique.
9. The method according to claim 1 further comprises: Based on the second indicator in the multi-view video bitstream, determining whether the loop filter parameters for the first picture corresponding to the first view are to be predicted from the corresponding loop filter parameters of the second picture corresponding to the second view.
10. The method according to claim 1, wherein The first indicator indicates that the shared loop filter parameter set is signaled jointly, and wherein a second loop filter parameter set is signaled separately for the first picture.
11. The method according to claim 1, wherein, Performing the second loop filtering process on the second picture includes using the reconstructed picture of the first picture as an input to the second loop filtering process.
12. The method according to claim 11, wherein, The performing the second loop filtering process on the second picture including using the reconstructed picture of the first picture includes: Obtaining an offset set based on the reconstructed picture of the first picture; Obtaining the reconstruction of the second picture; and Applying the offset set to the reconstruction of the second picture.
13. The method according to claim 11, wherein, The second picture is also used as an input to the second loop filtering process.
14. The method according to claim 11, wherein, Use a first component of a reconstructed picture of the first picture as an input for loop filtering a second component of the second picture, and wherein the first component is a component different from the second component.
15. The method according to claim 11, wherein, Performing the second loop filtering process on the second picture includes using the reconstructed picture of the first picture includes: using one or more samples from the reconstructed picture that are co-located with samples to be filtered in the first picture.
16. The method according to claim 1, wherein, The plurality of pictures have an encoding order and a filtering order, and wherein the encoding order is different from the filtering order.
17. A computing system, comprising: Control circuitry; A memory; And One or more instruction sets stored in the memory and configured to be executed by the control circuitry, the one or more instruction sets including instructions for: Receiving video data including a plurality of pictures, wherein the plurality of pictures include a first picture corresponding to a first view and a second picture corresponding to a second view; Determining whether loop filter parameters corresponding to the first picture corresponding to the first view and the second picture corresponding to the second view are to be signaled jointly or separately; and In accordance with determining that the loop filter parameters are to be signaled jointly, writing a first indicator into a multi-view video bitstream, the first indicator being for indicating that the loop filter parameters are signaled jointly for the first picture and the second picture.
18. The computing system according to claim 17, further comprising: Writing a second indicator into the multi-view video bitstream, the second indicator being for indicating whether a particular loop filtering technique is jointly enabled for the first picture and the second picture.
19. The computing system according to claim 17, further comprising: Writing a second indicator into the multi-view video bitstream, the second indicator being for indicating whether one or more loop filtering indicators of the first picture are applicable to the second picture.
20. A non-transitory computer-readable storage medium storing one or more instruction sets configured to be executed by a computing device including control circuitry and a memory, the one or more instruction sets including instructions for: Obtain a source video sequence corresponding to a view set; And Performing a conversion between the source video sequence and a multi-view video bitstream of visual media data, wherein, The multi-view video bitstream includes: A first plurality of encoded pictures corresponding to a first view, the first plurality of encoded pictures including a first picture; A second plurality of encoded pictures corresponding to a second view, the second plurality of encoded pictures including a second picture; and A first indicator for indicating whether loop filter parameters are signaled jointly for the first picture and the second picture.