Multi-view viewpoint position supplementation enhancement information message

By utilizing SEI messages for multi-view video coding, the method optimizes intra-prediction and motion vector prediction, addressing redundancy issues in multi-view video encoding and decoding, thereby enhancing compression efficiency and quality in video streaming and conferencing.

JP7815354B2Active Publication Date: 2026-02-17TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024125133
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-06-27
Filing Date
2024-07-31
Publication Date
2026-02-17
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently encoding and decoding multi-view video data, particularly in managing redundancy and optimizing intra-prediction directions, which affects compression efficiency and bitstream representation.

Method used

The introduction of supplemental enhancement information (SEI) messages, including scalability dimension information (SDI) and multi-view view position (MVP) information, allows for precise determination and display of viewpoint positions in video decoding, enhancing compression efficiency by optimizing intra-prediction and motion vector prediction mechanisms.

Benefits of technology

This approach improves video coding efficiency by reducing redundancy and bitstream size, enabling effective encoding and decoding of multi-view video data, particularly in applications like video conferencing and streaming, while maintaining high-quality video rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007815354000004
    Figure 0007815354000004
  • Figure 0007815354000005
    Figure 0007815354000005
  • Figure 0007815354000006
    Figure 0007815354000006
Patent Text Reader

Abstract

To provide a method, a device, and a non-transitory computer-readable storage medium for video decoding.SOLUTION: A device includes a video decoding device processing circuit. The processing circuit decodes pictures associated with a viewpoint of each layer in a coded video sequence (CVS) from the bitstream and determines a viewpoint position to be displayed in one dimension on the basis of a first supplemental enhancement information (SEI) message and a second SEI message. The first SEI message includes scalability dimension information (SDI) and the second SEI message includes multiview viewpoint position (MVP) information. The processing circuit also displays the decoded pictures on the basis of the viewpoint position in the one dimension.SELECTED DRAWING: Figure 14
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001]

[0001] Reference This application claims priority to U.S. patent application Ser. No. 17 / 850,581, filed June 27, 2022, entitled "Multi-view viewpoint position supplemental enhancement information message," which in turn claims priority to U.S. provisional application Ser. No. 63 / 215,944, filed June 28, 2021, entitled "Techniques for multi-view viewpoint position SEI message for coded video streams," the disclosures of which are incorporated herein by reference in their entirety.

[0002]

[0002] Technical Field This disclosure describes embodiments generally related to video coding. [Background technology]

[0003]

[0003] background The background discussion provided herein is intended to generally present the context of the present disclosure. Work under the names of the current inventors is not admitted, expressly or impliedly, as prior art to the present disclosure to the extent that that work is described in this background section or in a descriptive manner that would not otherwise qualify it as prior art at the time of filing.

[0004]

[0004] Coding and decoding of images and / or video can be performed using inter-picture prediction with motion compensation. Uncompressed digital images and / or video can include a series of pictures, each with spatial dimensions of, for example, 1920 x 1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally known as the frame rate), for example, 60 pictures per second, or 60 Hz. Uncompressed images and / or video have specific bitrate requirements. For example, 1080p60 4:2:0 video (1920 x 1080 luminance sample resolution at a 60 Hz frame rate) with 8 bits per sample requires a bandwidth approaching 1.5 Gbit / s. One hour of such video requires more than 600 Gbytes of storage space.

[0005] One of the goals of image and / or video coding and decoding can be said to be the reduction of redundancy in the input image and / or video signal through compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, in some cases by more than two orders of magnitude. While the description herein uses video encoding / decoding as an illustrative example, the same techniques can be applied to image encoding / decoding in a similar manner without departing from the spirit of this disclosure. Both lossless and non-lossless compression, as well as combinations thereof, can be used. Lossless compression refers to a technique in which an exact copy of the original signal can be reconstructed from the compressed original signal. When non-lossless compression is used, the reconstructed signal may not be identical to the original signal, but the distortion between the original and reconstructed signal is small enough that the reconstructed signal is useful for the intended application. In the case of video, non-lossless compression is widely used. The amount of acceptable distortion depends on the application; for example, users of a particular consumer streaming application may be able to tolerate higher distortion than users of a television distribution application. The achievable compression ratio may reflect that a higher acceptable / tolerable distortion may result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques in several broad categories, including, for example, motion compensation, transform, quantization, and entropy coding.

[0007]

[0007] Video codec technology can include a technique known as intra-coding. In intra-coding, sample values ​​are represented without reference to samples or other data from previously reconstructed reference pictures. In some video codecs, pictures are spatially subdivided into blocks of samples. If all blocks of samples are coded in intra mode, the picture can be considered an intra-picture. Intra-pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and therefore can be used as the first picture in a coded video bitstream and video session or as a still image. Samples in intra-blocks can be subjected to a transform, and the transform coefficients can be quantized before entropy coding. Intra-prediction can be a technique that minimizes sample values ​​in the pre-transform domain. In some cases, the smaller the DC value and AC coefficients after the transform, the fewer bits required for a given quantization step size to represent the block after entropy coding.

[0008]

[0008] Traditional intra-coding, such as that known from MPEG-2 generation coding techniques, does not use intra-prediction. However, some newer video compression techniques include techniques that attempt to predict intra-prediction from surrounding sample data and / or metadata obtained during encoding / decoding of blocks of data, e.g., spatially adjacent and preceding in decoding order. Such techniques are hereinafter referred to as "intra-prediction" techniques. Note that, at least in some cases, intra-prediction uses only reference data from the current picture being reconstructed, not from a reference picture.

[0009]

[0009] Many different forms of intra-prediction may exist. If more than one such technique can be used in a given video coding technique, the technique used may be coded as an intra-prediction mode. In some cases, a mode may have sub-modes and / or parameters, which may be coded separately or may be included in a mode codeword. The codeword used for a given mode, sub-mode, and / or parameter combination may affect the coding efficiency gain from intra-prediction and, therefore, the entropy coding technique used to convert the codeword into a bitstream.

[0010]

[0010] Certain modes of intra prediction were introduced in H.264, improved in H.265, and further refined in new coding techniques such as the Joint Search Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). A predictor block can be formed using neighboring sample values ​​belonging to already available samples. The sample values ​​of neighboring samples are copied into the predictor block according to a certain direction. The reference to the direction in use can be coded in the bitstream or can be predicted itself.

[0011]

[0011] Referring to FIG. 1A, shown at the bottom right is a subset of nine predictor directions known from the 33 possible predictor directions of H.265 (corresponding to the 33 angular modes out of the 35 intra modes). The point where the arrows converge (101) represents the predicted sample. The arrows indicate the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from one or more samples to the upper right at an angle of 45 degrees from horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from one or more samples to the lower left of sample (101) at an angle of 22.5 degrees from horizontal.

[0012] 1A, a square block (104) of 4x4 samples is shown in the upper left (indicated by a thick dashed line). The square block (104) contains 16 samples, each labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample (from the top) in the Y dimension and the first sample (from the left) in the X dimension. Similarly, sample S44 is the fourth sample in the block (104) in both the Y and X dimensions. Because the block is 4x4 samples in size, S44 is located in the lower right. Also shown are reference samples following a similar numbering scheme. The reference samples are labeled with R and their Y position (e.g., row index) and X position (column index) relative to the block (104). In both H.264 and H.265, the prediction samples are adjacent to the block being reconstructed; therefore, there is no need to use negative values.

[0013]

[0013] Intra-picture prediction can be performed by copying reference sample values ​​from neighboring samples, as assigned by the signaled prediction direction. For example, assume that the coded video bitstream includes signaling indicating, for this block, a prediction direction consistent with arrow (102)—i.e., the sample is predicted from one or more prediction samples pointing upward and to the right at a 45-degree angle from horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. And sample S44 is predicted from reference sample R08.

[0014] In some cases, the values ​​of multiple reference samples may be combined, for example by interpolation, to calculate the reference sample; especially when the direction is not evenly divisible by 45 degrees.

[0015] The number of possible directions has increased as video coding technology has evolved. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013), and as of the time of this disclosure, JEM / VVC / BMS can support up to 65 directions. Experiments are performed to identify the most likely directions, and specific techniques in entropy coding are used to represent those possible directions using fewer bits while accepting certain penalties for less likely directions. Furthermore, the direction itself can often be predicted from neighboring directions used in adjacent, already decoded blocks.

[0016] FIG. 1B shows a schematic diagram (105) illustrating 65 intra-prediction directions according to JEM to illustrate the increasing number of prediction directions over time.

[0017]

[0017] The mapping of intra-prediction direction bits in a coded video bitstream to represent directions can vary from one video coding technique to another; for example, they can range from simple, direct mappings of prediction directions to complex adaptive schemes involving intra-prediction modes, codewords, most-probable modes, and similar techniques. However, in any given case, there may be certain directions that are statistically less likely to occur in the video content than certain other directions. Because the goal of video compression is to reduce redundancy, these unlikely directions will be represented with more bits than more likely directions in a well-performing video coding technique.

[0018]

[0018] Motion compensation may be a non-lossless compression technique in which blocks of sample data from a previously reconstructed picture or part thereof (reference picture) are used to predict a newly reconstructed picture or part of a picture after being spatially shifted in the direction indicated by a motion vector (hereinafter referred to as MV). In some cases, the reference picture may be the same as the picture currently being reconstructed. MV may have two dimensions, X and Y, or three dimensions, the third of which is an indication of the reference picture in use (the latter may indirectly be considered a temporal dimension).

[0019] In some video compression techniques, the motion vector applicable to a particular area of ​​sample data can be predicted from other motion vectors, e.g., from other areas of sample data spatially adjacent to the area being reconstructed and preceding that MV in decoding order. This can significantly reduce the amount of data required to code the motion vectors, thereby eliminating redundancy and increasing compression. For example, when coding an input video signal derived from a camera (known as natural video), MV prediction can work effectively because there is a statistical likelihood that areas larger than the area to which a single motion vector is applicable will move in similar directions, and thus, in some cases, can be predicted using similar motion vectors derived from MVs in neighboring areas. This results in a motion vector for a given area that is found to be similar or identical to the motion vector predicted from surrounding MVs, which, after entropy coding, can be represented using fewer bits than would be used to code the motion vector directly. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., MV) derived from the original signal (i.e., sample stream). In other cases, the MV prediction itself may be non-lossless, for example due to rounding errors when computing a predictor from several surrounding MVs.

[0020]

[0020] Various MV prediction mechanisms are described in H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016). Among the many MV prediction mechanisms provided by H.265, the one described in this application is a technique hereafter referred to as "spatial merge".

[0021]

[0021] Referring to Figure 2, a current block (201) contains samples that have been found by the encoder during a motion search process to be predictable from a spatially shifted previous block of the same size. Instead of directly coding its MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the most recent reference picture (in decoding order) using MVs associated with any of five surrounding samples (202 to 206, respectively) denoted A0, A1, and B0, B1, B2. In H.265, MV prediction can use predictors from the same reference picture as neighboring blocks. Summary of the Invention

[0022]

[0022] Aspects of the disclosure provide methods and apparatuses for video encoding and decoding. In some examples, the video decoding apparatus includes a processing circuit configured to decode, from a bitstream, pictures associated with a viewpoint of each layer in a coded video sequence (CVS). The processing circuit is capable of determining a viewpoint position to be displayed in one dimension based on a first supplemental enhancement information (SEI) message and a second SEI message. The first SEI message includes scalability dimension information (SDI), and the second SEI message includes multiview view position (MVP) information. The processing circuit is capable of displaying the decoded picture based on the viewpoint position in the one dimension.

[0023] In an embodiment, for each viewpoint of each layer of the CVS, the processing circuitry can determine a viewpoint identifier for the respective viewpoint based on the first SEI message, and can determine a respective viewpoint position to be displayed in one dimension based on the viewpoint position variable in the second SEI message and the determined viewpoint identifier.

[0024] In an embodiment, the processing circuit derives a first value indicating the number of viewpoints from a first SEI message. The processing circuit is capable of obtaining a second value related to the number of viewpoints from a second SEI message. The second value is different from the first value. The processing circuit is capable of comparing the first value with 1 plus the second value in a compatibility check.

[0025] In an embodiment, the second SEI message is associated with an intra random access point (IRAP) access unit of the CVS.

[0026] In an embodiment, the viewpoint position is applied to the access unit of the CVS.

[0027] In an embodiment, the first SEI message is included in the CVS based on the inclusion of the first SEI message, and the second SEI message is included in the CVS based on the inclusion of the first SEI message and the second SEI message in the CVS.

[0028]

[0028] In an embodiment, the CVS includes a third SEI message including 3D reference view information. The 3D reference view information indicates a viewpoint identifier for a left view, a viewpoint identifier for a right view, and a reference view width for the reference view. The processing circuitry can derive a first value indicating the number of viewpoints from the first SEI message. The processing circuitry can obtain a third value related to the number of viewpoints from the third SEI message. The third value is different from the first value. The processing circuitry can compare the first value with 1 plus the third value in a compatibility test.

[0029] In an embodiment, the third SEI message indicates the number of reference indications signaled in the third SEI message, and the maximum value of the number of reference indications is based on a third value related to the number of viewpoints in the third SEI message.

[0030]

[0030] In an embodiment, the first SEI message is an SDI SEI message, and the second SEI message is an MVP SEI message.

[0031] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a video decoding method. [Brief explanation of the drawings]

[0032]

[0032] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1A]

[0033] FIG. 1A is a schematic diagram of an example subset of intra-prediction modes. [Figure 1B]

[0034] FIG. 1B is an example of an exemplary intra-prediction direction. [Figure 2]

[0035] FIG. 2 shows a current block (201) and surrounding samples in an embodiment. [Figure 3]

[0036] FIG. 3 is a simplified block diagram schematic of a communication system (300) according to an embodiment. [Figure 4]

[0037] FIG. 4 is a simplified block diagram schematic of a communication system (400) according to an embodiment. [Figure 5]

[0038] FIG. 5 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment. [Figure 6]

[0039] FIG. 6 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment. [Figure 7]

[0040] FIG. 7 shows a block diagram of an encoder according to another embodiment. [Figure 8]

[0041] FIG. 8 shows a block diagram of a decoder according to another embodiment. [Figure 9]

[0042] FIG. 9 shows a diagram of an autostereoscopic display according to an embodiment of the present disclosure. [Figure 10A]

[0043] FIG. 10A illustrates an example of reordering pictures based on multi-view viewpoint positions according to an embodiment of the present disclosure. [Figure 10B]

[0044] FIG. 10B illustrates an example of reordering pictures based on multi-view viewpoint positions according to an embodiment of the present disclosure. [Figure 10C] FIG. 0C illustrates an example of reordering pictures based on multi-view viewpoint positions according to an embodiment of the present disclosure. [Figure 11]

[0045] FIG. 11 illustrates example syntax for a supplemental enhancement information (SEI) message for indicating viewpoint positions for multi-view video according to an embodiment of the present disclosure. [Figure 12]

[0046] FIG. 12 illustrates an example syntax for an SEI message indicating 3D reference display information according to an embodiment of the present disclosure. [Figure 13]

[0047] FIG. 13 shows a flow chart outlining an encoding process according to an embodiment of the present disclosure. [Figure 14]

[0048] FIG. 14 shows a flow chart outlining a decoding process according to an embodiment of the present disclosure. [Figure 15]

[0049] FIG. 15 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0033]

[0050] Figure 3 illustrates a simplified block diagram of a communication system (300) according to one embodiment of the present disclosure. The communication system (300) includes multiple terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of Figure 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) can code video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The coded video data can be transmitted in the form of one or more coded video bitstreams. The terminal device (320) can receive the coded video data from the network (350), decode the coded video data to reconstruct a video picture, and display the video picture according to the reconstructed video data. Unidirectional data transmission may be common in media serving applications, etc.

[0034]

[0051] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) for bidirectional transmission of coded video data, such as may occur during a video conference. For bidirectional data transmission, for example, each of the terminal devices (330) and (340) can code video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) over the network (350). Each of the terminal devices (330) and (340) can also receive coded video data transmitted by the other of the terminal devices (330) and (340), decode the coded video data to reconstruct the video pictures, and display the video pictures on an accessible display device in accordance with the reconstructed video data.

[0035]

[0052] In the example of FIG. 3 , terminal devices 310, 320, 330, and 340 are shown as a server, a personal computer, and a smartphone, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure find application with laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network 350 represents any number of networks that carry coded video data between terminal devices 310, 320, 330, and 340, including, for example, wired and / or wireless communication networks. Communication network 350 can exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of the present disclosure, the architecture and topology of network 350 may not be important to the operation of the present disclosure, unless otherwise described below.

[0036]

[0053] Figure 4 illustrates the placement of a video encoder and video decoder in a streaming environment as an example application of the disclosed subject matter, which is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, and storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0037]

[0054] The streaming system may include a video source (401) and a capture subsystem (413), which may include, for example, a digital camera, capable of generating a stream of uncompressed video pictures (402). In one example, the video picture stream (402) includes samples captured by the digital camera. The video picture stream (402), depicted as a thicker line to emphasize the greater amount of data compared to the encoded video data (404) (or coded video bitstream), may be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) may include hardware, software, or a combination thereof, and may enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (404) (or encoded video bitstream (404)), depicted as a thin line to emphasize its smaller amount of data compared to the stream of video pictures (402), can be stored on the streaming server (405) for future use. One or more streaming client subsystems, such as the client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include a video decoder (410), for example, within the electronic device (430). The video decoder (410) decodes the incoming copy of the encoded video data (407) and generates an output stream of video pictures (411) that can be rendered on a display (412) (e.g., a display screen) or other rendering device (not shown).In some streaming systems, the encoded video data 404, 407, and 409 (e.g., video bitstreams) may be encoded according to a particular video coding / compression standard. Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0038]

[0055] It should be noted that electronic devices 420 and 430 may include other components (not shown). For example, electronic device 420 may include a video decoder (not shown), and electronic device 430 may include a video encoder (not shown).

[0039]

[0056] 5 shows a block diagram of a video decoder (510) according to one embodiment of the present disclosure. The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., receiving circuitry). The video decoder (510) can be used in place of the video decoder (410) in the example of FIG. 4.

[0040]

[0057] The receiver (531) can receive one or more coded video sequences to be decoded by the video decoder (510); in the same or another embodiment, it can receive one coded video sequence at a time, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences can be received from a channel (501), which can be a hardware or software link to a storage device that stores the coded video data. The receiver (531) can receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which can be transferred using respective entities (not shown). The receiver (531) can separate the coded video sequences from other data. To address network jitter, a buffer memory (515) can be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter referred to as the "parser (520)"). In certain applications, the buffer memory (515) is part of the video decoder (510). In other cases, it may be external to the video decoder (510) (not shown). In yet another example, there may be a buffer memory (not shown) external to the video decoder (510), for example, to deal with network jitter, and even another buffer memory (515) internal to the video decoder (510), for example, to handle playback timing. If the receiver (531) is receiving data from a store-and-forward device with sufficient bandwidth and controllability or from a synchronous network, the buffer memory (515) may not be needed or may be small.For use in best-effort packet networks such as the Internet, a buffer memory (515) may be required, which may be relatively large and may advantageously be adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) outside the video decoder (510).

[0041]

[0058] The video decoder (510) may include a parser (520) for reconstructing symbols (521) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (510) and, potentially, information for controlling a rendering device, such as a rendering device (512) (e.g., a display screen) that is not an integral part of the electronic device (530) but may be coupled to the electronic device (530), as shown in FIG. 5. The rendering device control information may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context effects, etc. The parser (520) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (520) can also extract from the coded video sequence information such as transform coefficients, quantization parameter values, motion vectors, etc.

[0042]

[0059] The parser (520) is capable of performing an entropy decoding / parsing process on the video sequence received from the buffer memory (515) to generate symbols (521).

[0043]

[0060] The reconstruction of the symbols (521) may include several different units depending on the type of coded video picture or part thereof (inter- and intra-picture, inter- and intra-block) and other factors. Which units are included and how can be controlled by subgroup control information parsed by the parser (520) from the coded video sequence. The flow of such subgroup control information between the parser (520) and the following units is not depicted for clarity.

[0044]

[0061] Beyond the functional blocks already described, the video decoder (510) may be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating within commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0045]

[0062] The first unit is a scalar / inverse transform unit (551), which receives quantized transform coefficients as well as control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) from the parser (520) as symbols (521). The scalar / inverse transform unit (551) can output blocks containing sample values ​​that can be input to an aggregator (555).

[0046]

[0063] In some cases, the output samples of the scaler / inverse transform (551) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) generates blocks of the same size and shape as the block being reconstructed using already reconstructed surrounding information retrieved from the current picture buffer (558). The current picture buffer (558), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (555) optionally adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (552) to the output sample information as provided by the scaler / inverse transform unit (551).

[0047]

[0064] In other cases, the output samples of the scalar / inverse transform unit (551) may relate to a block that may be inter-coded and motion-compensated. In such cases, the motion-compensated prediction unit (553) may access the reference picture memory (557) to retrieve samples used for prediction. After motion-compensating the retrieved samples according to the symbols (521) associated with the block, these samples are added by the aggregator (555) to the output of the scalar / inverse transform unit (551) to generate output sample information (called residual samples or residual signals in this case). The addresses in the reference picture memory (557) from which the motion-compensated prediction unit (553) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (553), for example, in the form of symbols (521), which may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​taken from a reference picture memory (557), motion vector prediction mechanisms, etc., where sub-sample accurate motion vectors are used.

[0048]

[0065] The output samples of the aggregator (555) may be subjected to various loop filtering techniques in a loop filter unit (556). The video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (556) as symbols (521) from the parser (520), which may be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.

[0049]

[0066] The output of the loop filter unit (556) may be a sample stream that can be output to a rendering device (512) or stored in a reference picture memory (557) for use in future inter-picture prediction.

[0050]

[0067] Once a given coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of a subsequent coded picture.

[0051]

[0068] The video decoder (510) can perform decoding operations according to a predetermined video compression technique in a standard, such as ITU-T Rec. H.265. A coded video sequence can conform to the syntax specified by the video compression technique or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technique or standard and the profile documented in the video compression technique or standard. Specifically, a profile can select certain tools from among all tools available in the video compression technique or standard as the only tools that can be used under that profile. Compliance also requires that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the levels may in some cases be further constrained by the Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.

[0052]

[0069] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0053]

[0070] 6 shows a block diagram of a video encoder (603) according to one embodiment of the present disclosure. The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmitting circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0054]

[0071] The video encoder (603) can receive video samples from a video source (601) (which in the example of FIG. 6 is not part of the electronic device (620)) that can capture video images to be coded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0055]

[0072] The video source (601) may provide a source video sequence to be coded by the video encoder (603) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (601) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (601) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, where each pixel may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion focuses on examples.

[0056]

[0073] According to one embodiment, the video encoder (603) is capable of coding and compressing pictures of a source video sequence into a coded video sequence (643) in real time or under any other time constraint required by the application. Imposing an appropriate coding rate is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below, the coupling of which is not depicted for clarity. Parameters set by the controller (650) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured with other appropriate functions associated with the video encoder (603) optimized for a particular system design.

[0057]

[0074] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As a simplified explanation, in one example, the coding loop may include a source coder (630) (e.g., responsible for generating a symbol stream based on an input picture to be coded and reference pictures) and a (local) decoder (533) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in a manner similar to that generated by a (remote) decoder (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques contemplated by the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference picture memory (634). Because decoding the symbol stream yields bit-exact results independent of the location (local or remote) of the decoder, the contents in the reference picture memory (634) are also bit-exact between the local and remote encoders. In other words, the predictor of the encoder "sees" the exact same sample values ​​for the reference picture samples that the decoder would "see" if it were using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is used in several related technologies as well.

[0058]

[0075] The operation of the "local" decoder (633) may be the same as that of a "remote" decoder, such as the video decoder (510), already described in detail above in connection with Figure 5. However, briefly referring also to Figure 5, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (645) and parser (520) may be lossless, the entropy decoding portion of the video decoder (510), including the buffer memory (515) and parser (520), may not be fully implemented in the local decoder (633).

[0059]

[0076] In embodiments, decoder techniques other than analysis / entropy decoding present in a decoder are also present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder techniques can be omitted, as they are the reverse of the decoder techniques described generically. Only in certain areas will a more detailed description be provided below.

[0060]

[0077] In operation, the source coder (630) may, in some instances, perform motion-compensated predictive coding, which predicts and encodes an input picture by reference to one or more previously coded pictures from a video sequence that have been designated as “reference pictures.” In this manner, the coding engine (632) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that can be selected as predictive references for the input picture.

[0061]

[0078] The local video decoder (633) can decode coded video data of pictures that can be designated as reference pictures based on symbols generated by the source coder (630). The operation of the coding engine (632) can advantageously be a non-lossless process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (633) can repeat the decoding process that can be performed by the video decoder with respect to the reference pictures, causing the reconstructed reference pictures to be stored in the reference picture cache (634). In this way, the video encoder (603) can locally store copies of reconstructed reference pictures that have common content with reconstructed reference pictures obtained by a far-end video decoder (assuming there are no transmission errors).

[0062]

[0079] The predictor (635) can perform a prediction search for the coding engine (632). That is, for a new picture to be coded, the predictor (635) can search the reference picture memory (534) for sample data (as candidate reference pixel blocks) or predetermined metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (535) can operate on a sample-block-pixel-block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (635), an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (634).

[0063]

[0080] The controller (650) may manage the coding operations of the source coder (630), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0064]

[0081] All outputs of the aforementioned functional units can be subjected to entropy coding in an entropy coder (645), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0065]

[0082] The transmitter (640) can buffer the coded video sequence, as produced by the entropy coder (645), and prepare it for transmission over a communication channel (660), which may be a hardware or software link to a storage device that stores the coded video data. The transmitter (640) can merge the coded video data from the video encoder (603) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0066]

[0083] The controller (650) may manage the operation of the video encoder (603). During coding, the controller (650) may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, a picture may frequently be assigned as one of the following picture types:

[0067]

[0084] An intra picture (I-picture) is one that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0068]

[0085] A predicted picture (P-picture) can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0069]

[0086] Bidirectionally predicted pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a block.

[0070]

[0087] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples) and can be coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture. For example, blocks of an I-picture can be non-predictively coded, or they can be predictively coded with reference to previously coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture can be predictively coded with spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture can be predictively coded with spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0071]

[0088] The video encoder (603) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In this operation, the video encoder (603) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard being used.

[0072]

[0089] In one embodiment, the transmitter (640) can transmit additional data along with the coded video. The source coder (630) can include such data as part of the coded video sequence. The additional data can include temporal, spatial, and SNR enhancement layers, as well as other forms of redundant data (such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.).

[0073]

[0090] Video can be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture under encoding / decoding, referred to as the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a reference picture that was previously coded and is still buffered in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0074]

[0091] In some embodiments, bi-prediction techniques may be used for inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, that both precede the current picture in decoding order (but may be past and future, respectively, in display order) in the video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block may be predicted by a combination of the first and second reference blocks.

[0075]

[0092] Furthermore, to improve coding efficiency, it is possible to use merge mode techniques for inter-picture prediction.

[0076]

[0093] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed on a block-by-block basis. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​(e.g., luma values) for pixels, such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0077]

[0094] 7 shows a diagram of a video encoder (703) according to another embodiment of the present disclosure. The video encoder (703) is configured to receive a processed block of sample values ​​(e.g., a predictive block) in a current video picture in a sequence of video pictures and encode the processed block into a coded picture that is part of a coded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) of the example of FIG. 4.

[0078]

[0095] In an HEVC example, the video encoder (703) receives a matrix of sample values ​​for a processing block, such as an 8x8 sample prediction block. The video encoder (703) determines whether the processing block is best coded using intra mode, inter mode, or bi-prediction mode, e.g., using rate-distortion optimization. If the processing block is to be coded in intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into a coded picture; if the processing block is to be coded in inter mode or bi-prediction mode, the video encoder (703) may use inter prediction techniques or bi-prediction techniques, respectively, to encode the processing block into a coded picture. In certain video coding techniques, the merge mode may be an inter prediction picture sub-mode, in which case the motion vector is derived from one or more motion vector predictors without the benefit of any motion vector components coded outside the predictor. In certain other video coding techniques, there may be motion vector components applicable to the current block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown), to determine the mode of the processing block.

[0079]

[0096] In the example of Figure 7, the video encoder (703) includes an inter-encoder (730), an intra-encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in Figure 7.

[0080]

[0097] The inter-encoder (730) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., a block in a previous picture and a block in a subsequent picture), generate inter-prediction information (e.g., a description of redundant information due to the inter-coding technique, motion vectors, merge mode information), and calculate an inter-prediction result (e.g., a predicted block) based on the inter-prediction information using any suitable technique. In some examples, the reference picture is a decoded reference picture that has been decoded based on coded video information.

[0081]

[0098] The intra encoder (722) is configured to receive samples of a current block (e.g., a processing block), optionally compare the block to previously coded blocks in the same picture, generate transformed and quantized coefficients, and optionally generate intra prediction information (e.g., intra prediction direction information according to one or more intra coding techniques). In one example, the intra encoder (722) also calculates intra prediction results (e.g., predicted blocks) based on the intra prediction information and reference blocks in the same picture.

[0082]

[0099] The general-purpose controller (721) is configured to determine general control data and control other components of the video encoder (703) based on the general control data. In one example, the general-purpose controller (721) determines the mode of the block and provides a control signal to the switch (726) based on the mode. For example, if the mode is intra mode, the general-purpose controller (721) controls the switch (726) to select intra mode results for use by the residual calculator (723) and the entropy encoder (725) to select intra prediction information and include it in the bitstream; if the mode is inter mode, the general-purpose controller (721) controls the switch (726) to select inter prediction results for use by the residual calculator (723) and the entropy encoder (725) to select inter prediction information and include it in the bitstream.

[0083]

[0100] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-encoder (722) or the inter-encoder (730). The residual encoder (724) is configured to operate on the residual data to encode the residual data and generate transform coefficients. In one example, the residual encoder (724) is configured to transform the residual data from the spatial domain to the frequency domain to generate transform coefficients. The transform coefficients are then subjected to a quantization process to obtain quantized transform coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform an inverse transform and generate decoded residual data. The decoded residual data can be used by the intra-encoder (722) and the inter-encoder (730), as appropriate. For example, the inter-encoder (730) may generate decoded blocks based on the decoded residual data and inter-prediction information, and the intra-encoder (722) may generate decoded blocks based on the decoded residual data and intra-prediction information. The decoded blocks are processed appropriately to generate decoded pictures, which may be buffered in memory circuitry (not shown) and, in some instances, used as reference pictures.

[0084]

[0101] The entropy encoder (725) is configured to format a bitstream to include the coded blocks. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra-prediction information or inter-prediction information), residual information, and other suitable information in the bitstream. Note that, in accordance with the disclosed subject matter, residual information is not present when coding blocks in a merged sub-mode of either an inter mode or a bi-prediction mode.

[0085]

[0102] 8 shows a diagram of a video decoder (810) according to another embodiment of the present disclosure. The video decoder (810) is configured to receive coded pictures that are part of a coded video sequence and decode the coded pictures to generate reconstructed pictures. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0086]

[0103] In the example of Figure 8, the video decoder (810) includes an entropy decoder (871), an inter-decoder (880), a residual decoder (873), a reconstruction module (874), and an intra-decoder (872) coupled together as shown in Figure 8.

[0087]

[0104] The entropy decoder (871) can be configured to reconstruct, from a coded picture, certain symbols that represent syntax elements from which the coded picture is constructed. Such symbols can include, for example, the mode in which the block is coded (e.g., intra mode, inter mode, bi-prediction mode, merge submode, or the latter two in another submode), prediction information (e.g., intra-prediction information or inter-prediction information) that can identify certain samples or metadata used for prediction by the intra decoder (872) or inter decoder (880), respectively, residual information (e.g., in the form of quantized transform coefficients), etc. In one example, if the prediction mode is an inter- or bi-prediction mode, the inter-prediction information is provided to the inter decoder (880); if the prediction type is an intra-prediction type, the intra-prediction information is provided to the intra decoder (872). The residual information can be inverse quantized and provided to the residual decoder (873).

[0088]

[0105] The inter decoder (880) is configured to receive inter prediction information and generate inter prediction results based on the inter prediction information.

[0089]

[0106] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0090]

[0107] The residual decoder (873) is configured to perform inverse quantization to extract unquantized transform coefficients and process the unquantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantization parameters (QPs)), which may be provided by the entropy decoder (871) (this may be only a small amount of control information, so a data path is not depicted).

[0091]

[0108] The reconstruction module (874) is configured to combine, in the spatial domain, the residual as output by the residual decoder (873) with the prediction result (possibly output by an inter- or intra-prediction module) to form a reconstructed block, which may be part of a reconstructed picture, which may be part of a reconstructed video. It should be noted that other appropriate processes, such as deblocking processes, may be performed to improve visual quality.

[0092]

[0109] It should be noted that the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using any suitable technology. In some embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (403), (603), and (703) and the video decoders (410), (510), and (810) may be implemented using one or more processors executing software instructions.

[0093]

[0110] According to embodiments of the present disclosure, a bitstream may include one or more coded video sequences (CVSs). A CVS may be coded independently of other CVSs. Each CVS may include one or more layers, where each layer may be a representation of video with a particular quality (e.g., spatial resolution) or a representation of a particular component interpretation property, such as a depth map, a transparency map, or a perspective view. In the temporal dimension, each CVS may include one or more access units (AUs). Each AU may include one or more pictures of different layers corresponding to the same time instance. A coded layer video sequence (CLVS) is a per-layer CVS that may include a sequence of picture units in the same layer. If a bitstream has multiple layers, the CVSs in the bitstream may include one or more CLVSs for each layer.

[0094]

[0111] In one embodiment, a CVS includes a sequence of AUs, where the sequence of AUs includes, in decoding order, an Intra Random Access Point (IRAP) AU followed by zero or more AUs that are not IRAP AUs. In one example, the zero or more AUs include all subsequent AUs, but not any subsequent AUs that are IRAP AUs. In one example, a CLVS includes a sequence of pictures and associated non-Video Coding Layer (VCL) Network Abstraction Layer units of the base layer of the CVS.

[0095]

[0112] Aspects of the present disclosure include a multi-view viewpoint supplemental enhancement information (SEI) message for a coded video stream.

[0096]

[0113] According to some aspects of the present disclosure, video can be classified as single-view video and multi-view video. For example, single-view video (e.g., monoscopic video) is a two-dimensional medium that provides a viewer with a single perspective of a scene. Multi-view video can provide multiple perspectives of a scene, creating a sense of realism for the viewer. In one example, 3D video can provide two perspectives, such as a left and right perspective corresponding to a human viewer. The two perspectives can be displayed simultaneously or nearly simultaneously using different polarizations of light, and the viewer wears polarized glasses so that each eye of the viewer receives a different perspective.

[0097]

[0114] In another example, a display device, such as an auto-stereoscopic display device, is configured to generate different pictures depending on the viewer's eye position and does not require glasses for viewing, and may be referred to as a glasses-less 3D display.

[0098]

[0115] Multi-view video can be created by simultaneously capturing a scene using multiple cameras, where the cameras are appropriately positioned so that each camera captures the scene from one viewpoint. The multiple cameras can capture multiple video sequences corresponding to the multiple viewpoints. To accommodate more viewpoints, more cameras can be used to generate a multi-view video with multiple video sequences associated with the viewpoints. Multi-view video can require a large amount of storage space and / or high bandwidth for transmission. Multi-view video coding techniques have been developed in the field to reduce the required storage space or transmission bandwidth.

[0099]

[0116] To improve the efficiency of multi-view video coding, similarities between views are exploited. In some implementations, one of multiple views, referred to as a base view, is coded like a monoscopic video. For example, intra-picture and / or temporal inter-picture prediction is used when coding the base view. The base view can be decoded using a monoscopic decoder (e.g., a monoscopic decoder) that performs intra-picture and inter-picture prediction. Views other than the base view in multi-view video can be referred to as dependent views. In addition to intra-picture and inter-picture prediction, inter-view prediction with disparity compensation may be used to code the dependent views. In one example, in inter-view prediction, a current block in a dependent view is predicted using a reference block of samples from a picture of another view at the same time instance. The location of the reference block is indicated by a disparity vector. Inter-view prediction is similar to inter-(picture) prediction, except that motion vectors are replaced with disparity vectors and temporal reference pictures are replaced with reference pictures from other views.

[0100]

[0117] According to some aspects of the present disclosure, multiview coding can employ a multi-layer approach. A multi-layer approach can multiplex different coded representations of a video sequence (e.g., HEVC coded ones), called layers, into one bitstream. The layers may be dependent on each other. The dependencies can be used in inter-layer prediction to exploit similarities between different layers and achieve increased compression performance. A layer can represent texture, depth, or other auxiliary information of a scene associated with a particular camera perspective. In some examples, all layers belonging to the same camera perspective are referred to as a view, and layers carrying the same type of information (e.g., texture or depth) are referred to as components within the scope of the multiview video.

[0101]

[0118] According to one aspect of the present disclosure, multi-view video coding may include the addition of high level syntax (HLS) (e.g., above the slice level) to an existing single-layer decoding core. In some examples, multi-view viewpoint coding does not change the syntax or decoding process required for single-layer coding below the slice level (e.g., of HEVC). This may allow existing implementations to be reused without significant modifications to build a multi-view video decoder. For example, the multi-view video decoder may be implemented based on the video decoder (510) or the video decoder (810).

[0102]

[0119] In some examples, all pictures related to the same capture or display time instance are included in an AU and have the same picture order count (POC). Multiview video coding enables inter-view prediction, which performs prediction from pictures within the same AU. For example, pictures decoded from other views can be inserted into one or both of the reference picture lists of the current picture. Furthermore, in some examples, motion vectors may be actual temporal motion vectors if they relate to temporal reference pictures of the same view, or disparity vectors if they relate to inter-view reference pictures. A block-level motion compensation module (e.g., block-level encoding software or hardware, block-level decoding software or hardware) can be used, which operates the same regardless of whether the motion vectors are temporal motion vectors or disparity vectors.

[0103]

[0120] According to one aspect of the present disclosure, multi-view video coding allows for coding pictures of different views at a display time instance in an order (e.g., decoding order) that is not necessarily related to the corresponding display position order or display order.

[0104]

[0121] FIG. 9 shows a diagram of an autostereoscopic display (900) in some examples. The autostereoscopic display (900) can display different viewpoint pictures at different positions. In one example, the autostereoscopic display (900) displays different viewpoint pictures depending on the detected eye position of the observer. In the example of FIG. 9, the eye position of the observer can be detected in one dimension, such as between a left-most position and a right-most position. For example, When the observer's eye position is at E0, the autostereoscopic display (900) displays a picture of the viewpoint identified by the viewpoint identifier ViewId[0]; When the observer's eye position is at E1, the autostereoscopic display (900) displays a picture of the viewpoint identified by the viewpoint identifier ViewId[1]; When the observer's eye position is at E2, the autostereoscopic display (900) displays a picture of the viewpoint identified by the viewpoint identifier ViewId[2]; When the observer's eye position is at E3, the autostereoscopic display (900) displays the picture of the viewpoint with viewpoint identifier ViewId[3]. The order of the observer's eye positions from left to right is E2, E0, E1, and E3.

[0105]

[0122] 10A illustrates an example of reordering pictures based on multi-view viewpoint positions according to an embodiment of the present disclosure. Given multiple or different viewpoints, viewpoint position information indicating a display order along one dimension (e.g., display from left to right along the horizontal axis) can be signaled for an intended user experience. The display order can be specified by the multiple viewpoints or multi-view viewpoint positions of the corresponding pictures. In one example, the multiple viewpoints include a first viewpoint to be displayed at a first multi-view viewpoint position and a second viewpoint to be displayed at a second multi-view viewpoint position. A first picture of the first viewpoint can be displayed at the first multi-view viewpoint position. A second picture of the second viewpoint can be displayed at the second multi-view viewpoint position.

[0106]

[0123] In some examples, pictures P0-P3 correspond to multiple views (e.g., ViewId[0]-ViewId[3]). Pictures P0-P3 are associated with the same capture time instance or the same display time instance. Pictures P0-P3 may be coded (encoded or decoded) in a coding order (e.g., encoding order or decoding order) that may be determined based on specific coding requirements, such as rate-distortion optimization. The decoding order may be specified based on pictures P0-P3 or based on corresponding views ViewId[0]-ViewId[3]. Referring to FIG. 10A , an example of a decoding order is P0 of view ViewId[0], P1 of view ViewId[1], P2 of view ViewId[2], and P3 of view ViewId[3], where P0 of view ViewId[0] is decoded prior to decoding P1-P3 of ViewId[1]-ViewId[3], respectively. Alternatively, the decoding order is ViewId[0], ViewId[1], ViewId[2], and ViewId[3].

[0107]

[0124] In one embodiment, the decoding order is different from the display order of pictures (e.g., P0-P3) or corresponding views (e.g., ViewId[0]-ViewId[3]). According to an embodiment of the present disclosure, the decoded pictures (e.g., P0-P3) or corresponding views (e.g., ViewId[0]-ViewId[3]) can be rearranged according to the display order. The display order can be signaled for an intended user experience. Referring to FIG. 10A , the display order is ViewId[2], ViewId[0], ViewId[1], and ViewId[3], where ViewId[2] is displayed as the leftmost view and ViewId[3] is displayed as the rightmost view. The corresponding display order based on the pictures is P2, P0, P1, P3. The views (and decoded pictures) are rearranged according to the display order.

[0108]

[0125] In some examples, an AU includes coded information of pictures of different views associated with the same capturing or display time instance. For example, an AU includes coded information of picture P0 of view ViewId[0], picture P1 of view ViewId[1], picture P2 of view ViewId[2], and picture P3 of view ViewId[3]. The coding order of pictures P0-P3 does not necessarily follow the order of the observer's eye positions.

[0109]

[0126] 10B-10C show an example of rearranging pictures according to multi-view viewpoint positions in an example. FIG. 10B shows pictures in an AU decoded according to decoding order. In the example of FIG. 10B, pictures P0-P3 in the AU are decoded in the order of P0, P1, P2, and P3. Because the order of the observer's eye positions from left to right is E2, E0, E1, and E3, the order of the decoded pictures does not correspond to the order of the observer's eye positions from left to right.

[0110]

[0127] In some examples, the decoded pictures can be reordered according to a display order signaled for the intended user experience, for example, in one example, the display order is related to the order of the observer's eye positions, such as from left to right.

[0111]

[0128] 10C shows the decoded pictures in the AUs, in some examples, rearranged according to display order, e.g., the display order is related to the order of the observer's eye positions from left to right.

[0112]

[0129] According to some aspects of the present disclosure, supplemental enhancement information (SEI) messages may be included in an encoded bitstream, for example, to aid in decoding and / or display of the encoded bitstream, or for another purpose. The SEI message(s) may include information not necessary for decoding, such as decoding samples of a coded picture from a VCL NAL unit. The SEI message(s) may be optional for constructing luma or chroma samples by the decoding process. In some examples, the SEI message(s) are not required to reconstruct luma or chroma samples during the decoding process. Furthermore, decoders that comply with video coding standards that support SEI messages are not required to process the SEI message(s) in a compliant manner. Some coding standards may require some SEI message information to check bitstream conformance or output timing decoder conformance. The SEI message(s) may be selectively processed by conforming decoders for output order conformance to a particular standard (e.g., HEVC 265 or VVC). In an embodiment, the SEI message is present in the bitstream.

[0113]

[0130] SEI messages can contain various types of data that indicate the timing of video pictures, or that describe various properties of the coded video, or that indicate how various properties can be used or extended. In some instances, SEI messages do not affect the core decoding process, but can indicate how the video is recommended to be post-processed or displayed.

[0114]

[0131] The SEI message can be used to provide additional information about the encoded bitstream, which can be used to modify the presentation of the bitstream once it is decoded or to provide information to the decoder. For example, SEI messages have been used to provide frame packing information (e.g., describing how video data is arranged within a video frame), content descriptions (e.g., indicating that the encoded bitstream is, for example, 360-degree video), and color information (e.g., color gamut and / or color range), among others.

[0115]

[0132] In some examples, the SEI message can be used to signal to a decoder that the encoded bitstream contains 360-degree video. The decoder can use this information to render the video data for a 360-degree presentation. Alternatively, if the decoder is not capable of rendering 360-degree video, the decoder can use this information to not render the video data.

[0116]

[0133] According to an embodiment of the present disclosure, a first SEI message may indicate scalability dimension information (SDI) of multi-view video. For example, the first SEI message (e.g., an SDI SEI message) may include the number and type of scalability dimensions, such as information indicating the number of views of the multi-view video. According to an embodiment of the present disclosure, a view identifier (ID) (e.g., denoted as ViewID) of the i-th layer in the current CVS may be signaled in the first SEI message (e.g., an SDI SEI message). In one example, the syntax element sdi_view_id_val[i] specifies the value of the view ID (e.g., ViewID) of the i-th layer in the current CVS. The length of the syntax element sdi_view_id_val[i] (e.g., the number of bits used to represent the syntax element sdi_view_id_val[i]) may be specified by the syntax element sdi_view_id_len_minus1. The unit of length may be bits. In one example, the value of the length of the syntax element sdi_view_id_val[i] is equal to the value of the syntax element sdi_view_id_len_minus1 plus one.

[0117]

[0134] According to one aspect of the present disclosure, the display order can be signaled using a second SEI message, such as an SEI message indicating multi-view viewpoint position (MVP) information. In some related examples, the second SEI message can include information indicating a viewpoint position in one dimension (e.g., along a horizontal axis). For example, the second SEI message can include a viewpoint position in one dimension corresponding to the position of the observer's eyes. The second SEI message (e.g., an MVP SEI message) can be useful, for example, for a glasses-less (or naked-eye) 3D display, to indicate a relative viewpoint position along the horizontal axis of a 3D rendering view.

[0118]

[0135] FIG. 11 shows example syntax (1100) for a second SEI message (e.g., an MVP SEI message) indicating viewpoint position for multi-view video. For example, the MVP SEI message specifies the relative viewpoint position within the CVS along the horizontal axis of the 3D rendering view. When multiple viewpoints are present, MVP information in the MVP SEI message indicating left-to-right display order can be signaled for the intended user experience.

[0119]

[0136] The MVP SEI message can be associated with an IRAP AU of the CVS. The MVP information signaled in the MVP SEI message can apply to the entire CVS. In some examples, the MVP SEI message can signal the number of viewpoints and then signal the positions of the viewpoints, respectively.

[0120]

[0137] In the example of Figure 11, a parameter denoted by num_views_minus1 can be signaled by an MVP SEI message as shown at (1110). The parameter num_views_minus1 indicates the number of views (e.g., a parameter denoted as NumViews), for example, in an access unit. For example, the number of views is equal to the sum of 1 and the value of the parameter num_views_minus1.

[0121]

[0138] The order of the viewpoints in one dimension (e.g., display order), such as the position of the viewpoint in the display order, can be signaled in the MVP SEI message using the parameter view_positions[i], where the value i is an integer between 0 and num_views_minus1, as shown at (1120) in Figure 11.

[0122]

[0139] The parameter view_position[i] may indicate the order of a view (e.g., display order) among the left-to-right views (e.g., all views) for display purposes, with the view identifier (ID) ViewId equal to the value of the syntax element sdi_view_id_val[i] in the first SEI message. The order value of the leftmost view may be equal to 0, and the order value may increase by 1 for the next view from left to right. The value of the parameter view_position[i] may be in the range from 0 to 62, inclusive.

[0123]

[0140] As an example, to signal the viewpoint positions in the example of FIG. 10A, "3" can be signaled as the parameter num_views_minus1 to indicate that the number of viewpoints is four. "1" is signaled as view_position[0] for ViewId=0, "2" is signaled as view_position[1] for ViewId=1, "0" is signaled as view_position[2] for ViewId 2, For ViewId 3, "3" is signaled as view_position[3], where "0" is the leftmost view and "3" is the rightmost view in left-to-right position. Then, when views ViewId[0]-ViewId[3] are decoded from the access unit, Viewpoint ViewId[0] has view_position[0], The viewpoint ViewId[1] has view_position[1], Viewpoint ViewId[2] has view_position[2], Viewpoint ViewId[3] has view_position[3]. The viewpoints ViewId[0]-ViewId[3] are rearranged according to the corresponding viewpoint positions view_position[0] to view_position[3] to obtain the display order in FIG. 10A.

[0124]

[0141] According to an embodiment of the present disclosure, pictures associated with a view point of each layer in the CVS can be decoded from the bitstream. A position of the view point to be displayed in one dimension (e.g., along the horizontal axis) can be determined based on the first SEI message and the second SEI message. The decoded pictures can be displayed based on the position of the view point in one dimension (e.g., along the horizontal axis). For each of the view points of each layer in the CVS, a view ID (e.g., ViewId[i]) of the respective view point can be determined based on the first SEI message. A position of each view point to be displayed in one dimension (e.g., along the horizontal axis) can be determined based on the view point position variable (e.g., view_position[i]) in the second SEI message and the determined view identifier (e.g., ViewId[i]).

[0125]

[0142] 10A and 11, in one example, the value of the syntax element num_view_minus1 is 3, and the number of views to be displayed is 3 plus 1, or 4. The integer i of the syntax element sdi_view_id_val[i] in the first SEI message (e.g., SDI SEI message) is 0, 1, 2, or 3, which corresponds to the syntax elements di_view_id_val[0], sdi_view_id_val[1], sdi_view_id_val[2], or sdi_view_id_val[3], respectively. In one example, sdi_view_id_val[0] is 0, sdi_view_id_val[1] is 1, sdi_view_id_val[2] is 2, and sdi_view_id_val[3] is 3. Based on the second SEI message (eg, MVP SEI message), the syntax element view_position[i] indicates the display order or viewpoint position with the ViewId equal to the syntax element sdi_view_id_val[i] in the first SEI message. The syntax element view_position[0] indicates the viewpoint position of ViewId equal to sdi_view_id_val[0] (eg, 0), and the viewpoint ViewId[0] is at view_position[0] (eg, 1). The syntax element view_position[1] indicates the viewpoint position of ViewId equal to sdi_view_id_val[1] (eg, 1), and the viewpoint ViewId[1] is at view_position[1] (eg, 2). The syntax element view_position[2] indicates the viewpoint position of ViewId equal to sdi_view_id_val[2] (eg, 2), and the viewpoint ViewId[2] is at view_position[2] (eg, 0). The syntax element view_position[3] indicates the viewpoint position of ViewId equal to sdi_view_id_val[3] (eg, 3), and the viewpoint ViewId[3] is at view_position[3] (eg, 3).

[0126]

[0143] According to one aspect of the present disclosure, an MVP SEI message includes one or more parameters that can have a semantic dependency with respect to information in an SDI SEI message. For example, the syntax element num_views_minus1 in an MVP SEI message has a semantic dependency with respect to the value of the parameter NumViews, which is derived from the SDI SEI message. In one example, a first value indicating the number of views (e.g., the value of the parameter NumViews) is derived from the SDI SEI message. A second value associated with the number of views (e.g., the value of the syntax element num_views_minus1) is obtained from the MVP SEI message. The second value may differ from the first value. The number “1” plus the second value can be compared to the first value in a conformance check. Referring to FIG. 10A , the first value is 4 and the second value is 3. The second value (e.g., 3) plus 1 is equal to the first value (e.g., 4), satisfying the conformance check. Therefore, MVP SEI messages may be subject to the constraints associated with SDI SEI messages.

[0127]

[0144] In one example, if the CVS does not include an SDI SEI message, the CVS does not include an MVP SEI message. In some examples, if the SDI SEI message does not exist, the MVP SEI message does not exist. In one example, the MVP SEI message is included in the CVS based on the SDI SEI message being included in the CVS.

[0128]

[0145] In an embodiment, the third SEI message specifies that 3D reference display information is signaled. Figure 12 shows example syntax (1200) in the third SEI message for specifying 3D reference display information according to an embodiment of the present disclosure.

[0129]

[0146] The third SEI message (e.g., the 3D Reference View Information SEI message shown in FIG. 12) may include information about the reference view width, the reference viewing distance, and the corresponding reference stereo pair (e.g., a pair of viewpoints to be displayed to the viewer's left and right eyes relative to the reference view at the reference viewing distance). The reference view information may enable the viewpoint renderer to generate an appropriate stereo pair for the target screen width and viewing distance. The values ​​of the reference view width and the reference viewing distance may be signaled in centimeters. The viewpoint pair specified in the third SEI message may be used to extract or estimate parameters related to the camera center-to-center distance of the reference stereo pair, which may be used to generate the viewpoint of the target view. In the case of a multi-view display, the reference stereo pair may correspond to a pair of viewpoints that may be simultaneously observed by the viewer's left and right eyes.

[0130]

[0147] If present, the third SEI message may be associated with an AU (e.g., an IRAP AU or a non-IRAP AU) if all AUs that follow it in decoding order follow it in output order. The third SEI message may apply to the current AU and all AUs that follow it in both output order and decoding order, up to but not including the next IRAP AU or the next AU that contains another reference display information SEI message.

[0131]

[0148] A third SEI message (e.g., a 3D Reference View Information SEI message) can specify the view parameters for which the 3D video sequence is optimized and the corresponding reference parameters. Each reference view (e.g., a reference view width and optionally a corresponding viewing distance) can be associated with one reference pair of viewpoints by signaling a respective viewpoint identifier (e.g., denoted as ViewId). The difference between the ViewId values ​​is referred to as the baseline distance (e.g., the center-to-center distance of the cameras used to acquire the 3D video sequence).

[0132]

[0149] The following formulas can be used to determine the baseline distance and horizontal shift for a receiver's display when the ratio between the receiver's viewing distance and the reference viewing distance is equal to the ratio between the receiver's screen width and the reference screen width:

[0133] baseline[i] = refBaseline[i] × (refDisplayWidth[i]÷displayWidth) Eq.1 shift[i] = refShift[i] * (refDisplayWidth[i]÷displayWidth) Eq.2 where refBaseline[i] is equal to (right_view_id[i] - left_view_id[i]) signaled in the third SEI message. Other parameters related to viewpoint generation may be obtained and determined by using similar formulas.

[0134] parameter[i] = refParameter[i] * (refDisplayWidth[i]÷displayWidth) Eq.3

[0150] refParameter[i] is a parameter related to viewpoint synthesis corresponding to the reference pair signaled by left_view_id[i] and right_view_id[i]. In the above formula, the width of the visible portion of the display used to display the video sequence can be understood under "display width". The above formula can be used to determine the viewpoint pair and horizontal shift or other viewpoint synthesis parameters when the viewing distance is not scaled proportionally to the screen width compared to the reference display parameter, where the effect of applying the above formula is to maintain the perceived depth at the viewing distance in the same proportion as at the reference setting.

[0135]

[0151] If the view synthesis-related parameters corresponding to the reference stereo pair change from one AU to another, the view synthesis-related parameters may be scaled by the same scaling factor as the parameters in the AU to which the third SEI message is associated. The above formula may be applied to obtain the parameters for the subsequent AU, where refParameter[i] is the parameter related to the reference stereo pair associated with the subsequent access unit.

[0136]

[0152] The horizontal shift for the receiver's display can be corrected by scaling it by the same factor used to scale the baseline distance (or other viewpoint synthesis parameter).

[0137]

[0153] If the CVS does not include the first SEI message (eg, the SDI SEI message), the CVS does not include the third SEI message (eg, the 3D reference information SEI message).

[0138]

[0154] 12, the syntax element prec_ref_display_width may specify the exponent of the maximum allowable truncation error for refDisplayWidth[i] as given by (2-prec_ref_display_width). The value of the syntax element prec_ref_display_width may be in the range of 0 to 31, inclusive.

[0139]

[0155] The syntax element ref_viewing_distance_flag equal to 1 may indicate the presence of a reference viewing distance. The syntax element ref_viewing_distance_flag equal to 0 indicates the absence of a reference viewing distance.

[0140]

[0156] The syntax element prec_ref_viewing_dist may specify the exponent of the maximum allowable truncation error for refViewingDist[i] as given by (2-prec_ref_viewing_dist). The value of prec_ref_viewing_dist may be in the range of 0 to 31 inclusive.

[0141]

[0157] The syntax element num_views_minus1 plus 1 may be equal to NumViews derived from a first SEI message (e.g., an SDI SEI message) for CVS. According to one aspect of the disclosure, the third SEI message includes one or more parameters having semantic dependencies regarding information in the first SEI message (e.g., an SDI SEI message). In one example, a first value (e.g., NumView) specifying the number of views is derived from the first SEI message. A third value (e.g., num_views_minus1) related to the number of views is obtained from the third SEI message. The third value may differ from the first value. One (e.g., the value “1”) plus the third value may be compared to the first value in a conformance check.

[0142]

[0158] The syntax element num_ref_displays_minus1 plus 1 may specify the number of reference displays signaled in the third SEI message. The maximum value of the number of reference displays may be based on a third value (e.g., num_views_minus1) related to the number of views in the third SEI message. The value of num_ref_displays_minus1 may be in the range of 0 to num_views_minus1, inclusive.

[0143]

[0159] The syntax element left_view_id[i] may indicate the ViewId of the left viewpoint of the stereo pair corresponding to the i-th reference view.

[0144]

[0160] The syntax element right_view_id[i] may indicate the ViewId of the right view of the stereo pair corresponding to the i-th reference view.

[0145]

[0161] The syntax element exponent_ref_display_width[i] may specify the exponent part of the reference display width for the i-th reference display. The value of exponent_ref_display_width[i] may be in the range 0 to 62, inclusive. The value 63 is reserved for future use by ITU-T|ISO / IEC. Decoders may treat the value 63 as indicating an unspecified reference display width.

[0146]

[0162] The syntax element mantissa_ref_display_width[i] can specify the mantissa part of the reference display width for the i-th reference display. The variable refDispWidthBits, which specifies the number of bits for the syntax element mantissa_ref_display_width[i], can be derived as follows:

[0163] If the syntax element exponent_ref_display_width[i] is equal to 0, refDispWidthBits is set equal to Max(0, prec_ref_display_width - 30).

[0147]

[0164] Otherwise (0 < exponent_ref_display_width[i] < 63), refDispWidthBits is set equal to Max(0, exponent_ref_display_width[i] + prec_ref_display_width - 31).

[0148]

[0165] The syntax element exponent_ref_viewing_distance[i] may specify the exponent part of the reference viewing distance for the i-th reference view. The value of exponent_ref_viewing_distance[i] may be in the range 0 to 62, inclusive. The value 63 is reserved for future use by ITU-T|ISO / IEC. Decoders may treat the value 63 as indicating an unspecified reference view width.

[0149]

[0166] The syntax element mantissa_ref_viewing_distance[i] can specify the mantissa of the reference viewing distance of the i-th reference view. The syntax element refViewDistBits is a variable that specifies the number of bits of mantissa_ref_viewing_distance[i], and can be derived as follows:

[0150]

[0167] If the syntax element exponent_ref_viewing_distance[i] is equal to 0, refViewDistBits is set equal to Max(0, prec_ref_viewing_distance - 30).

[0151]

[0168] Otherwise (0 < exponent_ref_viewing_distance[i] < 63), efViewDistBits is set equal to Max(0, exponent_ref_viewing_distance[i] + prec_ref_viewing_distance - 31).

[0152]

[0169] Table 1 shows an example association between syntax elements and camera parameter variables in the third SEI message.

[0153] Table 1 - Correlation between camera parameter variables and syntax elements

[0154]

number

[0170] The variables in the x row of Table 1 are derived from the respective variables or values ​​in the e, n, and v rows of Table 1 as follows: If e is not equal to 0, the following applies: x=2 (e-31) ×(1+n÷2 v ) Otherwise (e.g., if e is equal to 0), the following applies: x=2 -(30+v) ×n

[0171] 12, a syntax element additional_shift_present_flag[i] equal to 1 may indicate that information regarding an additional horizontal shift of the left and right viewpoints of the i-th reference view is present in the third SEI message. A syntax element additional_shift_present_flag[i] equal to 0 may indicate that information regarding an additional horizontal shift of the left and right viewpoints of the i-th reference view is not present in the third SEI message.

[0155]

[0172] The syntax element num_sample_shift_plus512[i] may indicate a recommended additional horizontal shift for the stereo pair corresponding to the ith reference baseline and the ith reference view.

[0156]

[0173] If (num_sample_shift_plus512[i] - 512) is less than 0, the left viewpoint of the stereo pair corresponding to the ith reference baseline and the ith reference view can be shifted left by (512 - num_sample_shift_plus512[i]) samples relative to the right viewpoint of the stereo pair.

[0157]

[0174] Otherwise, if the syntax element num_sample_shift_plus512[i] is equal to 512, no shift operation may be applied.

[0158]

[0175] Otherwise, if (num_sample_shift_plus512[i] - 512) is greater than 0, the left viewpoint of the stereo pair corresponding to the ith reference view and the ith reference baseline can be shifted rightward by (512 - num_sample_shift_plus512[i]) samples relative to the right viewpoint of the stereo pair.

[0159]

[0176] The value of num_sample_shift_plus512[i] can be in the range of 0 to 123 inclusive.

[0160]

[0177] In one embodiment, shifting the left viewpoint by x samples in the left (or right) direction can be performed by a two-step process: 1) shifting the left viewpoint by x / 2 samples in the left (or right) direction and shifting the right viewpoint by x / 2 samples in the right (or left) direction; 2) filling the left and right image margins of x / 2 samples within the width of both the left and right viewpoints with the background color.

[0161]

[0178] The following pseudocode shows an example of a shift operation where the left eye is shifted left by x samples relative to the right eye:

[0162]

number

[0179] The following pseudocode shows an example of a shift operation where the left eye is shifted right by x samples relative to the right eye:

[0163]

number

[0180] The variable backgroundColour may take on different values ​​in different systems, for example black or grey.

[0164]

[0181] The syntax element three_dimensional_reference_displays_extension_flag equal to 0 may indicate that no additional data follows in the third SEI message (e.g., the 3D reference displays SEI message). In one example, the value of three_dimensional_reference_displays_extension_flag is equal to 0 in a bitstream that conforms to some standard. A value of 1 for three_dimensional_reference_displays_extension_flag is reserved for future use by ITU-T|ISO / IEC. A decoder may ignore all data following a 1 for three_dimensional_reference_displays_extension_flag in the third SEI message (e.g., the 3D reference displays SEI message).

[0165]

[0182] FIG. 13 shows a flow chart outlining an encoding process (1300) according to an embodiment of the present disclosure. In various embodiments, the process (1300) is performed by processing circuitry, such as the processing circuitry of terminal devices (310), (320), (330), and (340), or processing circuitry performing the functions of a video encoder (e.g., (403), (603), (703)). In some embodiments, the process (1300) is implemented in software instructions, and thus, the processing circuitry performs the process (1300) when it executes the software instructions. The process begins and proceeds at (S1310).

[0166]

[0183] In (S1310), pictures associated with a view in the bitstream (e.g., ViewId[0]-ViewId[3] in Figures 10A-C) can be coded. In one example, pictures associated with a view may be coded in coding order. In one example, pictures are included in an AU of a CVS, and the pictures (and the view) are associated with the same capture time instance or display time instance.

[0167]

[0184] At 1320, a first capture enhancement information (SEI) message including scalability dimension information (SDI) and a second SEI message including multi-view viewpoint position (MVP) information may be formed. In one example, the first SEI message is denoted as an SDI SEI message and the second SEI message is denoted as an MVP SEI message, as described in this disclosure.

[0168]

[0185] The first SEI message (e.g., an SDI SEI message) may indicate the number of views for the CVS. The first SEI message may specify a view ID of the ith layer in the CVS, where the value i is an integer. The value i may be greater than or equal to 0. In one example, the value i is an integer from 0 to num_views_minus1, as described with reference to FIG. 11. In one example, the syntax element sdi_view_id_val[i] specifies the value of the view ID (e.g., ViewID) of the ith layer in the CVS. The view IDs of each layer in the CVS may be referenced by a second SEI message (e.g., an MVP SEI message).

[0169]

[0186] As mentioned above, the coding order (e.g., encoding order or decoding order) may differ from the display order used to display pictures associated with a view. The second SEI message may indicate the display order. As described in the disclosure, in one example, the second SEI message indicates the view position (e.g., view_position[i]) of the view in one dimension (e.g., the horizontal axis), where the view may be specified by ViewIds in the first SEI message.

[0170]

[0187] In one example, the second SEI message is associated with an IRAP AU in the CVS. In one example, the second SEI message is included in the CVS based on the first SEI message included in the CVS.

[0171]

[0188] In (S1330), the first SEI message and the second SEI message can be included in a bitstream. In one example, the first SEI message and the second SEI message are encoded and transmitted. A viewpoint ID of each layer corresponding to a viewpoint in the CVS can be signaled. In one example, each layer corresponds to a viewpoint. A viewpoint position of a viewpoint in one dimension can be signaled.

[0172]

[0189] The process (1300) proceeds to (S1399) and ends.

[0173]

[0190] The process 1300 can be adapted appropriately for various scenarios, and the steps of the process 1300 can be adjusted accordingly. One or more steps of the process 1300 can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process 1300. Additional steps can be added.

[0174]

[0191] In one example, the CVS includes a third SEI message (e.g., the 3D reference view information SEI message shown in FIG. 12) including 3D reference view information indicating a view identifier for the left view, a view identifier for the right view, and a reference view width for the reference view. The third SEI message may indicate the number of reference views and the number of view points signaled in the third SEI message. The maximum value of the number of reference views is based on the number of view points indicated in the third SEI message.

[0175]

[0192] In one example, a first value indicating the number of views (e.g., NumViews) is derived from a first SEI message. A third value related to the number of views from a third SEI message is obtained. The third value can be different from the first value. The third value plus one (e.g., the value "1") can be compared to the first value in a compatibility check.

[0176]

[0193] FIG. 14 shows a flow chart outlining a decoding process (1400) according to an embodiment of the present disclosure. In various embodiments, the process (1400) is performed by a processing circuit, such as the processing circuitry of the terminal devices (310), (320), (330), and (340), the processing circuitry performing the functions of the video encoder (403), the processing circuitry performing the functions of the video decoder (410), the processing circuitry performing the functions of the video decoder (510), the processing circuitry performing the functions of the video encoder (603), etc. In some embodiments, the process (1400) is implemented in software instructions, and thus, the processing circuitry performs the process (1400) when it executes the software instructions. The process begins at (S1401) and proceeds to (S1410).

[0177]

[0194] At (S1410), pictures associated with a view point of each layer in a coded video sequence (CVS) can be decoded from the bitstream. In one example, the pictures associated with a view point can be decoded in decoding order. In one example, the pictures are included in an AU of the CVS, and the pictures (and the view point) are associated with the same capture time instance or display time instance.

[0178]

[0195] At (S1420), a position of a viewpoint to be displayed in one dimension (e.g., along a horizontal axis) may be determined based on a first supplemental enhancement information (SEI) message and a second SEI message. The first SEI message may include scalability dimension information (SDI). The second SEI message may include multi-view viewpoint position (MVP) information.

[0179]

[0196] For example, the first SEI message includes an SDI SEI message and the second SEI message includes an MVP SEI message.

[0180]

[0197] The first SEI message (e.g., an SDI SEI message) may indicate the number of views for the CVS. The first SEI message may specify a view ID (e.g., ViewID[i]) of the ith layer in the CVS, where the value i is an integer. The value i may be equal to or greater than 0. In one example, the value i is an integer from 0 to num_views_minus1, as described with reference to FIG. 11 . In one example, a syntax element sdi_view_id_val[i] in the first SEI message specifies the value of the view ID (e.g., ViewID) of the ith layer in the CVS. The view IDs of each layer in the CVS may be referenced by a second SEI message (e.g., an MVP SEI message).

[0181]

[0198] The second SEI message (e.g., MVP SEI message) may indicate the position of the viewpoint in one dimension (e.g., view_position[i]), and the viewpoint may be specified by the ViewIds in the first SEI message. In one example, the one dimension is the horizontal axis, and the position of the viewpoint to be displayed is along the horizontal axis.

[0182]

[0199] In an embodiment, for each of the viewpoints in each layer of the CVS, a viewpoint ID (e.g., ViewId[i]) of the respective viewpoint is determined based on the first SEI message. For example, a syntax element sdi_view_id_val[i] in the first SEI message specifies a viewpoint ID value (e.g., denoted as ViewID). The position of each viewpoint to be displayed in one dimension (e.g., along the horizontal axis) can be determined based on the viewpoint position variable (e.g., view_position[i]) in the second SEI message and the determined viewpoint ID (e.g., ViewId[i]).

[0183]

[0200] In an embodiment, the second SEI message is associated with an intra random access point (IRAP) AU of the CVS. In one example, the viewpoint position applies to the AU of the CVS.

[0184]

[0201] At (S1410), the decoded picture can be displayed based on the determined position of the viewpoint in one dimension (e.g., along the horizontal axis), as described with reference to Figures 10A-10C.

[0185]

[0202] The process (1400) proceeds to (S1499) and ends.

[0186]

[0203] The process 1400 can be adapted appropriately for various scenarios, and the steps of the process 1400 can be adjusted accordingly. One or more steps of the process 1400 can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement the process 1400. Additional steps can be added.

[0187]

[0204] In an embodiment, whether the second SEI message is included in the CVS is based on whether the first SEI message is included in the CVS, and the second SEI message may be present in the CVS if the first SEI message is present in the CVS.

[0188]

[0205] If the CVS does not include the first SEI message (e.g., an SDI SEI message), the CVS does not include the second SEI message (e.g., an MVP SEI message). If the first SEI message is not included in the CVS, (S1420) is omitted. (S1430) can be configured to display the decoded pictures based on the decoding order, a default order, etc.

[0189]

[0206] In an embodiment, a first value (e.g., NumViews) specifying the number of viewpoints is derived from a first SEI message. A second value (e.g., num_views_minus1) associated with the number of viewpoints can be obtained from a second SEI message. The second value can be different from the first value. A compatibility check can be performed by comparing 1 plus the second value with the first value.

[0190]

[0207] In one example, the CVS includes a third SEI message (e.g., the 3D reference view information SEI message shown in FIG. 12 ) including 3D reference view information indicating a view identifier and a view distance of a left view, a view identifier of a right view, and a reference view width of a reference view. The third SEI message may indicate the number of reference views and the number of view points signaled in the third SEI message. In one example, a first value (e.g., NumViews) indicating the number of view points is derived from the first SEI message. A third value (e.g., num_views_minus1) related to the number of view points from the third SEI message is obtained. The third value may be different from the first value. One (e.g., the value “1”) plus the third value may be compared with the first value in a compatibility check. The maximum value of the number of reference views is based on the third value related to the number of view points in the third SEI message.

[0191]

[0208] The embodiments of the present disclosure can be used separately or in combination in any order. Furthermore, each method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors, or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0192]

[0209] The techniques described above may be implemented as computer software using computer-readable instructions and may be physically stored on one or more computer-readable media. For example, Figure 15 illustrates a computer system (1500) suitable for implementing certain embodiments of the disclosed subject matter.

[0193]

[0210] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that contains instructions that may be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that may go through interpretation, microcode execution, etc.

[0194]

[0211] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0195]

[0212] 15 with respect to computer system (1500) are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (1500).

[0196]

[0213] The computer system (1500) may include certain human interface input devices that may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), auditory input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that do not necessarily involve direct human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still-image cameras), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic pictures).

[0197]

[0214] The input human interface devices may include one or more of (only one of each is depicted) a keyboard (1501), a mouse (1502), a trackpad (1503), a touch screen (1510), a data glove (not shown), a joystick (1505), a microphone (1506), a scanner (1507), and a camera (1508).

[0198]

[0215] The computer system (1500) may also include certain human interface output devices that may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices can include haptic output devices (e.g., haptic feedback via a touch screen (1510), data gloves (not shown), joystick (1505), although there can be haptic feedback devices that do not serve as input), auditory output devices (e.g., speakers (1509), headphones (not shown)), visual output devices (e.g., screens (1510), including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output, three-dimensional or higher output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0199]

[0216] The computer system (1500) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1520) using media such as CD / DVD (1521), thumb drives (1522), removable hard drives or solid state drives (1523), legacy magnetic media (not shown) such as tape and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0200]

[0217] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitional signals.

[0201]

[0218] The computer system 1500 may also include a network interface 1554 to one or more communication networks 1555. The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial TV), vehicular, and industrial networks including CANBus, etc. Particular networks typically require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (1549) (e.g., a USB port on the computer system (1500)); others are typically integrated into the core of the computer system (1500) by attaching to a system bus, as described below (e.g., an Ethernet interface is integrated in a PC computer system, and a cellular network interface is integrated in a smartphone computer system). Using any of these networks, the computer system (1500) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast television), one-way transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional, such as with other computer systems using local or wide-area digital networks. Specific protocols and protocol stacks, as described above, can be used with each of these networks and network interfaces.

[0202]

[0219] The aforementioned human interface devices, human accessible storage devices, and network interfaces may be attached to a core (1540) of the computer system (1500).

[0203]

[0220] A core (1540) may include one or more central processing units (CPUs) (1541), graphics processing units (GPUs) (1542), specialized programmable processing devices in the form of field programmable gate arrays (FPGAs) (1543), task-specific hardware accelerators (1544), graphics adapters (1550), etc. These devices, along with read-only memory (ROM) (1545), random access memory (1546), and internal mass storage devices (e.g., internal non-user-accessible hard drives, SSDs, etc.) (1547), may be connected via a system bus (1548). In some computer systems, the system bus (1548) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1548) or via a peripheral bus (1549). In one example, a screen (1510) can be connected to a graphics adapter (1550). Peripheral bus architectures include PCI, USB, etc.

[0204]

[0221] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) may combine to execute specific instructions that may constitute the aforementioned computer code. The computer code may be stored in ROM (1545) or RAM (1546). Temporary data may be stored in RAM (1546), while persistent data may be stored in, for example, internal mass storage (1547). Rapid storage and retrieval from any memory device may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (1541), GPU (1542), mass storage (1547), ROM (1545), RAM (1546), etc.

[0205]

[0222] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those having skill in the computer software arts.

[0206]

[0223] By way of example, and not limitation, a computer system having the architecture (1500), specifically the core (1540), may provide functionality by executing software embodied in one or more tangible computer-readable media as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such computer-readable media may be media associated with user-accessible mass storage as described above, as well as specific storage of the core (1540) that is non-transitory in nature, such as the core's internal mass storage (1547) or ROM (1545). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (1540). The computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause the core (1540), specifically the processor (including a CPU, GPU, FPGA, etc.) therein, to perform specific processes or portions of specific processes described herein, including defining data structures stored in RAM (1546) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1544)), which may execute in place of or in conjunction with software to perform the particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0207]

[0224] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise many systems and methods that, while not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.

[0208]

[0225] Additional notes (Appendix 1) 1. A video decoding method in a video decoder, comprising: decoding from the bitstream pictures associated with each layer viewpoint in the coded video sequence (CVS); determining the viewpoint position to be displayed in one dimension based on a first supplemental enhancement information (SEI) message and a second SEI message, wherein the first SEI message includes scalability dimension information (SDI) and the second SEI message includes multi-view viewpoint position (MVP) information; a display step of displaying the decoded picture based on the viewpoint position in the one dimension; A method comprising:

[0209] (Appendix 2) 10. The method of claim 1, wherein the determining step determines, for each of the respective layer perspectives of the CVS: determining a view identifier for each of the views based on the first SEI message; and determining respective viewpoint positions in the one dimension to be displayed based on viewpoint position variables in the second SEI message and the determined viewpoint identifier.

[0210] (Appendix 3) The method according to Appendix 1, further comprising: deriving a first value indicative of the number of viewpoints from the first SEI message; obtaining a second value related to the number of viewpoints from the second SEI message, the second value being different from the first value; comparing the first value with 1 plus the second value in a compatibility test; A method comprising:

[0211] (Appendix 4) 2. The method of claim 1, wherein the second SEI message is associated with an Intra-Random Access Point (IRAP) access unit of the CVS.

[0212] (Appendix 5) 2. The method of claim 1, wherein the viewpoint position is applied to an access unit of the CVS.

[0213] (Appendix 6) 2. The method of claim 1, wherein the second SEI message is included in the CVS based on the first SEI message being included in the CVS; The method, wherein the determining and displaying steps are performed based on the first SEI message and the second SEI message being included in the CVS.

[0214] (Appendix 7) 13. The method according to claim 1, wherein the CVS includes a third SEI message including 3D reference view information, the 3D reference view information indicating a viewpoint identifier of a left view point, a viewpoint identifier of a right view point, and a reference view width of a reference view, the method further comprising: deriving a first value indicative of the number of viewpoints from the first SEI message; obtaining a third value related to the number of viewpoints from the third SEI message, the third value being different from the first value; comparing the first value with 1 plus the third value in a compatibility test; A method comprising:

[0215] (Appendix 8) In the method according to Appendix 7, the third SEI message indicates the number of reference indications signaled in the third SEI message; The method, wherein the maximum number of reference indications is based on the third value related to the number of viewpoints in the third SEI message.

[0216] (Appendix 9) 2. The method of claim 1, wherein the first SEI message is an SDI SEI message and the second SEI message is an MVP SEI message.

[0217] (Appendix 10) 1. A video decoding apparatus comprising a processing circuit, the processing circuit comprising: decoding, from the bitstream, pictures associated with each layer viewpoint in the coded video sequence (CVS); determining the viewpoint position to be displayed in one dimension based on a first supplemental enhancement information (SEI) message and a second SEI message, where the first SEI message includes scalability dimension information (SDI) and the second SEI message includes multi-view viewpoint position (MVP) information; and displaying the decoded picture based on the viewpoint position in the one dimension.

[0218] (Appendix 11) 11. The apparatus of claim 10, wherein the processing circuitry, for each of the respective layer viewpoints of the CVS: determining a view identifier for each of the views based on the first SEI message; and and determining, based on a viewpoint position variable in the second SEI message and the determined viewpoint identifier, each viewpoint position in the one dimension to be displayed.

[0219] (Appendix 12) 11. The apparatus of claim 10, wherein the processing circuitry: deriving a first value indicative of the number of viewpoints from the first SEI message; obtaining a second value related to the number of viewpoints from the second SEI message, the second value being different from the first value; comparing the first value to one plus the second value in a compatibility test.

[0220] (Appendix 13) 11. The apparatus of claim 10, wherein the second SEI message is associated with an Intra Random Access Point (IRAP) access unit of the CVS.

[0221] (Appendix 14) 11. The apparatus of claim 10, wherein the viewpoint position is applied to an access unit of the CVS.

[0222] (Appendix 15) 11. The apparatus of claim 10, wherein the second SEI message is included in the CVS based on the first SEI message being included in the CVS; The apparatus, wherein the processing circuitry is configured to perform the determining and the displaying based on the first SEI message and the second SEI message being included in the CVS.

[0223] (Appendix 16) 11. The apparatus of claim 10, wherein the CVS includes a third SEI message including 3D reference view information, the 3D reference view information indicating a viewpoint identifier of a left view point, a viewpoint identifier of a right view point, and a reference view width of a reference view, and the processing circuitry: deriving a first value indicative of the number of viewpoints from the first SEI message; obtaining a third value related to the number of viewpoints from the third SEI message, the third value being different from the first value; comparing the first value to one plus the third value in a compatibility test.

[0224] (Appendix 17) 17. The apparatus of claim 16, the third SEI message indicates the number of reference indications signaled in the third SEI message; The apparatus, wherein the maximum value of the number of reference indications is based on the third value related to the number of viewpoints in the third SEI message.

[0225] (Appendix 18) 11. The apparatus of claim 10, wherein the first SEI message is an SDI SEI message and the second SEI message is an MVP SEI message.

[0226] (Appendix 19) A non-transitory computer-readable storage medium storing a program executable by at least one processor, the program comprising: decoding, from the bitstream, pictures associated with each layer viewpoint in the coded video sequence (CVS); determining the viewpoint position to be displayed in one dimension based on a first supplemental enhancement information (SEI) message and a second SEI message, where the first SEI message includes scalability dimension information (SDI) and the second SEI message includes multi-view viewpoint position (MVP) information; and displaying the decoded picture based on the viewpoint position in the one dimension.

[0227] (Appendix 20) 20. The non-transitory computer-readable storage medium of claim 19, wherein the program executable by the at least one processor performs, for each of the respective layer perspectives of the CVS: determining a view identifier for each of the views based on the first SEI message; and determining each viewpoint position to be displayed in the one dimension based on a viewpoint position variable in the second SEI message and the determined viewpoint identifier.

[0228]

[0226] Appendix A: Acronyms [Explanation of symbols]

[0229] JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOPs: Groups of Pictures TUs: Transform Units PUs: Prediction Units CTUs: Coding Tree Units CTBs: Coding Tree Blocks PBs: Prediction Blocks HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPUs: Central Processing Units GPUs: Graphics Processing Units CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Array SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit RD: Rate-Distortion

Claims

1. 1. A video encoding method in a video encoder, comprising: encoding pictures associated with each layer's viewpoint in a coded video sequence (CVS); a determining step of forming a first supplemental enhancement information (SEI) message and a second SEI message to determine a viewpoint position of the viewpoint to be displayed in one dimension, wherein the first SEI message includes scalability dimension information (SDI), the SDI including information indicating the number of viewpoints, the first SEI message including a viewpoint identifier for each layer, and the second SEI message includes multi-view viewpoint position (MVP) information, the MVP indicating a viewpoint position in a display order along the one dimension; encoding the first SEI message and the second SEI message; and when displaying decoded pictures based on the viewpoint positions in the one dimension, if a decoding order of the pictures differs from a display order of the pictures, the determining step includes a step of determining, for each of the view points of the respective layers of the CVS, a respective viewpoint position to be displayed in the one dimension from a view point identifier corresponding to the view point of the respective layer, the view point identifier being included in the first SEI message, and a view point position variable in the second SEI message, so that the decoded pictures can be rearranged according to the display order.

2. 1. A video encoding method in a video encoder, comprising: encoding pictures associated with each layer's viewpoint in a coded video sequence (CVS); a determining step of forming a first supplemental enhancement information (SEI) message and a second SEI message to determine a viewpoint position of the viewpoint to be displayed in one dimension, wherein the first SEI message includes scalability dimension information (SDI), the SDI including information indicating the number of viewpoints, the first SEI message including a viewpoint identifier for each layer, and the second SEI message includes multi-view viewpoint position (MVP) information, the MVP indicating a viewpoint position in a display order along the one dimension; generating an encoded bitstream including the picture, the first SEI message, and the second SEI message, and transmitting the bitstream from the video encoder; and when displaying decoded pictures based on the viewpoint positions in the one dimension, if a decoding order of the pictures differs from a display order of the pictures, the determining step includes a step of determining, for each of the view points of the respective layers of the CVS, a respective viewpoint position to be displayed in the one dimension from a view point identifier corresponding to the view point of the respective layer, the view point identifier being included in the first SEI message, and a view point position variable in the second SEI message, so that the decoded pictures can be rearranged according to the display order.

3. 2. The method of claim 1, wherein the second SEI message is associated with an Intra Random Access Point (IRAP) access unit of the CVS.

4. The method of claim 1 , wherein the viewpoint position is applied to an access unit of the CVS.

5. The method of claim 1 , wherein the second SEI message is included in the CVS based on the first SEI message being included in the CVS.

6. 10. The method of claim 1, wherein the CVS includes a third SEI message including 3D reference view information, the 3D reference view information indicating a viewpoint identifier for a left view, a viewpoint identifier for a right view, and a reference view width for a reference view.

7. 7. The method of claim 6, the third SEI message indicates the number of reference indications signaled in the third SEI message; The method, wherein the maximum number of reference indications is based on a third value related to the number of viewpoints in the third SEI message.

8. 2. The method of claim 1, wherein the first SEI message is an SDI SEI message and the second SEI message is an MVP SEI message.

9. A video encoding device comprising a processing circuit configured to perform the method of any one of claims 1 to 8.

10. A computer program product which, when executed by at least one processor, causes the computer to perform the method according to any one of claims 1 to 8.

11. A storage medium storing the computer program according to claim 10.