Coding and decoding method based on visual volume video, encoder and decoder
By developing a visual volumetric video encoding and decoding method and system, the complexity and applicability issues of immersive video coding in existing technologies are solved, enabling more efficient volumetric video reconstruction and transmission.
Patent Information
- Application Number
- CN202480025142.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-11
- Filing Date
- 2024-04-10
- Publication Date
- 2025-11-14
AI Technical Summary
Existing MPEG immersive video coding methods cannot effectively handle a wide range of immersive video input data and are complex, making them difficult to use in various applications.
A method and system for encoding and decoding visual volumetric video is provided, which reconstructs volumetric video content by processing decoding and encoding syntax elements, including setting default values and enabling/disabling extended syntax elements.
It improves versatility for various applications and encoding/decoding efficiency, and is suitable for the reconstruction and transmission of large-format videos.
Smart Images

Figure CN120958792A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 458,644, filed April 11, 2023. The entire contents of the aforementioned patent application are incorporated herein by reference and form part of this specification. Technical Field
[0003] This disclosure generally relates to computer-implemented methods and systems for video processing. Specifically, this disclosure relates to Visual Volumetric Video-based Coding (V3C). Background Technology
[0004] In recent years, digital video technology has made progress in many areas. With the emergence of three-dimensional (3D), augmented reality (AR), and virtual reality (VR) applications, immersive video content is gaining widespread acceptance. For these applications, volumetric content is often desired, and sometimes even necessary. For example, the Moving Picture Experts Group (MPEG) Immersive Video (MIV) format has been proposed for storing and transmitting volumetric video content with a Visual Volumetric Video Coding (V3C) format. For instance, a V3C video stream typically includes texture data, depth data, and metadata. Texture and depth data images are stored as tiles in one or more atlases, which are two-dimensional (2D) placeholders of predetermined sizes used to store the tiles.
[0005] MPEG Immersive Video (MIV) can include multi-viewport video. To efficiently compress such immersive video, one or more base viewport videos are selected. For the remaining viewport videos, redundancy between the remaining viewport videos and the base videos is first removed, retaining only the non-overlapping portions. The base viewport videos and the non-overlapping viewport videos are then stitched together to form a larger spliced video. This spliced video and its corresponding information are encoded and decoded by existing video codecs and other encoding / decoding methods, respectively.
[0006] Unlike traditional video, volumetric video consists of a series of frames, each a 3D representation of a real-world object or scene captured at a specific moment. The MPEG Visual Volumetric Video Coding (V3C) standard defines the general mechanisms for encoding and streaming volumetric content. The two main codecs associated with the MPEG V3C standard are V-PCC for point cloud data transmission and MPEG Immersive Video (MIV) for multi-view content with depth.
[0007] However, existing MIVs are not well-suited for a wide range of immersive video input data and also present additional complexities. The aim is to design a universal MIV system and methodology that can be used in various applications. Summary of the Invention
[0008] Embodiments of this disclosure provide a visual volumetric video encoding / decoding (V3C) method, encoder, and decoder that can improve MIV in a V3C system.
[0009] In a first aspect, embodiments of this disclosure provide a visual volumetric video decoding (V3C) method applied to a decoder. The method includes the following steps: acquiring a bitstream of a volumetric video; decoding a first syntax element from the bitstream, the first syntax element indicating whether occupancy information exists in a V3C sub-bitstream component corresponding to a geometric component; in response to the absence of the first syntax element in the bitstream, setting the first syntax element to a first default value to indicate whether occupancy information exists or not in the V3C sub-bitstream component corresponding to the geometric component; and decoding volumetric content from the bitstream based on the value of the first syntax element to reconstruct the volumetric video.
[0010] According to one embodiment, the method further includes the steps of: decoding a second syntax element from the bitstream, the second syntax element indicating whether there is a recommended depth of tiles derived based on tile data unit parameters rather than geometric video data; in response to the absence of the second syntax element in the bitstream, setting the second syntax element to a second default value to indicate that when the relevant syntax element is equal to zero, the depth of the derived tiles is determined by external means; and decoding volumetric content from the bitstream based on the value of the second syntax element to reconstruct the volumetric video.
[0011] According to one embodiment, the method further includes the steps of: decoding a third syntax element from the bitstream, the third syntax element indicating whether a relevant syntax element exists in the multi-view video extended syntax structure; setting the syntax element to a third default value in response to the absence of a second syntax element in the bitstream to indicate the presence of a relevant syntax element in the multi-view video extended syntax structure; and decoding volumetric content from the bitstream based on the value of the third syntax element to reconstruct the volumetric video.
[0012] According to one embodiment, the method further includes the steps of: decoding a fourth syntax element from the bitstream, the fourth syntax element indicating the number of bits used to represent the relevant syntax element; setting the fourth syntax element to a fourth default value in response to the absence of a second syntax element in the bitstream, to indicate that one bit is used to represent the relevant syntax element; and decoding volumetric content from the bitstream based on the value of the fourth syntax element to reconstruct the volumetric video.
[0013] According to one embodiment, the method further includes the steps of: decoding a fifth syntax element from the bitstream, the fifth syntax element indicating whether an extended syntax structure associated with volumetric content of at least one specific format exists in the bitstream; decoding a sixth syntax element from the bitstream, the sixth syntax element indicating whether a multi-view video extended syntax element exists, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled; the multi-view video extended syntax element is decoded from the bitstream in response to the sixth syntax element being enabled; and the volumetric content is decoded from the bitstream using the multi-view video extended syntax element to reconstruct the volumetric video.
[0014] According to one embodiment, the method further includes the steps of: decoding a seventh syntax element from the bitstream, the seventh syntax element indicating whether a video-based point cloud compression (V-PCC) extension syntax element for a MIV profile exists; and in response to the seventh syntax element being disabled, decoding a multi-view video extension syntax element from the bitstream, and confirming that no V-PCC extension syntax element for a MIV profile exists.
[0015] According to one embodiment, the method further includes the steps of: decoding an eighth syntax element from the bitstream, the eighth syntax element indicating the presence of a seventh syntax element, wherein the seventh syntax element indicates the presence of a V-PCC extended syntax element for a MIV profile; and in response to the eighth syntax element being disabled, decoding a multi-view video extended syntax element from the bitstream, and confirming that no V-PCC extended syntax element for a MIV profile exists.
[0016] In a second aspect, embodiments of this disclosure provide a decoder. The decoder includes a communication interface, a storage device, and a processor. The communication interface is configured to acquire a bitstream of volumetric video. The storage device is configured to store the bitstream of volumetric video. The processor is coupled to the communication interface and the storage device and is configured to: decode a first syntax element from the bitstream, the first syntax element indicating whether occupancy information exists in a V3C sub-bitstream component corresponding to a geometric component; in response to the absence of the first syntax element in the bitstream, set the first syntax element to a first default value to indicate whether occupancy information exists or not in the V3C sub-bitstream component corresponding to the geometric component; and decode volumetric content from the bitstream based on the value of the first syntax element to reconstruct the volumetric video.
[0017] According to one embodiment, the processor is further configured to: decode a second syntax element from the bitstream, the second syntax element indicating whether there is a recommended depth of tiles derived based on tile data unit parameters rather than geometric video data; in response to the absence of the second syntax element in the bitstream, set the second syntax element to a second default value to indicate that the depth of the derived tiles is determined by external means when the relevant syntax element is equal to zero; and decode volumetric content from the bitstream based on the value of the second syntax element to reconstruct volumetric video.
[0018] According to one embodiment, the processor is further configured to: decode a third syntax element from the bitstream, the third syntax element indicating whether a relevant syntax element exists in the multiview video extended syntax structure; in response to the absence of a second syntax element in the bitstream, set the third syntax element to a third default value to indicate the presence of the relevant syntax element in the multiview video extended syntax structure; and decode volumetric content from the bitstream based on the value of the third syntax element to reconstruct the volumetric video.
[0019] According to one embodiment, the processor is further configured to: decode a fourth syntax element from the bitstream, the fourth syntax element indicating the number of bits used to represent the relevant syntax element; set the fourth syntax element to a fourth default value in response to the absence of a second syntax element in the bitstream, to indicate that one bit is used to represent the relevant syntax element; and decode volumetric content from the bitstream based on the value of the fourth syntax element to reconstruct the volumetric video.
[0020] According to one embodiment, the processor is further configured to: decode a fifth syntax element from the bitstream, the fifth syntax element indicating whether an extended syntax structure associated with volumetric content of at least one specific format exists in the bitstream; decode a sixth syntax element from the bitstream, the sixth syntax element indicating whether a multi-view video extended syntax element exists, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled; the multi-view video extended syntax element is decoded from the bitstream in response to the sixth syntax element being enabled; and the volumetric content is decoded from the bitstream using the multi-view video extended syntax element to reconstruct the volumetric video.
[0021] According to one embodiment, the processor is further configured to: decode a seventh syntax element from the bitstream, the seventh syntax element indicating whether a video-based point cloud compression (V-PCC) extension syntax element for the MIV profile exists; and in response to the seventh syntax element being disabled, decode a multi-view video extension syntax element from the bitstream, where no V-PCC extension syntax element for the MIV profile exists.
[0022] According to one embodiment, the processor is further configured to: decode an eighth syntax element from the bitstream, the eighth syntax element indicating the presence of a seventh syntax element, wherein the seventh syntax element indicates the presence of a V-PCC extended syntax element for a MIV profile; and, in response to the eighth syntax element being disabled, decode a multi-view video extended syntax element from the bitstream, where no V-PCC extended syntax element for a MIV profile exists.
[0023] In a third aspect, embodiments of this disclosure provide a visual volumetric video encoding (V3C) method applied to an encoder. The method includes the following steps: acquiring volumetric video data; determining, based on the volumetric video data, whether occupancy information exists in the V3C sub-stream components corresponding to geometric components to encode a first syntax element; and generating a volumetric video bitstream, wherein the volumetric video bitstream contains, or does not contain, a first syntax element indicating whether occupancy information exists in the V3C sub-stream components corresponding to geometric components.
[0024] According to one embodiment, the method further includes the step of: determining, based on the volumetric video data, whether there is a recommended depth of tiles derived from tile data unit parameters rather than geometric video data, to encode a second syntax element, wherein the presence or absence of a second syntax element in the bitstream of the volumetric video indicates whether there is a recommended depth of tiles derived from tile data unit parameters rather than geometric video data.
[0025] According to one embodiment, the method further includes the following steps: determining whether a relevant syntax element exists in the multi-view video extended syntax structure based on the volumetric video data, so as to encode a third syntax element, wherein the presence or absence of a third syntax element in the bitstream of the volumetric video indicates whether a relevant syntax element exists in the multi-view video extended syntax structure.
[0026] According to one embodiment, the method further includes the step of: determining the number of bits used to represent the relevant syntax element based on data from the volumetric video to encode the fourth syntax element, wherein the bitstream of the volumetric video contains or does not contain a fourth syntax element indicating the number of bits used to represent the relevant syntax element.
[0027] According to one embodiment, the method further includes the steps of: encoding a fifth syntax element indicating the presence of an extended syntax structure associated with volume content of at least one particular format into the bitstream; and encoding a sixth syntax element indicating the presence of a multi-view video extended syntax element in response to the fifth syntax being enabled, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled.
[0028] According to one embodiment, the method further includes the step of encoding whether a seventh syntax element is disabled into the bitstream, wherein the seventh syntax element is used to indicate whether a video-based point cloud compression (V-PCC) extension syntax element for MIV profiles is present.
[0029] According to one embodiment, the method further includes the step of encoding whether an eighth syntax element is disabled into the bitstream, wherein the eighth syntax element is used to indicate whether a seventh syntax element exists, and the seventh syntax element indicates whether a V-PCC extension syntax element for the MIV profile exists.
[0030] In a fourth aspect, embodiments of this disclosure provide an encoder. The encoder includes a communication interface, a storage device, and a processor. The communication interface is configured to acquire volumetric video data. The storage device is configured to store the volumetric video data. The processor is coupled to the communication interface and the storage device and is configured to: determine, based on the volumetric video data, whether occupancy information exists in a V3C sub-stream component corresponding to a geometric component to encode a first syntax element; and generate a volumetric video bitstream, wherein the volumetric video bitstream contains, or does not contain, a first syntax element indicating whether occupancy information exists in a V3C sub-stream component corresponding to a geometric component.
[0031] According to one embodiment, the processor is further configured to: determine, based on the volumetric video data, whether there is a recommended depth of tiles derived from tile data unit parameters rather than geometric video data, to encode a second syntax element, wherein the presence or absence of a second syntax element in the bitstream of the volumetric video indicates whether there is a recommended depth of tiles derived from tile data unit parameters rather than geometric video data.
[0032] According to one embodiment, the processor is further configured to: determine, based on volumetric video data, whether a relevant syntax element exists in the multiview video extended syntax structure, to encode a third syntax element, wherein the presence or absence of a third syntax element in the bitstream of the volumetric video indicates whether a relevant syntax element exists in the multiview video extended syntax structure.
[0033] According to one embodiment, the processor is further configured to: determine the number of bits used to represent the relevant syntax element based on data from the volumetric video to encode the fourth syntax element, wherein the bitstream of the volumetric video contains or does not contain a fourth syntax element indicating the number of bits used to represent the relevant syntax element.
[0034] According to one embodiment, the processor is further configured to: encode a fifth syntax element indicating the presence of an extended syntax structure associated with volumetric content of at least one particular format into the bitstream; and, in response to the fifth syntax being enabled, encode a sixth syntax element indicating the presence of a multi-view video extended syntax element, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled.
[0035] According to one embodiment, the processor is further configured to encode whether a seventh syntax element is disabled into the bitstream, wherein the seventh syntax element is used to indicate whether a video-based point cloud compression (V-PCC) extension syntax element for the MIV profile exists.
[0036] According to one embodiment, the processor is further configured to encode whether an eighth syntax element is disabled into the bitstream, wherein the eighth syntax element is used to indicate whether a seventh syntax element exists, and the seventh syntax element indicates whether a V-PCC extension syntax element for the MIV profile exists.
[0037] In a fifth aspect, embodiments of this disclosure provide a non-transitory computer-readable recording medium storing a program that causes a computer to perform the following steps: decoding a first syntax element from a bitstream, the first syntax element indicating whether occupancy information exists in a V3C sub-bitstream component corresponding to a geometric component; setting the first syntax element to a first default value in response to the absence of the first syntax element in the bitstream to indicate whether occupancy information exists or not in the V3C sub-bitstream component corresponding to the geometric component; and decoding volumetric content from the bitstream based on the value of the first syntax element to reconstruct a volumetric video.
[0038] In a sixth aspect, embodiments of the present disclosure provide a non-transitory computer-readable recording medium storing a program that causes a computer to perform the following operations: determining, based on data from a volumetric video, whether occupancy information exists in a V3C sub-stream component corresponding to a geometric component, to encode a first syntax element; and generating a bitstream of the volumetric video, wherein the bitstream of the volumetric video contains, or does not contain, a first syntax element indicating whether occupancy information exists in a V3C sub-stream component corresponding to a geometric component. Attached Figure Description
[0039] The various aspects of this disclosure can be best understood by reading the following detailed description in conjunction with the accompanying drawings. Note that, in accordance with standard industry practice, the various features are not drawn to scale. In fact, for clarity of explanation, the dimensions of the various features may be arbitrarily increased or decreased.
[0040] Figure 1 This is a schematic block diagram of a video encoding and decoding system related to embodiments of this disclosure.
[0041] Figure 2A This is a schematic block diagram of a video encoder associated with embodiments of this disclosure.
[0042] Figure 2B This is a schematic block diagram of a video decoder related to embodiments of this disclosure.
[0043] Figure 3 This is a block diagram illustrating an immersive video encoder in relation to embodiments of this disclosure.
[0044] Figure 4 This is a block diagram illustrating an immersive video decoder in relation to embodiments of the present disclosure.
[0045] Figure 5 A schematic diagram of the hardware structure of an encoder provided for an embodiment of the present invention.
[0046] Figure 6This is a flowchart of a visual volumetric video encoding (V3C) method applied to an encoder according to embodiments of the present disclosure.
[0047] Figure 7 This is a schematic diagram of the hardware structure of the decoder provided in an embodiment of the present invention.
[0048] Figure 8 This is a flowchart of the V3C method applied to the decoder according to an embodiment of the present invention.
[0049] Figure 9A and Figure 9B This is a syntax table in the Visual Volumetric Video Coding (V3C) according to an embodiment of the present invention.
[0050] Figures 10A to 10C This is a syntax table in the Visual Volumetric Video Decoding (V3C) according to an embodiment of the present invention. Detailed Implementation
[0051] To gain a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The drawings are for reference and explanation purposes only and are not intended to limit the embodiments of this disclosure.
[0052] This disclosure proposes several improvements to the representation of MPEG Immersive Video (MIV). The proposed method can be used in future MIV standards. With the implementation of the proposed method, modifications to the bitstream structure, syntax, constraints, and mappings used for generating point clouds for decoding are considered for standardization.
[0053] The encoding and decoding involved in the embodiments of this disclosure mainly include video encoding and video decoding. For ease of understanding, please refer to the following references first. Figure 1 This disclosure introduces a video encoding and decoding system relating to embodiments thereof.
[0054] Figure 1 This is a schematic block diagram of a video encoding and decoding system related to embodiments of this disclosure. (See reference...) Figure 1 The video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device 110 encodes video data (which can be understood as compression) to generate a bitstream and transmits the bitstream to the decoding device 120. The decoding device 120 decodes the bitstream generated by the encoding device 110 to generate decoded video data.
[0055] The encoding device 110 in the embodiments of this disclosure can be understood as a device with video encoding capabilities, and the decoding device 120 can be understood as a device with video decoding capabilities. That is, the embodiments of this disclosure include a wider range of devices for the encoding device 110 and the decoding device 120, including but not limited to, for example, smartphones, desktop computers, mobile computing devices, laptop computers (e.g., tablet computers), tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, and in-vehicle computers.
[0056] In some embodiments, encoding device 110 may transmit encoded video data (e.g., a bitstream) to decoding device 120 via channel 130. Channel 130 may include one or more media and / or devices capable of transmitting encoded video data from encoding device 110 to decoding device 120.
[0057] In one embodiment, channel 130 includes one or more communication media that enable encoding device 110 to directly transmit encoded video data to decoding device 120 in real time. In this embodiment, encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission cables.
[0058] In another example, channel 130 includes a storage medium for storing video data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical discs, digital video discs (DVDs), flash memory, etc. In this example, decoding device 120 can obtain the encoded video data from the storage medium.
[0059] In another embodiment, channel 130 may include a storage server that can store video data encoded by encoding device 110. In this embodiment, decoding device 120 can download the stored encoded video data from the storage server. Optionally, the storage server can store encoded video data and can transmit the encoded video data to decoding device 120; the storage server may be, for example, a web server, a File Transfer Protocol (FTP) server, etc.
[0060] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. In some embodiments, the output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0061] In some embodiments, in addition to the video encoder 112 and the input interface 113, the encoding device 110 may include a video source 111.
[0062] Video source 111 may include at least one of the following: video capture device (e.g., camera); video archive; video input interface for receiving video data from video content providers; computer graphics system for generating video data.
[0063] Video encoder 112 encodes video data from video source 111 to generate a bitstream. The video data from video source 111 may include not only texture information but also depth information corresponding to the texture information; the texture information may be a three-channel color image, and the depth information may be a depth map. Furthermore, video source 111 may provide camera parameters to video encoder 112. For another view, one or more images (pictures) or image sequences (picture sequences) to be encoded can be generated based on the texture information, depth information, and camera parameters. The bitstream contains encoding information in the form of a bitstream of images or image sequences. The encoding information may include encoded image data and related data. The syntax structure refers to the set of zero or more syntax elements arranged in an indicative order within the bitstream.
[0064] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data can also be stored on a storage medium or storage server for subsequent retrieval by the decoding device 120.
[0065] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122.
[0066] In some embodiments, in addition to the input interface 121 and the video decoder 122, the decoding device 120 may also include a display device 123.
[0067] Input interface 121 includes a receiver and / or a modem. Input interface 121 can receive encoded video data through channel 130.
[0068] The video decoder 122 is used to decode the encoded video data to obtain decoded video data, and to transmit the decoded video data to the display device 123.
[0069] Display device 123 displays the decoded video data. In some embodiments, display device 123 may be integrated with decoding device 120, or it may be external to decoding device 120. Display device 123 may include various display devices, such as liquid crystal display (LCD), plasma display, organic light emitting diode (OLED) display, or other types of display devices.
[0070] Notice, Figure 1 This is merely an example; the technical solutions of the embodiments disclosed herein are not limited to those described above. Figure 1 For example, the techniques disclosed herein can also be applied to one-way video encoding or one-way video decoding.
[0071] The following describes the video coding framework involved in the embodiments of this disclosure.
[0072] Figure 2A This is a schematic block diagram of a video encoder related to embodiments of this disclosure. It should be understood that the video encoder 200 can be used to perform lossy compression on images, or to perform lossless compression on images. Lossless compression can be visual lossless compression or mathematical lossless compression, and the embodiments are not limited thereto.
[0073] The video encoder 200 can be applied to image data in luminance-chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance used to describe color and saturation.
[0074] For example, a video encoder 200 reads video data and, for each frame of the video data, divides it into several coding tree units (CTUs). In some examples, a CTU may be referred to as a "tree block," "Largest Coding Unit (LCU)," or "coding tree block (CTB)." Each CTU can be associated with a pixel block of equal size in the image. Each pixel can correspond to one luminance (luma) sample and two chrominance (chroma) samples. Therefore, each CTU can be associated with one luminance sample block and two chrominance sample blocks. The size of a CTU may be, for example, 128×128, 64×64, 32×32, etc. A CTU can also be divided into several coding units (CUs) for encoding. CUs can be rectangular or square blocks. CUs can be further divided into prediction units (PUs) and transform units (TUs), allowing encoding, prediction, and transformation to be separated and making processing more flexible. In one example, the CTU is divided into multiple CUs using a quadtree, and the CUs are further divided into multiple TUs and multiple PUs using a quadtree.
[0075] Video encoders and decoders can support various PU sizes. Assuming a specific CU size is 2N×2N, the video encoder and decoder can support PUs of 2N×2N or N×N size for intra-frame prediction, and symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.
[0076] In some embodiments, such as Figure 2A As shown, the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filtering unit 260, a decoded image buffer memory 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
[0077] Optionally, in this disclosure, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as the predicted image block or the image prediction block, and the reconstructed image block may also be referred to as the reconstruction block or the image reconstruction block.
[0078] In some embodiments, the prediction unit 210 includes an intra-frame prediction unit 211 and an inter-frame estimation and inter-frame prediction unit 212. Because there is strong correlation between adjacent pixels in a video frame, intra-frame prediction methods are used in video encoding and decoding techniques to eliminate spatial redundancy between adjacent pixels. Because there is strong similarity between adjacent frames in a video, inter-frame prediction methods are used in video encoding and decoding techniques to eliminate temporal redundancy between adjacent frames, thereby improving encoding and decoding efficiency.
[0079] Inter-frame estimation and inter-frame prediction 212 can be used for inter-frame prediction. Inter-frame prediction can include motion estimation and motion compensation, which can reference image information from different frames. Inter-frame prediction uses motion information to find a reference block from the reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The frame used for inter-frame prediction can be a P-frame and / or a B-frame, where a P-frame refers to a forward prediction frame and a B-frame refers to a bidirectional prediction frame. Inter-frame prediction uses motion information to find a reference block from the reference frame and generates a prediction block based on the reference block. Motion information includes a list of reference frames, a reference frame index, and a motion vector. The motion vector can be in all pixels or sub-pixels. If the motion vector is in a sub-pixel, interpolation filtering needs to be used in the reference frame to generate the desired sub-pixel block. Here, the block of all pixels or sub-pixels in the reference frame found based on the motion vector is called the reference block. Some techniques directly use the reference block as the prediction block, while others reprocess the reference block to generate the prediction block. The reprocessing of generating prediction blocks based on reference blocks can also be understood as using reference blocks as prediction blocks and then generating new prediction blocks based on the prediction blocks.
[0080] Intra-prediction unit 211 refers only to information from the same frame image and predicts pixel information in the currently encoded image block to eliminate spatial redundancy. The frame used in intra-prediction can be an I-frame.
[0081] Intra-frame prediction has multiple prediction modes. Taking the international digital video coding standards H-series as an example, the H.264 / AVC standard has 8 angular prediction modes and 1 non-angular prediction mode, while H.265 / HEVC extends this to 33 angular prediction modes and 2 non-angular prediction modes. HEVC uses intra-frame prediction modes including planar mode, DC mode, and 33 angular modes, for a total of 35 prediction modes. VVC uses intra-frame prediction modes including planar mode, DC mode, and 65 angular modes, for a total of 67 prediction modes.
[0082] It is worth noting that with the increase in angle modes, intra-frame prediction will be more accurate and better meet the development needs of high-definition and ultra-high-definition digital video.
[0083] The residual unit 220 can generate a residual block of the CU based on the pixel blocks of the CU and the prediction blocks of the PU of the CU. For example, the residual unit 220 can generate a residual block of the CU such that the value of each sample in the residual block is equal to the difference between the sample in the pixel block of the CU and the corresponding sample in the prediction block of the PU of the CU.
[0084] The transform / quantization unit 230 can quantize the transform coefficients. The transform / quantization unit 230 can quantize the transform coefficients associated with the TU of the CU based on the quantization parameter (QP) value associated with the CU. The video encoder 200 can adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0085] The inverse transform / quantization unit 240 can apply inverse quantization and inverse transform to the quantized transform coefficients to reconstruct the residual block from the quantized transform coefficients.
[0086] The reconstruction unit 250 can add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by the prediction unit 210 to produce a reconstructed image block associated with the TU. By reconstructing sample blocks of each TU of the CU in this way, the video encoder 200 can reconstruct the pixel blocks of the CU.
[0087] The loop filtering unit 260 is used to process pixels undergoing inverse transform and inverse quantization to compensate for distortion information and provide a better reference for subsequent encoding of pixels. For example, a deblocking filtering operation can be performed to reduce the blocking artifacts of pixel blocks associated with the CU.
[0088] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive compensation / adaptive loop filtering (SAO / ALF) unit, wherein the deblocking filtering unit is used to remove blocking effects, and the SAO / ALF unit is used to remove ringing effects.
[0089] The decoded image cache memory 270 can store reconstructed pixel blocks. Inter-frame estimation and inter-frame prediction 21211 can use a reference image containing the reconstructed pixel blocks to perform inter-frame prediction on the PUs of other images. Additionally, the intra-frame prediction unit 211 can use the reconstructed pixel blocks in the decoded image cache memory 270 to perform intra-frame prediction on other PUs in the same image as the CU.
[0090] The entropy coding unit 280 can receive quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 can perform one or more entropy coding operations on the quantized transform coefficients to generate entropy-coded data.
[0091] Figure 2B This is a schematic block diagram of a video decoder related to embodiments of this disclosure.
[0092] like Figure 2B As shown, the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filtering unit 350, and a decoded image cache memory 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.
[0093] The video decoder 300 can receive a bitstream. The entropy decoding unit 310 can parse the bitstream to extract syntax elements from it. As part of parsing the bitstream, the entropy decoding unit 310 can parse entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filtering unit 350 can decode the video data based on the syntax elements extracted from the bitstream and generate decoded video data.
[0094] In some embodiments, the prediction unit 320 includes an intra-frame prediction unit 321 and an inter-frame prediction unit 322.
[0095] Intra-prediction unit 321 can perform intra-prediction to generate prediction blocks for the PU. Intra-prediction unit 321 can use an intra-prediction mode to generate prediction blocks for the PU based on pixel blocks of spatially adjacent PUs. Intra-prediction unit 321 can also determine the intra-prediction mode of the PU based on one or more syntax elements parsed from the bitstream.
[0096] Inter-frame prediction unit 322 can construct a first reference image list (list 0) and a second reference image list (list 1) based on the syntax elements parsed from the bitstream. Additionally, if the PU uses inter-frame predictive coding, the entropy decoding unit 310 can parse the motion information of the PU. Inter-frame prediction unit 322 can determine one or more reference blocks of the PU based on the motion information. Inter-frame prediction unit 322 can generate prediction blocks for the PU based on one or more reference blocks.
[0097] The inverse quantization / transform unit 330 can inverse quantize (i.e., dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 can use the QP value associated with the CU of the TU to determine the degree of quantization.
[0098] After inverse quantization of the transform coefficients, the inverse quantization / transformation unit 330 can apply one or more inverse transforms to the inverse quantized transform coefficients to produce a residual block associated with the TU.
[0099] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 can add samples of the residual block to the corresponding samples of the prediction block to reconstruct the pixel block of the CU and obtain the reconstructed image block.
[0100] The loop filter unit 350 can perform deblocking filtering operations to reduce the blocking artifacts of pixel blocks associated with the CU.
[0101] The video decoder 300 can store the reconstructed image from the CU in the decoded image cache memory 360. The video decoder 300 can use the reconstructed image in the decoded image cache memory 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0102] The basic processes of video encoding and video decoding are as follows.
[0103] At the encoding end, the image frame is divided into multiple blocks. For the current block, prediction unit 210 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. Residual unit 220 can calculate a residual block based on the prediction block of the current block and the original block, that is, the difference between the prediction block of the current block and the original block. The residual block can also be referred to as residual information. The residual block undergoes processing such as transformation and quantization performed by transform / quantization unit 230, which can remove information that is not sensitive to the human eye and eliminate visual redundancy. Optionally, the residual block before transformation and quantization by transform / quantization unit 230 can be referred to as a temporal residual block, and the temporal residual block after transformation and quantization by transform / quantization unit 230 can be referred to as a frequency residual block or a frequency domain residual block. Entropy coding unit 280 receives the quantized change coefficients output from change quantization unit 230 and can perform entropy coding on the quantized change coefficients to output a bitstream. For example, entropy coding unit 280 can eliminate character redundancy based on the target context model and the probability information of the binary bitstream.
[0104] At the decoding end, the entropy decoding unit 310 can parse the bitstream to obtain prediction information and quantization coefficient matrix for the current block. The prediction unit 320 generates a prediction block for the current block based on the prediction information using intra-frame prediction or inter-frame prediction of the current block. The inverse quantization / transform unit 330 performs inverse quantization and inverse transform on the quantization coefficient matrix obtained from the bitstream to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed block constitutes the reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or based on the block to obtain a decoded image. The decoded image can also be referred to as the reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction of subsequent frames.
[0105] It should be noted that the block partitioning information determined at the encoding end, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, are carried in the bitstream when necessary. The decoding end determines the same block partitioning information, prediction, transform, quantization, entropy coding, and loop filtering mode information or parameter information as the encoding end by parsing the bitstream and analyzing existing information, thereby ensuring that the image encoded by the encoding end is the same as the image decoded by the decoding end.
[0106] The above describes the basic workflow of a video codec within a block-based hybrid coding framework. As technology advances, some modules or steps of this framework or workflow may be optimized. This disclosure applies to, but is not limited to, the basic workflow of a video codec within a block-based hybrid coding framework.
[0107] In some public scenarios, multiple heterogeneous contents appear simultaneously in the same 3D scene, such as multi-view video and point clouds. For multi-view video, MPEG Immersive Video (MIV) technology is used for encoding and decoding, and for point clouds, video-based point cloud compression (V-PCC) technology is used for encoding and decoding. In some embodiments, in visual volumetric video encoding and decoding (V3C), frame packing technology is used to encode and decode multi-view video and point clouds.
[0108] Figure 3 This is a block diagram illustrating an immersive video encoder in relation to embodiments of this disclosure. Reference Figure 3 The encoding apparatus includes, wholly or partially, a view optimizer 510, an atlas builder 520, a texture encoder 530, a depth encoder 540, and a metadata synthesizer 550. The encoding apparatus sequentially uses the view optimizer 510 and the atlas builder 520 to generate MPEG immersive video (MIV) format from the input multi-view video, and then uses the texture encoder 530 and the depth encoder 540 to encode the MIV format data.
[0109] The view optimizer 510 categorizes all views included in the input multi-view video into one or more base views and one or more supplementary views. For this view optimization, the view optimizer 510 calculates how many base views are needed and selects as many base views as determined. The view optimizer 510 can determine the base views and supplementary views by using physical positions (e.g., angular differences between views) and overlaps between views. Therefore, the view with the most common scene among all views can be selected as the base view. After selecting the base views and supplementary views, the base views are retained and directly input into the encoder.
[0110] Atlas builder 520 constructs atlases from base views and additional views. As described above, the base view selected by view optimizer 510 is included in the atlas as a complete image. Atlas builder 520 generates tiles from additional views representing parts that are difficult to predict based on the base views, and then constructs the tiles generated from multiple additional views into an atlas. Atlas builder 520 is configured to generate one or more atlases based on input from view optimizer 510. To generate atlases, atlas builder 520 includes a trimmer 522, an aggregator 524, and a tile packer 526.
[0111] The trimmer 522 removes overlapping portions of the additional views while preserving the base view, and generates a binary mask indicating whether there is overlap between pixels included in the additional views. In other words, the trimmer 522 is configured to identify and remove redundancy between views. For example, a mask for an additional view may have the same resolution as the additional view, with a value of "1" indicating that the depth image value at the corresponding pixel is valid, and a value of "0" indicating that pixels need to be removed to overlap with the base view. The trimmer 522 searches for overlap information by performing a warp in 3D coordinates based on depth information. Here, warp refers to the process of using depth information to predict and compensate for the displacement vector between the two views.
[0112] Aggregator 524 accumulates the masks generated for each additional view in chronological order. This accumulation of masks can reduce the construction information of the final atlas.
[0113] Tile packer 526 packs tiles from the base view and additional views to ultimately generate an atlas. When processing the texture and depth information of the base view, tile packer 526 constructs the atlas of the base view using the original image as tiles. Regarding the texture and depth information of the additional views, tile packer 526 constructs the atlas of the additional views by generating block tiles using masks and then packing the block tiles.
[0114] Texture encoder 530 encodes the texture atlas. Depth encoder 540 encodes the depth atlas. As described above, texture encoder 530 and depth encoder 540 can be implemented using existing encoders, such as High-Efficiency Video Coding (HEVC) or VVC. For example, texture encoder 530 and depth encoder 540 can use... Figure 2A The video coding framework in the video is used to encode texture atlases and depth atlases.
[0115] The metadata synthesizer 550 generates sequence parameters related to encoding, metadata from multi-view cameras, and parameters related to atlases.
[0116] The encoding device generates and transmits a bitstream obtained by multiplexing the encoded texture, the encoded depth, and the metadata.
[0117] Figure 4 This is a block diagram illustrating an immersive video decoder in relation to embodiments of this disclosure. Reference Figure 4 The immersive video decoding device includes, in whole or in part, a texture decoder 410, a depth decoder 420, a metadata analyzer 430, an atlas tile occupancy graph generator 440, and a renderer 450.
[0118] Texture decoder 410 decodes a texture atlas from the bitstream. Depth decoder 420 decodes a depth atlas from the bitstream. As described above, texture decoder 410 and depth decoder 420 can be implemented using existing decoders (e.g., High Efficiency Video Coding (HEVC) or VVC). For example, texture decoder 410 and depth decoder 420 can use... Figure 2B The video encoding framework in the video is used to decode texture atlases and depth atlases.
[0119] Metadata analyzer 430 parses metadata from the bitstream. Occupancy map generator 440 generates an occupancy map using atlas-related parameters included in the metadata. The occupancy map is information related to the position of block maps; it can be generated by the encoding device and then transmitted to the decoding device, or generated by the decoding device using the metadata.
[0120] Renderer 450 uses texture atlases, depth atlases, and occupancy maps to reconstruct the immersive video to be delivered to the user.
[0121] [Encoder]
[0122] Figure 5 This is a schematic diagram of the hardware structure of the encoder provided in an embodiment of this disclosure. (Reference) Figure 5 The encoder 30 includes a communication interface 32, a storage device 34, and a processor 36.
[0123] For example, communication interface 32 can be a network card that supports wired network connections such as Ethernet, a wireless network card that supports wireless communication standards such as IEEE 802.11n / b / g / ac / ax / be, or any other network connection device, but this embodiment is not limited to these. Communication interface 32 is configured to acquire volumetric video data.
[0124] Storage device 34 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM) used as an external cache memory. Storage device 34 described in this disclosure is configured to store volumetric video data acquired by communication interface 32. In some embodiments, storage device 34 is a non-transitory computer-readable recording medium configured to store a program that causes processor 36 to perform the visual volumetric video encoding / decoding (V3C) method as shown below.
[0125] Processor 36 is coupled to communication interface 32 and storage device 34 via bus system 38. It is understood that bus system 38 is used as a data bus to enable connection and communication between these components. In addition to a data bus, bus system 38 can also be a power bus, control bus, status signal bus, or a combination thereof, but the embodiments are not limited thereto.
[0126] Figure 6 This is a flowchart of a visual volumetric video encoding (V3C) method applied to an encoder according to embodiments of the present disclosure. Also refer to... Figure 5 and Figure 6 The method in this embodiment is applied to Figure 5 The encoder 30 in the present disclosure. The detailed steps of the V3C method of the accompanying elements in the encoder 30 will now be described below.
[0127] In step S610, the processor 36 can acquire volumetric video data using the communication interface 32. In some embodiments, the volumetric video consists of a series of frames accompanied by metadata, and each frame contains volumetric content, which is a 3D representation of a real-world object or scene captured at a certain moment. For example, the frame can be a volumetric frame. The volumetric content can be represented in the format of a point cloud or multi-view video, but is not limited thereto.
[0128] Specifically, in 3D applications (e.g., virtual reality (VR), augmented reality (AR), or mixed reality (MR)), visual volumetric content with different representation formats can appear in the same scene. For example, multiple media objects can exist in the same 3D scene. In some embodiments, the background and some objects in the 3D scene are represented as multi-view video, while some objects are represented as 3D point clouds.
[0129] In some embodiments, volumetric content includes multiple media contents presented simultaneously in the same 3D space. In some embodiments, volumetric content includes multiple media contents presented in the same 3D space at different times. In some embodiments, volumetric content includes media content in different 3D spaces. However, there are no specific limitations on the volumetric content described above in the embodiments of this disclosure.
[0130] In some embodiments, volumetric content can be represented by point clouds or multi-view video, and various point cloud extension syntax elements and multi-view video extension syntax elements can be used to encode the volumetric content.
[0131] In step S620, the processor 36 can determine whether there is occupancy information in the V3C sub-stream component corresponding to the geometric component based on the volumetric video data, so as to encode the first syntax element. For example, the first syntax element can be the syntax element "vme_embedded_occupancy_enabled_flag" or the syntax element "asme_embedded_occupancy_enabled_flag" defined in the MIV extended syntax table in the V3C standard.
[0132] In some embodiments, processor 36 may encode a specific syntax element indicating the presence of a MIV extension syntax structure in the bitstream of the volumetric video. For example, this specific syntax element may be the syntax element "asps_miv_extension_present_flag" defined in the raw byte sequence payload (RBSP) syntax table of the General Atlas Sequence Parameter Set (ASPS) in the V3C standard. Next, processor 36 may determine whether this specific syntax element is enabled, i.e., whether its value is equal to 1. In response to the specific syntax element being equal to 1, processor 36 may encode a first syntax element indicating the presence of occupancy information in the V3C sub-bitstream component corresponding to the geometric component into the bitstream. In response to the specific syntax element not being equal to 1, processor 36 may not encode the first syntax element indicating the presence of occupancy information in the V3C sub-bitstream component corresponding to the geometric component into the bitstream.
[0133] In step S630, processor 36 can generate a bitstream of volumetric video. The presence or absence of a first syntax element in the bitstream of volumetric video indicates whether occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component. In some embodiments, the first syntax element is absent in the bitstream of volumetric video in response to the bitstream not containing a corresponding extended syntax structure.
[0134] In some embodiments, a syntax element "vme_embedded_occupancy_enabled_flag" equal to 1 indicates the presence of occupancy information in the V3C sub-stream component corresponding to the geometric component. This is determined by checking whether vuh_unit_type is equal to V3C_GVD or V3C_PVD, or by external means if the V3C unit header is unavailable. A syntax element "vme_embedded_occupancy_enabled_flag" equal to 0 indicates the absence of occupancy information or the presence of occupancy information in the V3C sub-stream component corresponding to the occupancy component. This is determined by checking whether vuh_unit_type is equal to V3C_OVD or V3C_PVD, or by external means if the V3C unit header is unavailable. When processor 36 generates a bitstream that does not contain the extended syntax structure "vps_miv_extension()", the syntax element "vme_embedded_occupancy_enabled_flag" is not present in the bitstream. If the syntax element "vme_embedded_occupancy_enabled_flag" does not exist in the bitstream, it is inferred that the value of "vme_embedded_occupancy_enabled_flag" is equal to 0.
[0135] In some embodiments, the syntax element "asme_embedded_occupancy_enabled_flag" equal to 1 indicates the presence of occupancy information in the V3C sub-stream component corresponding to the geometric component. This is determined by checking whether vuh_unit_type equals V3C_GVD or V3C_PVD, or by external means if the V3C unit header is unavailable. The syntax element "asme_embedded_occupancy_enabled_flag" equal to 0 indicates the absence of occupancy information or the presence of occupancy information in the V3C sub-stream component corresponding to the occupancy component. This is determined by checking whether vuh_unit_type equals V3C_OVD or V3C_PVD, or by external means if the V3C unit header is unavailable. When processor 36 generates a bitstream that does not contain the extended syntax structure "asps_miv_extension()", the syntax element "asme_embedded_occupancy_enabled_flag" is not present in the bitstream. If the syntax element "asme_embedded_occupancy_enabled_flag" does not exist, it is inferred that the value of asme_embedded_occupancy_enabled_flag is equal to 0.
[0136] The requirement for stream consistency is that when a V3C VPS is available, the value of asme_embedded_occupancy_enabled_flag should be equal to the value of vme_embedded_occupancy_enabled_flag.
[0137] In some embodiments, the processor 36 may determine, based on the volumetric video data, whether a recommended depth for tiles derived from tile data unit parameters rather than geometric video data exists, in order to encode the second syntax element. For example, the second syntax element may be the syntax element "asme_patch_constant_depth_flag" defined in the MIV extended syntax table in the V3C standard. The second syntax element indicates whether a recommended depth for tiles derived from tile data unit parameters rather than geometric video data exists in the volumetric video bitstream. Specifically, in response to a specific syntax element indicating the presence of a MIV extended syntax structure in the volumetric video bitstream being enabled, the processor 36 may encode the second syntax element into the bitstream. In response to a specific syntax element indicating the presence of a MIV extended syntax structure in the volumetric video bitstream being disabled, the processor 36 may not encode the second syntax element into the bitstream.
[0138] In some embodiments, the syntax element "asme_patch_constant_depth_flag" being equal to 1 indicates that the recommended depth of the tile is derived based on tile data unit parameters rather than geometric video data. The syntax element "asme_patch_constant_depth_flag" being equal to 0 indicates that the depth is determined by external means when vps_geometry_video_present_flag[aspsAtlasID] is equal to 0. When processor 36 generates a bitstream that does not contain the extended syntax structure "asps_miv_extension()", the syntax element "asme_patch_constant_depth_flag" is not present in the bitstream. When the syntax element "asme_patch_constant_depth_flag" is absent, it is inferred that the value of asme_patch_constant_depth_flag is equal to 0.
[0139] In some embodiments, the processor 36 can determine whether a relevant syntax element exists in the multi-view video extended syntax structure based on the volumetric video data, in order to encode the third syntax element. For example, the third syntax element may be the syntax element "asme_patch_texture_offset_enabled_flag" defined in the MIV extended syntax table in the V3C standard. For example, the relevant syntax element may be the syntax element "asme_patch_texture_offset_bit_depth_minus1" in the syntax structure "asps_miv_extension()", or the syntax element "pdu_texture_offset[tileID][p][c]" in the syntax structure "pdu_miv_extension(tileID, p)". The presence or absence of a third syntax element in the volumetric video bitstream that indicates the presence of a relevant syntax element in the multi-view video extended syntax structure determines whether the third syntax element exists. Specifically, in response to the enabling of a specific syntax element indicating whether the MIV extended syntax structure exists in the volumetric video bitstream, the processor 36 may encode the third syntax element into the bitstream. In response to the presence of a specific syntax element of the MIV extended syntax structure in the bitstream indicating the volumetric video being disabled, the processor 36 may not encode the third syntax element into the bitstream.
[0140] In some embodiments, a syntax element "asme_patch_texture_offset_enabled_flag" equal to 1 indicates that the syntax structure "asps_miv_extension()" contains the syntax element "asme_patch_texture_offset_bit_depth_minus1", and the syntax structure "pdu_miv_extension(tileID, p)" contains the syntax element "pdu_texture_offset[tileID][p][c]". A syntax element "asme_patch_texture_offset_enabled_flag" equal to 0 indicates that the syntax structure "asps_miv_extension()" does not contain the syntax element "asme_patch_texture_offset_bit_depth_minus1", and the syntax structure "pdu_miv_extension(tileID, p)" does not contain the syntax element "pdu_texture_offset[tileID][p][c]". When processor 36 generates a bitstream that does not contain the extended syntax structure "asps_miv_extension()", the syntax element "asme_patch_texture_offset_enabled_flag" does not exist in the bitstream. If the syntax element "asme_patch_texture_offset_enabled_flag" does not exist in the bitstream, it is inferred that the value of the syntax element "asme_patch_texture_offset_enabled_flag" is equal to 0.
[0141] In some embodiments, the processor 36 may determine the number of bits used to represent the relevant syntax element based on the volumetric video data in order to encode the fourth syntax element. For example, the fourth syntax element may be the syntax element "asme_patch_texture_offset_bit_depth_minus1" defined in the MIV extended syntax table in the V3C standard. For example, the relevant syntax element may be the syntax element "pdu_texture_offset[tileID][p][c]" in the syntax structure "pdu_miv_extension(tileID, p)". The fourth syntax element indicating the number of bits used to represent the relevant syntax element may or may not exist in the volumetric video bitstream. Specifically, in response to the activation of a specific syntax element indicating the presence of a MIV extended syntax structure in the volumetric video bitstream, the processor 36 may encode the fourth syntax element into the bitstream. In response to the deactivation of a specific syntax element indicating the presence of a MIV extended syntax structure in the volumetric video bitstream, the processor 36 may not encode the fourth syntax element into the bitstream.
[0142] In some embodiments, incrementing the syntax element “asme_patch_texture_offset_bit_depth_minus1” by 1 indicates the number of bits used to represent the syntax element “pdu_texture_offset[tileID][p][c]”. For all values of attrIdx (where ai_attribute_type_id[aspsAtlasID[attrIdx] equals ATTR_TEXTURE), if the syntax element “vps_attribute_video_present_flag[aspsAtlasID]” equals 1, then the syntax element “asme_patch_texture_offset_bit_depth_minus1” should be within the range of 0 to the minimum value of ai_attribute_2d_bit_depth_minus1[aspsAtlasID][attrIdx], inclusive. For all values of attrIdx (where pin_attribute_type_id[aspsAtlasID][attrIdx] equals ATTR_TEXTURE), if the syntax element "pin_attribute_present_flag[aspsAtlasID]" equals 1, then the syntax element "asme_patch_texture_offset_bit_depth_minus1" should be within the range of 0 to the minimum value of pin_attribute_2d_bit_depth_minus1[aspsAtlasID][attrIdx], inclusive. When processor 36 generates a bitstream that does not contain the extended syntax structure "asps_miv_extension()", the syntax element "asme_patch_texture_offset_bit_depth_minus1" does not exist in the bitstream. When the syntax element "asme_patch_texture_offset_bit_depth_minus1" does not exist in the bitstream, it is inferred that the value of the syntax element "asme_patch_texture_offset_bit_depth_minus1" is equal to 0.
[0143] In some embodiments, processor 36 may encode a fifth syntax element indicating the presence of an extended syntax structure associated with volumetric content of at least one specific format into the bitstream. Processor 36 may encode a sixth syntax element indicating the presence of a multi-view video extended syntax element in response to the fifth syntax element being enabled. The sixth syntax element is enabled in response to the fifth syntax element being enabled. For example, the fifth syntax element is the syntax element "asps_extension_present_flag" defined in the general ASPS RBSP syntax table in the V3C standard. For example, the sixth syntax element is the syntax element "asps_miv_extension_present_flag" defined in the general ASPS RBSP syntax table in the V3C standard.
[0144] A syntax element "asps_extension_present_flag" equal to 1 indicates the presence of the syntax elements "asps_vpcc_extension_present_flag", "asps_miv_extension_present_flag", and "asps_extension_6bits" in the atlas_sequence_parameter_set_rbsp() syntax structure. A syntax element "asps_extension_present_flag" equal to 0 indicates the absence of the syntax elements "asps_vpcc_extension_present_flag", "asps_miv_extension_present_flag", and "asps_extension_6bits" in the bitstream. Bitstream consistency requires that when the value of the syntax element "asps_extension_present_flag" is equal to 1, the value of the syntax element "asps_miv_extension_present_flag" should also be equal to 1.
[0145] In some embodiments, the processor 36 may encode whether a seventh syntax element is disabled into the bitstream, wherein the seventh syntax element is used to indicate the presence of video-based point cloud compression (V-PCC) extension syntax elements for the MIV profile. For example, the seventh syntax element is the extension presence syntax element "aaps_vpcc_extension_present_flag" defined in the Common Atlas Frame Parameter Set (AFPS) RBSP syntax table in the V3C standard and encoded in the MIV toolset profile to indicate the presence of V-PCC extension syntax elements. Because the extension presence syntax element "aaps_vpcc_extension_present_flag" is set to 0 for the MIV toolset profile, there are no V-PCC related syntax elements for parsing the MIV bitstream.
[0146] In some embodiments, the processor is further configured to: encode whether an eighth syntax element is disabled into the bitstream, wherein the eighth syntax element is used to indicate the presence of a seventh syntax element, which indicates the presence of a V-PCC extension syntax element for the MIV profile. For example, the eighth syntax element is defined in the general AFPS RBSP syntax table in the V3C standard and encoded in the MIV toolset profile to indicate the presence of the extension presence syntax element "aaps_extension_present_flag". Because the extension presence syntax element "aaps_extension_present_flag" is set to 0 for the MIV toolset profile, the extension presence syntax element "aaps_vpcc_extension_present_flag" is also set to 0, and therefore there are no V-PCC related syntax elements for parsing the MIV bitstream.
[0147] [Decoder]
[0148] Figure 7 This is a schematic diagram of the hardware structure of the decoder provided in an embodiment of this disclosure. (See reference...) Figure 7 The decoder 50 includes a communication interface 52, a storage device 54, and a processor 56, which is coupled to the communication interface 52 and the storage device 54 via a bus system 58.
[0149] It is understood that the hardware structure of communication interface 52, storage device 54, processor 56, and bus system 58 is similar to that of communication interface 32, storage device 34, processor 36, and bus system 38, and therefore will not be described again here. In some embodiments, storage device 54 is a non-transitory computer-readable recording medium configured to store a program that causes processor 56 to execute the visual volumetric video encoding (V3C) method as shown below.
[0150] In this embodiment, the communication interface 52 is configured to acquire the bitstream of the volumetric video, and the storage device 54 is configured to store the bitstream of the volumetric video.
[0151] Figure 8 This is a flowchart of a V3C method applied to a decoder according to embodiments of this disclosure. See also... Figure 7 and Figure 8 The method in this embodiment is applied to Figure 7 The decoder 50 is described below. The detailed steps of the V3C method of an exemplary embodiment of this disclosure, accompanying the elements in the decoder 50, will now be described.
[0152] In step S810, the processor 56 can acquire the bitstream of the volumetric video.
[0153] In step S820, the processor 56 can decode a first syntax element from the bitstream, which indicates whether occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component. For example, the first syntax element is the syntax element "vme_embedded_occupancy_enabled_flag" or the syntax element "asme_embedded_occupancy_enabled_flag" defined in the MIV syntax table of the V3C standard.
[0154] In step S830, processor 56 determines whether a first syntax element exists during the decoding of the bitstream. When a specific syntax element in the bitstream indicates that the bitstream does not contain an extended syntax structure, processor 56 can determine that the first syntax element does not exist during the decoding of the bitstream.
[0155] In step S840, the processor 56 may, in response to the absence of a first syntax element in the bitstream, set the first syntax element to a first default value to indicate whether occupancy information is absent or present in the V3C sub-bitstream component corresponding to the geometric component. In some embodiments, the first default value may be zero.
[0156] In step S850, the processor 56 can decode the volumetric content from the bitstream based on the value of the first syntax element to reconstruct the volumetric video. Specifically, when the first syntax element is absent, a value of zero for the first syntax element indicates that there is no occupancy information or that occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component. Otherwise, a value of one for the first syntax element indicates that occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component.
[0157] In some embodiments, processor 56 can decode a second syntax element from the bitstream that indicates whether a recommended depth of tiles derived based on tile data unit parameters rather than geometric video data exists. For example, the second syntax element could be the syntax element "asme_patch_constant_depth_flag" defined in the MIV extended syntax table of the V3C standard. Processor 56 can determine whether the second syntax element exists in the bitstream. When a specific syntax element in the bitstream indicates that the bitstream does not contain an extended syntax structure, processor 56 can determine that the second syntax element does not exist when decoding the bitstream. Processor 56 can set the second syntax element to a second default value in response to the absence of the second syntax element in the bitstream, indicating that the depth of the derived tiles is determined by an external means when the relevant syntax element is equal to zero. For example, the relevant syntax element could be the syntax element "vps_geometry_video_present_flag[aspsAtlasID]". The second default value can be zero. That is, when the second syntax element does not exist when decoding the bitstream, processor 56 can set the value of the second syntax element to zero. Processor 56 can decode volumetric content from the bitstream based on the value of the second syntax element to reconstruct the volumetric video. When the second syntax element is absent, a value of zero indicates that the depth of the tile to be derived is determined by external means when the relevant syntax element is zero. Otherwise, a value of one indicates that the recommended depth of the tile is derived based on tile data unit parameters rather than geometric video data.
[0158] In some embodiments, processor 56 can decode a third syntax element from the bitstream that indicates the presence of a relevant syntax element in the multi-view video extended syntax structure. For example, the third syntax element could be the syntax element "asme_patch_texture_offset_enabled_flag" defined in the MIV extended syntax table of the V3C standard. For example, the relevant syntax element could be the syntax element "asme_patch_texture_offset_bit_depth_minus1" in the syntax structure "asps_miv_extension()", or the syntax element "pdu_texture_offset[tileID][p][c]" in the syntax structure "pdu_miv_extension(tileID, p)". Processor 56 can determine whether a third syntax element exists in the bitstream. When a specific syntax element in the bitstream indicates that the bitstream does not contain an extended syntax structure, processor 56 can determine that a third syntax element does not exist when decoding the bitstream. In response to the absence of a third syntax element in the bitstream, processor 56 can set the third syntax element to a third default value to indicate the presence of a relevant syntax element in the multi-view video extended syntax structure. The third default value can be zero. In other words, when a third syntax element is not present during the decoding of the bitstream, processor 56 can set the value of the third syntax element to zero. Processor 56 can then decode the volumetric content from the bitstream based on the value of the third syntax element to reconstruct the volumetric video. When a third syntax element is absent, in response to the absence of a second syntax element in the bitstream, the value of the second syntax element being zero indicates the presence of a relevant syntax element in the multi-view video extended syntax structure. Otherwise, the value of the second syntax element being one indicates the absence of a relevant syntax element in the multi-view video extended syntax structure.
[0159] In some embodiments, processor 56 can decode a fourth syntax element from the bitstream, which indicates the number of bits used to represent the relevant syntax element. For example, the fourth syntax element could be the syntax element "asme_patch_texture_offset_bit_depth_minus1" defined in the MIV extended syntax table of the V3C standard. For example, the relevant syntax element could be the syntax element "pdu_texture_offset[tileID][p][c]" in the syntax structure "pdu_miv_extension(tileID, p)". Processor 56 can determine whether a fourth syntax element exists in the bitstream. When a specific syntax element in the bitstream indicates that the bitstream does not contain an extended syntax structure, processor 56 can determine that a fourth syntax element does not exist when decoding the bitstream. In response to the absence of a fourth syntax element in the bitstream, processor 56 can set the fourth syntax element to a fourth default value to indicate that one bit is used to represent the relevant syntax element. The fourth default value can be zero. That is, when a fourth syntax element does not exist when decoding the bitstream, processor 56 can set the value of the fourth syntax element to zero. Processor 56 can decode volumetric content from the bitstream based on the value of the fourth syntax element to reconstruct the volumetric video. When the fourth syntax element is absent, a value of zero indicates that the number of bits used to represent the pdu_texture_offset[tileID][p][c] syntax element is 1.
[0160] In some embodiments, processor 56 can decode a fifth syntax element from the bitstream, the fifth syntax element indicating the presence of an extended syntax structure associated with volumetric content of at least one specific format in the bitstream. For example, the fifth syntax element is the syntax element "asps_extension_present_flag" defined in the generic ASPS RBSP syntax table in the V3C standard. Processor 56 can decode a sixth syntax element from the bitstream, the sixth syntax element indicating the presence of a multi-view video extended syntax element, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled. For example, the sixth syntax element is the syntax element "asps_miv_extension_present_flag" defined in the generic ASPS RBSP syntax table in the V3C standard. When the fifth syntax element is disabled, the sixth syntax element is not present in the bitstream. Furthermore, when the sixth syntax element is disabled, the multi-view video extended syntax element associated with the sixth syntax element is not present in the bitstream. Processor 56 can decode the multi-view video extended syntax element from the bitstream in response to the sixth syntax element being enabled. The requirement for stream consistency is that when the value of the syntax element "asps_extension_present_flag" is equal to 1, the value of the syntax element "asps_miv_extension_present_flag" should also be equal to 1. Processor 56 can use the Multi-View Video Extension syntax element to decode volumetric content from the stream to reconstruct the volumetric video.
[0161] In some embodiments, processor 56 can decode a seventh syntax element from the bitstream, which indicates the presence of a video-based point cloud compression (V-PCC) extension syntax element for the MIV profile. In response to the seventh syntax element being disabled, processor 56 can decode multi-view video extension syntax elements from the bitstream, and no V-PCC extension syntax element is present for the MIV profile. For example, the seventh syntax element is the extension presence syntax element "aaps_vpcc_extension_present_flag" defined in the Common Atlas Frame Parameter Set (AFPS) RBSP syntax table in the V3C standard and encoded in the MIV toolset profile to indicate the presence of a V-PCC extension syntax element. Since the extension presence syntax element "aaps_vpcc_extension_present_flag" in the MIV toolset profile is set to 0, no V-PCC-related syntax elements are present for parsing the bitstream of volumetric video. Therefore, processor 56 only decodes MIV extension syntax elements, and no V-PCC extension syntax element is present for the MIV profile.
[0162] In some embodiments, processor 56 can decode an eighth syntax element from the bitstream, the eighth syntax element indicating the presence of a seventh syntax element, wherein the seventh syntax element indicates the presence of a V-PCC extension syntax element for the MIV profile. Processor 56 can decode a multi-view video extension syntax element from the bitstream in response to the eighth syntax element being disabled, and the absence of a V-PCC extension syntax element for the MIV profile. For example, the eighth syntax element is an extension presence syntax element "aaps_extension_present_flag" defined in the general AFPS RBSP syntax table in the V3C standard and encoded in the MIV toolset profile to indicate the presence of the extension presence syntax element "aaps_vpcc_extension_present_flag". Since the extension presence syntax element "aaps_extension_present_flag" is set to 0 for the MIV toolset profile, the extension presence syntax element "aaps_vpcc_extension_present_flag" is also set to 0, therefore there are no V-PCC related syntax elements for parsing the MIV bitstream.
[0163] Based on the above, even if the syntax elements “vme_embedded_occupancy_enabled_flag”, “asme_embedded_occupancy_enabled_flag”, “asme_patch_constant_depth_flag”, “asme_patch_texture_offset_enabled_flag”, and “asme_patch_texture_offset_bit_depth_minus1” are not present in the bitstream, the decoder 50 of this embodiment can still decode volumetric content from the volumetric video bitstream, thereby facilitating the encoding of the MIV bitstream.
[0164] Figure 9A and Figure 9B It is a syntax table in the visual volumetric video encoding (V3C) according to embodiments of the present disclosure.
[0165] refer to Figure 9ASyntax Table 92 is the general ASPS RBSP syntax table, in which, in section 922, the extension presence element “asps_extension_present_flag” is encoded to indicate whether the bitstream contains extended data for volume content. If the extension presence element “asps_extension_present_flag” is enabled (i.e., equal to 1), the extension presence element “asps_miv_extension_present_flag” is encoded to indicate the presence of MIV extension syntax elements, and the extension presence element “asps_vpcc_extension_present_flag” is encoded to indicate the presence of V-PCC extension syntax elements.
[0166] refer to Figure 9B Syntax table 94 is the ASPS MIV extended syntax table, which encodes the syntax element "asme_embedded_occupancy_enabled_flag" to indicate whether occupancy information exists in the V3C sub-stream component corresponding to the geometric component. In this embodiment, when the syntax element "asme_embedded_occupancy_enabled_flag" is not present in the bitstream, the decoder sets the value of the syntax element "asme_embedded_occupancy_enabled_flag" to zero to facilitate the encoding of the V-PCC bitstream.
[0167] Furthermore, as shown in syntax table 94, the syntax element "asme_patch_constant_depth_flag" is encoded to indicate whether a recommended depth for tiles derived based on tile data unit parameters rather than geometric video data exists. In this embodiment, when the syntax element "asme_patch_constant_depth_flag" is not present in the bitstream, the decoder sets the value of this syntax element "asme_patch_constant_depth_flag" to zero to facilitate the encoding of the V-PCC bitstream.
[0168] Additionally, as shown in syntax table 94, the syntax element "asme_patch_texture_offset_enabled_flag" is encoded to indicate whether the asme_patch_texture_offset_bit_depth_minus1 exists in the asps_miv_extension() syntax structure, and whether the pdu_texture_offset[tileID][p][c] syntax element exists in the pdu_miv_extension(tileID, p) syntax structure. In this embodiment, when the syntax element "asme_patch_texture_offset_enabled_flag" does not exist in the bitstream, the decoder sets the value of the syntax element "asme_patch_texture_offset_enabled_flag" to zero to facilitate the encoding of the V-PCC bitstream.
[0169] Additionally, as shown in syntax table 94, the syntax element “asme_patch_texture_offset_bit_depth_minus1” is encoded to indicate the number of bits used to represent the pdu_texture_offset[tileID][p][c] syntax element. In this embodiment, when the syntax element “asme_patch_texture_offset_bit_depth_minus1” is not present in the bitstream, the decoder sets the value of this syntax element “asme_patch_texture_offset_bit_depth_minus1” to zero to facilitate the encoding of the V-PCC bitstream.
[0170] Figures 10A to 10C It is a syntax table in the visual volumetric video encoding (V3C) according to embodiments of the present disclosure.
[0171] refer to Figure 10A Syntax table 1002 is a general AAPS RBSP syntax table, in which the extended presence syntax element “aaps_extension_present_flag” is encoded to indicate whether an extended syntax structure associated with volume content of at least one specific format exists in the bitstream. If the extended presence syntax element “aaps_extension_present_flag” is enabled (i.e., equal to 1), the extended presence syntax element “aaps_vpcc_extension_present_flag” is encoded to indicate whether a point cloud extended video syntax element is encoded into the bitstream.
[0172] refer to Figure 10BSyntax table 1004 is a table of syntax element values used for constructing MIV toolset configuration files. This table includes the maximum allowed syntax element values for MIV toolset configuration file construction. In this table, the value of the extension presence syntax element "aaps_vpcc_extension_present_flag" is set to disabled, i.e., equal to 0. That is, for MIV configuration files, the extension presence syntax element "aaps_vpcc_extension_present_flag" is set to zero, and therefore there are no V-PCC related syntax elements for parsing MIV bitstreams, thus facilitating the encoding of MIV bitstreams.
[0173] refer to Figure 10C Syntax table 1006 is also a table of syntax element values used to construct the MIV toolset configuration file, in which the value of the extension presence syntax element "aaps_extension_present_flag" is set to disabled (i.e. equal to 0). This indicates that the extension presence syntax element "aaps_vpcc_extension_present_flag" is disabled (equal to 0), so there are no V-PCC related syntax elements for parsing the MIV bitstream, thereby facilitating the encoding of the MIV bitstream.
[0174] In summary, in the visual volumetric video encoding / decoding (V3C) method, encoder, and decoder disclosed herein, default values are defined for certain MIV-related syntax elements when they are absent from the bitstream, and modifications are proposed to the syntax element value table used to construct the MIV toolset configuration file. Therefore, atlas extension-related syntax elements can be parsed to aid in decoding the MIV bitstream.
[0175] It will be apparent to those skilled in the art that various modifications and variations can be made to the disclosed embodiments without departing from the scope or spirit of this disclosure. In view of the foregoing, this disclosure is intended to cover modifications and variations falling within the scope of the appended claims and their equivalents.
Claims
1. A V3C decoding method based on visual volumetric video, applied to a decoder, the method comprising: Obtain the bitstream of the video file. Decode a first syntax element from the bitstream, the first syntax element indicating whether there is occupancy information in the V3C sub-bitstream component corresponding to the geometric component; In response to the absence of the first syntax element in the bitstream, the first syntax element is set to a first default value to indicate whether the occupancy information is absent or present in the V3C sub-bitstream component corresponding to the geometric component; as well as Based on the value of the first syntax element, the volumetric content is decoded from the bitstream to reconstruct the volumetric video.
2. The method according to claim 1, further comprising: Decode a second syntax element from the bitstream, the second syntax element indicating whether there is a recommended depth for tiles derived based on tile data unit parameters rather than geometric video data; In response to the absence of the second syntax element in the bitstream, the second syntax element is set to a second default value to indicate that when the relevant syntax element is equal to zero, the depth of the exported tile is determined by external means. as well as The volumetric content is decoded from the bitstream based on the value of the second syntax element to reconstruct the volumetric video.
3. The method according to claim 1, further comprising: Decode a third syntax element from the bitstream, the third syntax element indicating whether a relevant syntax element exists in the multi-view video extended syntax structure; In response to the absence of the second syntax element in the bitstream, the third syntax element is set to a third default value to indicate the presence of the relevant syntax element in the multi-view video extended syntax structure; as well as The volumetric content is decoded from the bitstream based on the value of the third syntax element to reconstruct the volumetric video.
4. The method according to claim 1, further comprising: Decode a fourth syntax element from the bitstream, the fourth syntax element indicating the number of bits used to represent the relevant syntax element; In response to the absence of the second syntax element in the bitstream, the fourth syntax element is set to a fourth default value to indicate that one bit is used to represent the relevant syntax element; as well as The volumetric content is decoded from the bitstream based on the value of the fourth syntax element to reconstruct the volumetric video.
5. The method according to claim 1, further comprising: Decode a fifth syntax element from the bitstream, the fifth syntax element indicating whether there is an extended syntax structure in the bitstream associated with volume content of at least one specific format; The sixth syntax element is decoded from the bitstream, the sixth syntax element indicating whether a multi-view video extension syntax element exists, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled; In response to the sixth syntax element being enabled, the multi-view video extension syntax element is decoded from the bitstream; and The volumetric content is decoded from the bitstream using the multi-view video extension syntax element to reconstruct the volumetric video.
6. The method according to claim 1, further comprising: Decode the seventh syntax element from the bitstream, the seventh syntax element indicating whether there is a video-based point cloud compression V-PCC extension syntax element for the MIV profile; as well as In response to the seventh syntax element being disabled, the multi-view video extension syntax element is decoded from the bitstream, and the V-PCC extension syntax element for the MIV profile is not present.
7. The method according to claim 1, further comprising: Decode an eighth syntax element from the bitstream, the eighth syntax element indicating the presence of a seventh syntax element, wherein the seventh syntax element indicating the presence of a V-PCC extended syntax element for the MIV profile; and In response to the eighth syntax element being disabled, the multi-view video extension syntax element is decoded from the bitstream, and the V-PCC extension syntax element for the MIV profile is not present.
8. A decoder, comprising: A communication interface configured to acquire the bitstream of a volumetric video; A storage device configured to store the bitstream of the volumetric video; as well as A processor, coupled to the communication interface and the storage device, and configured to: Decode a first syntax element from the bitstream, the first syntax element indicating whether there is occupancy information in the V3C sub-bitstream component corresponding to the geometric component; In response to the absence of the first syntax element in the bitstream, the first syntax element is set to a first default value to indicate whether there is no occupancy information or the occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component; as well as Based on the value of the first syntax element, the volumetric content is decoded from the bitstream to reconstruct the volumetric video.
9. The decoder according to claim 8, wherein, The processor is also configured to: Decode a second syntax element from the bitstream, the second syntax element indicating whether there is a recommended depth for tiles derived based on tile data unit parameters rather than geometric video data; In response to the absence of the second syntax element in the bitstream, the second syntax element is set to a second default value to indicate that when the relevant syntax element is equal to zero, the depth of the exported tile is determined by external means. as well as The volumetric content is decoded from the bitstream based on the value of the second syntax element to reconstruct the volumetric video.
10. The decoder according to claim 8, wherein, The processor is also configured to: Decode a third syntax element from the bitstream, the third syntax element indicating whether a relevant syntax element exists in the multi-view video extended syntax structure; In response to the absence of the second syntax element in the bitstream, the third syntax element is set to a third default value to indicate the presence of the relevant syntax element in the multi-view video extended syntax structure; as well as The volumetric content is decoded from the bitstream based on the value of the third syntax element to reconstruct the volumetric video.
11. The decoder according to claim 8, wherein, The processor is also configured to: Decode a fourth syntax element from the bitstream, the fourth syntax element indicating the number of bits used to represent the relevant syntax element; In response to the absence of the second syntax element in the bitstream, the fourth syntax element is set to a fourth default value to indicate that one bit is used to represent the relevant syntax element; as well as The volumetric content is decoded from the bitstream based on the value of the fourth syntax element to reconstruct the volumetric video.
12. The decoder according to claim 8, wherein, The processor is also configured to: Decode a fifth syntax element from the bitstream, the fifth syntax element indicating whether there is an extended syntax structure in the bitstream associated with volume content of at least one specific format; The sixth syntax element is decoded from the bitstream, the sixth syntax element indicating whether a multi-view video extension syntax element exists, wherein the sixth syntax element is enabled in response to the fifth syntax element being enabled; In response to the sixth syntax element being enabled, the multi-view video extension syntax element is decoded from the bitstream; and The volumetric content is decoded from the bitstream using the multi-view video extension syntax element to reconstruct the volumetric video.
13. The decoder according to claim 8, wherein, The processor is also configured to: Decode the seventh syntax element from the bitstream, the seventh syntax element indicating whether there is a video-based point cloud compression V-PCC extension syntax element for the MIV profile; as well as In response to the seventh syntax element being disabled, the multi-view video extension syntax element is decoded from the bitstream, and the V-PCC extension syntax element for the MIV profile is not present.
14. The decoder according to claim 8, wherein, The processor is also configured to: Decode an eighth syntax element from the bitstream, the eighth syntax element indicating the presence of a seventh syntax element, wherein the seventh syntax element indicating the presence of a V-PCC extended syntax element for the MIV profile; and In response to the eighth syntax element being disabled, the multi-view video extension syntax element is decoded from the bitstream, and there is no V-PCC extension syntax element for the MIV profile.
15. A V3C encoding method based on visual volumetric video, applied to an encoder, the method comprising: Acquire data from the volumetric video; Based on the data from the volumetric video, determine whether there is occupancy information in the V3C sub-stream component corresponding to the geometric component, so as to encode the first syntax element indication; as well as Generate the bitstream of the video of the stated volume. The presence or absence of the first syntax element in the bitstream of the volumetric video, indicating whether the occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component.
16. The method of claim 15, further comprising: Based on the data from the volumetric video, determine whether there exists a recommended depth for tiles derived from tile data unit parameters rather than geometric video data, in order to encode the second syntax element. The presence or absence of a second syntax element in the bitstream of the volumetric video that indicates the existence of the recommended depth of the tile derived from tile data unit parameters rather than geometric video data.
17. The method of claim 15, further comprising: Based on the data from the volumetric video, determine whether there are relevant syntax elements in the multi-view video extended syntax structure, in order to encode the third syntax element. The presence or absence of the third syntax element in the bitstream of the volumetric video indicates whether the relevant syntax element exists in the multi-view video extended syntax structure.
18. The method of claim 15, further comprising: Based on the data from the volumetric video, the number of bits used to represent the relevant syntax elements is determined for encoding the fourth syntax element. The presence or absence of a fourth syntax element in the bitstream of the volume video that indicates the number of bits used to represent the relevant syntax element.
19. The method of claim 15, further comprising: The fifth syntax element, which indicates the presence of an extended syntax structure associated with volume content of at least one specific format, is encoded into the bitstream; In response to the fifth syntax being enabled, a sixth syntax element indicating the presence of a multi-view video extension syntax element is encoded. In response to the fifth syntax element being enabled, the sixth syntax element is enabled.
20. The method of claim 15, further comprising: The seventh syntax element is encoded into the bitstream to indicate whether the seventh syntax element is disabled, wherein the seventh syntax element is used to indicate whether there is a video-based point cloud compression (V-PCC) extension syntax element for the MIV profile.
21. The method of claim 15, further comprising: The eighth syntax element is encoded into the bitstream to indicate whether the eighth syntax element is disabled, wherein the eighth syntax element is used to indicate whether the seventh syntax element is present, and the seventh syntax element indicates whether the V-PCC extension syntax element for the MIV profile is present.
22. An encoder, comprising: A communication interface configured to acquire volumetric video data; A storage device configured to store the data of the volumetric video; as well as A processor, coupled to the communication interface and the storage device, and configured to: Based on the data from the volumetric video, determine whether there is occupancy information in the V3C sub-stream component corresponding to the geometric component, so as to encode the first syntax element; as well as Generate the bitstream of the video of the stated volume. The presence or absence of the first syntax element in the bitstream of the volumetric video, indicating whether the occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component.
23. The encoder according to claim 22, wherein, The processor is also configured to: Based on the data from the volumetric video, determine whether there exists a recommended depth for tiles derived from tile data unit parameters rather than geometric video data, in order to encode the second syntax element. The presence or absence of a second syntax element in the bitstream of the volumetric video that indicates the existence of the recommended depth of the tile derived from tile data unit parameters rather than geometric video data.
24. The encoder according to claim 22, wherein, The processor is also configured to: Based on the data from the volumetric video, determine whether there are relevant syntax elements in the multi-view video extended syntax structure, in order to encode the third syntax element. The presence or absence of the third syntax element in the bitstream of the volumetric video indicates whether the relevant syntax element exists in the multi-view video extended syntax structure.
25. The encoder according to claim 22, wherein, The processor is also configured to: Based on the data from the volumetric video, the number of bits used to represent the relevant syntax elements is determined for encoding the fourth syntax element. The presence or absence of a fourth syntax element in the bitstream of the volume video that indicates the number of bits used to represent the relevant syntax element.
26. The encoder according to claim 22, wherein, The processor is also configured to: The fifth syntax element, which indicates the presence of an extended syntax structure associated with volume content of at least one specific format, is encoded into the bitstream; In response to the fifth syntax being enabled, a sixth syntax element indicating the presence of a multi-view video extension syntax element is encoded. In response to the fifth syntax element being enabled, the sixth syntax element is enabled.
27. The encoder according to claim 22, wherein, The processor is also configured to: The seventh syntax element is encoded into the bitstream to indicate whether the video-based point cloud compression (V-PCC) extension syntax element for the MIV profile is disabled.
28. The encoder according to claim 22, wherein, The processor is also configured to: The eighth syntax element is encoded into the bitstream to indicate whether the eighth syntax element is disabled, wherein the eighth syntax element is used to indicate whether the seventh syntax element is present, and the seventh syntax element indicates whether the V-PCC extension syntax element for the MIV profile is present.
29. A non-transitory computer-readable recording medium, the non-transitory computer-readable recording medium storing a program that causes a computer to perform the following operations: Decode the first syntax element from the bitstream, the first syntax element indicating whether there is occupancy information in the V3C sub-bitstream component corresponding to the geometric component; In response to the absence of the first syntax element in the bitstream, the first syntax element is set to a first default value to indicate whether there is no occupancy information or the occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component; as well as Based on the value of the first syntax element, the volumetric content is decoded from the bitstream to reconstruct the volumetric video.
30. A non-transitory computer-readable recording medium storing a program that causes a computer to perform the following operations: Acquire data from the volumetric video; Based on the data from the volumetric video, the first syntax element indicating whether occupancy information exists in the V3C sub-stream component corresponding to the geometric component is encoded; and Generate the bitstream of the video of the stated volume. in, The presence or absence of the first syntax element in the bitstream of the volumetric video, indicating whether the occupancy information exists in the V3C sub-bitstream component corresponding to the geometric component, indicates the presence or absence of the first syntax element.