Adaptive Quantization for Instance-Based Mesh Coding

By deriving sub-meshes with varying bit depths and quantization levels based on areal densities, the solution addresses inefficiencies in compressing complex 3D meshes, improving compression efficiency and reducing errors.

JP7727126B2Active Publication Date: 2025-08-20TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024547120
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-05-05
Filing Date
2023-05-24
Publication Date
2025-08-20
Estimated Expiration
2043-05-24

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing complex 3D meshes due to varying importance of mesh regions, number of faces, and bit depth requirements, leading to large quantization errors.

Method used

The proposed solution involves deriving sub-meshes from volumetric data, setting different bit depths for overlapping sub-meshes based on their areal densities, and applying varying quantization levels to achieve efficient compression.

Benefits of technology

This approach reduces quantization errors by adapting bit depth and quantization levels to mesh characteristics, enhancing compression efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727126000023
    Figure 0007727126000023
  • Figure 0007727126000024
    Figure 0007727126000024
  • Figure 0007727126000025
    Figure 0007727126000025
Patent Text Reader

Abstract

1. A method and apparatus, comprising: computer code configured to: cause one or more processors to obtain an input mesh including volumetric data for at least one three-dimensional (3D) visual content; derive a plurality of submeshes of the input mesh from frames of the volumetric data; set a bit depth for a first submesh and a second submesh from the submeshes, the first bit depth being different from the second bit depth; quantize the first submesh and the second submesh based on the first bit depth and the second bit depth, respectively; and signal a result of quantizing the first submesh and the second submesh.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 359,749, filed July 8, 2022, and U.S. Patent Application No. 18 / 312,675, filed May 5, 2023, the disclosures of which are incorporated herein by reference in their entireties.

[0002] This disclosure is directed to a set of advanced video coding techniques, including both lossless and lossy mesh coding techniques based on mesh instances and bit depth. [Background technology]

[0003] Advances in 3D capture, modeling, and rendering are facilitating the ubiquitous presence of 3D content across several platforms and devices. Today, it is possible to film a baby's first steps on one continent, while the baby's grandparents on another continent can view (and possibly interact with) that child and enjoy a fully immersive experience with them. Yet, to achieve such realism, models are becoming ever more sophisticated, and significant amounts of data are linked to the creation and consumption of these models.

[0004] VMesh is an ongoing MPEG standard for compressing static and dynamic meshes. VMesh separates the input mesh into a simplified base mesh and a residual mesh. The base mesh can be encoded at high quality, while the remaining mesh can be encoded using subdivision surface fitting and displacement encoding to take advantage of local characteristics.

[0005] However, complex meshes often contain information about multiple instances to associate and link texture maps. This information is available at the time of encoding. On the other hand, meshes can also be segmented into parts based on their characteristics. For example, a human mesh will have more polygons in the facial region.

[0006] Therefore, a constant quantization step size applied to all instances, objects, or parts in a mesh leads to large quantization errors, mesh regions may not be equally important, the number of faces may vary significantly in different parts of the mesh, the base mesh may be simpler than the original mesh and displacements may therefore require less bit depth precision, etc. Therefore, for any of these reasons, a technical solution to such problems that have arisen in video coding technology is desired. Summary of the Invention [Means for solving the problem]

[0007] Methods and apparatuses are included that include a memory configured to store computer program code and one or more processors configured to access the computer program code and to operate as instructed by the computer program code. The computer program is configured to cause the processor to execute: acquisition code configured to cause the at least one processor to acquire volumetric data of at least one three-dimensional (3D) visual content; derivation code configured to cause the at least one processor to derive a plurality of sub-meshes from a frame of the volumetric data, each of the sub-meshes including a respective one of the instances of the object, and at least two of the sub-meshes overlap each other in the frame; setting code configured to cause the at least one processor to set a first bit depth to a first sub-mesh of at least two of the sub-meshes and a second bit depth to a second sub-mesh of the at least two of the sub-meshes, the first bit depth being different from the second bit depth; quantization code configured to cause the at least one processor to quantize at least two of the sub-meshes based on each of the first bit depth and the second bit depth; and signaling code configured to cause the at least one processor to signal a result of quantizing at least two of the sub-meshes.

[0008] According to various embodiments, the volumetric data includes at least two of the sub-meshes that overlap each other and are surrounded by bounding boxes, and deriving the plurality of sub-meshes includes surrounding each of the object instances with a respective second bounding box, and in either the first bit depth or the second bit depth, the quantization step size of any of at least two of the sub-meshes is smaller than the quantization step size of at least two of the sub-meshes that overlap each other and are surrounded by bounding boxes.

[0009] According to various embodiments, setting the first bit depth is based on determining a first areal density of a first submesh of the at least two of the submeshes, and setting the second bit depth is based on determining a second areal density of a second submesh of the at least two of the submeshes.

[0010] According to various embodiments, quantizing at least two of the submeshes includes applying a first level of quantization to a first submesh of the at least two of the submeshes based on a first bit depth, and applying a second level of quantization to a second submesh of the at least two of the submeshes based on a second bit depth, the first level being set to be lower than the second level based on determining that the first bit depth indicates that the first submesh of the at least two of the submeshes has one of the areal densities that is greater than that indicated by the second bit depth set for the second submesh of the at least two submeshes.

[0011] According to various embodiments, a first of the at least two sub-meshes is surrounded by a first of the second bounding boxes and a second of the at least two sub-meshes is surrounded by a second of the second bounding boxes, and determining the first and second areal densities includes comparing the number of faces of the first of the at least two sub-meshes and the second of the at least two sub-meshes to the volumes of the first of the second bounding boxes and the second of the second bounding boxes, respectively.

[0012] According to various embodiments, signaling the results of quantizing at least two of the sub-meshes includes signaling that each of the at least two of the sub-meshes share the same bounding box offset.

[0013] According to various embodiments, signaling the results of quantizing at least two of the sub-meshes includes signaling a total number of instances of the object in the frame.

[0014] According to various embodiments, signaling the results of quantizing at least two of the submeshes includes signaling differences between the quantizations applied to at least two of the submeshes and sorting the differences compared to other quantizations applied to other submeshes of the frame.

[0015] According to various embodiments, the first bit depth is applied to at least one other submesh of the other submeshes, and signaling the results of quantizing at least two of the submeshes includes grouping a first submesh of the at least two of the submeshes with at least one other submesh of the other submeshes based on determining that the first bit depth is applied to both the first submesh of the at least two of the submeshes and the at least one other submesh of the other submeshes.

[0016] According to various embodiments, setting the first bit depth and the second bit depth includes deriving a mean squared error for each of at least two sub-meshes at each of the first bit depth and the second bit depth. Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]

[0017] [Figure 1]1 is a schematic diagram according to an embodiment. [Figure 2] FIG. 1 is a simplified block diagram according to an embodiment. [Figure 3] 1 is a diagram according to an embodiment. [Figure 4] 1 is a diagram according to an embodiment. [Figure 5] 1 is a diagram according to an embodiment. [Figure 6] 1 is a diagram according to an embodiment. [Figure 7] 1 is a diagram according to an embodiment. [Figure 8] 1 is a diagram according to an embodiment. [Figure 9] 1 is a diagram according to an embodiment. [Figure 10] 1 is a diagram according to an embodiment. [Figure 11] 1 is a diagram according to an embodiment. [Figure 12] 1 is a diagram according to an embodiment. [Figure 13] FIG. 1 is a simplified flow diagram according to an embodiment. [Figure 14] FIG. 1 is a simplified flow diagram according to an embodiment. [Figure 15] FIG. 1 is a simplified flow diagram according to an embodiment. [Figure 16] 1 is a diagram according to an embodiment. [Figure 17] 1 is a diagram according to an embodiment. [Figure 18] 1 is a diagram according to an embodiment. [Figure 19] 1 is a diagram according to an embodiment. [Figure 20] FIG. 1 is a simplified flow diagram according to an embodiment. [Figure 21] 1 is a diagram according to an embodiment. [Figure 22] FIG. 1 is a simplified flow diagram according to an embodiment. [Figure 23] 1 is a diagram according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0018] The proposed features discussed below may be used separately or combined in any order. Furthermore, the embodiments may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0019] 1 illustrates a simplified block diagram of a communication system 100 according to one embodiment of the present disclosure. The communication system 100 may include at least two terminals 102, 103 interconnected via a network 105. In the case of unidirectional data transmission, a first terminal 103 may code video data at a local location for transmission to the other terminal 102 via the network 105. The second terminal 102 may receive the other terminal's coded video data from the network 105, decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media serving applications, for example.

[0020] 1 illustrates a second pair of terminals 101, 104 provided to support bidirectional transmission of coded video, such as may occur during a video conference. For bidirectional transmission of data, each terminal 101, 104 may code video data captured at a local location for transmission to the other terminal over network 105. Each terminal 101, 104 may also receive coded video data transmitted by the other terminal, decode the coded data, and display the recovered video data on a local display device.

[0021] In FIG. 1 , terminals 101, 102, 103, and 104 may be illustrated as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. Embodiments of the present disclosure also apply to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. Network 105 represents any number of networks that convey coded video data among terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this discussion, the architecture and topology of network 105 may not be important to the operation of the present disclosure, unless otherwise described herein below.

[0022] 2 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital television, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0023] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates an uncompressed video sample stream 213. The sample stream 213 may be emphasized as a high data volume when compared to an encoded video bitstream and may be processed by an encoder 202 coupled to the video source 201, which may be a camera, as described above. The encoder 202 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream 204 may be emphasized as a lower data volume when compared to the sample stream and may be stored on a streaming server 205 for future use. One or more streaming clients 212, 207 can access the streaming server 205 to retrieve copies 208, 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes a copy 208 of an input encoded video bitstream and creates an output video sample stream 210 that can be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, 208 may be encoded according to a particular video coding / compression standard. Examples of these standards are mentioned above and further described herein.

[0024] FIG. 3 may be a functional block diagram of a video decoder 300 according to one embodiment of the present invention.

[0025] Receiver 302 may receive one or more codec video sequences to be decoded by decoder 300, in the same or other embodiments, one coded video sequence at a time, with decoding of each coded video sequence independent of other coded video sequences. The coded video sequences may be received from channel 301, which may be a hardware / software link to a storage device that stores the encoded video data. Receiver 302 may receive the encoded video data along with other data, such as coded audio data and / or auxiliary data streams, that may be forwarded to a respective using entity (not shown). Receiver 302 may separate the coded video sequences from other data. To combat network jitter, buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter, “parser”). If receiver 302 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability or from an isosynchronous network, buffer 303 may not be needed or may be small. For use over a best effort packet network such as the Internet, buffer 303 may be required and may be relatively large, and may advantageously be adaptively sized.

[0026] The video decoder 300 may include a parser 304 for reconstructing symbols 313 from the entropy-coded video sequence. These symbol categories include information used to manage the operation of the decoder 300 and, potentially, information for controlling a rendering device, such as a display 312, that is not an integral part of the decoder but may be coupled to the decoder. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The parser 304 may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may be in accordance with a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context-sensitivity, etc. The parser 304 may extract from the coded video sequence a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. Subgroups may include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0027] The parser 304 may perform an entropy decoding / parsing operation on the video sequence received from the buffer 303 to create symbols 313. The parser 304 may receive the encoded data and selectively decode particular symbols 313. Additionally, the parser 304 may determine whether a particular symbol 313 should be provided to the motion compensated prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.

[0028] The reconstruction of symbols 313 may involve several different units, depending on the type of coded video picture or portion thereof (inter-picture and intra-picture, inter-block and intra-block, etc.), as well as other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by parser 304. The flow of such subgroup control information between parser 304 and the following units is not shown for clarity.

[0029] In addition to the functional blocks already mentioned, decoder 300 can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is adequate:

[0030] The first unit is a scalar / inverse transform unit 305. The scalar / inverse transform unit 305 receives quantized transform coefficients and control information from the parser 304 as symbol(s) 313, including the transform to use, block size, quantization coefficients, quantization scaling matrix, etc. The scalar / inverse transform unit 305 may output blocks containing sample values that may be input to an aggregator 310.

[0031] In some cases, the output samples of the scaler / inverse transform unit 305 may also relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture 309. The aggregator 310 may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit 307 to the output sample information provided by the scaler / inverse transform unit 305.

[0032] In other cases, the output samples of the scalar / inverse transform unit 305 may relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit 306 may access the reference picture memory 308 to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols 313 related to the block, these samples may be added by the aggregator 310 to the output of the scalar / inverse transform unit to generate output sample information (in this case, referred to as residual samples or residual signals). The addresses within the reference picture memory from which the motion compensation unit fetches the prediction samples may be controlled by a motion vector and made available to the motion compensation unit in the form of symbols 313, which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory when sub-sample accurate motion vectors are in use, motion vector prediction mechanisms, etc.

[0033] The output samples of aggregator 310 may be subjected to various loop filtering techniques in loop filter unit 311. Video compression techniques may include in-loop filter techniques controlled by parameters contained in the coded video bitstream and made available to loop filter unit 311 as symbols 313 from parser 304, but may also respond to meta-information obtained during decoding of a previous portion (in decoding order) of the coded picture or coded video sequence, or may respond to previously reconstructed, loop-filtered sample values.

[0034] The output of the loop filter unit 311 may be a sample stream that can be output to the rendering device 312 and also stored in the reference picture memory 557 for use in future inter-picture prediction.

[0035] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by parser 304), the current reference picture 309 can become part of reference picture buffer 308, and a new current picture memory can be reallocated before starting reconstruction of a subsequent coded picture.

[0036] Video decoder 300 may perform decoding operations according to a given video compression technique documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense of adhering to the syntax of the video compression technique or standard as specified in the video compression technique document or standard, specifically in a profile document therein. Compliance may also require that the complexity of the coded video sequence be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained by a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0037] In one embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder 300 to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal layer, a spatial layer, or a signal-to-noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0038] FIG. 4 may be a functional block diagram of a video encoder 400 according to one embodiment of the present disclosure.

[0039] The encoder 400 may receive video samples from a video source 401 (not part of the encoder) that may capture the video image(s) to be coded by the encoder 400 .

[0040] The video source 401 may provide a source video sequence to be coded by the encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media serving system, the video source 401 may be a storage device storing previously prepared video. In a video conferencing system, the video source 401 may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art can readily understand the relationship between pixels and samples. The following description focuses on samples.

[0041] According to one embodiment, the encoder 400 may code and compress pictures of a source video sequence into a coded video sequence 410 in real time, or under any other time constraint required by the application. Imposing an appropriate coding rate is one function of the controller 402. The controller controls and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, ...), picture size, Groups of Pictures (GOP) layout, maximum motion vector search range, etc. Other functions of the controller 402 may be relevant to optimizing the video encoder 400 for a particular system design, and therefore will be readily identifiable by those skilled in the art.

[0042] Some video encoders operate in what those skilled in the art readily recognize as a "coding loop." As an overly simplified explanation, the coding loop can consist of an encoding portion of an encoder 400 (hereinafter "source coder") (responsible for creating symbols based on an input picture to be coded and one or more reference pictures) and a (local) decoder 406 embedded in the encoder 400 that reconstructs the symbols to create sample data that a (remote) decoder would also create (since any compression between the symbols and the coded video bitstream is lossless in the video compression techniques contemplated by the disclosed subject matter). The reconstructed sample stream is input to a reference picture memory 405. Because decoding of the symbol stream leads to bit-exact results regardless of the location of the decoder (local or remote), the contents of the reference picture buffer are also bit-exact between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift when synchronism cannot be maintained, eg, due to channel errors) is well known to those skilled in the art.

[0043] The operation of the "local" decoder 406 may be the same as that of the "remote" decoder 300, which has already been described in detail above in connection with Figure 3. However, with brief reference also to Figure 4, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder 408 and parser 304 may be lossless, the entropy decoding portion of the decoder 300, including the channel 301, receiver 302, buffer 303, and parser 304, may not be fully implemented in the local decoder 406.

[0044] At this point, it can be said that any decoder technology, other than parsing / entropy decoding, present in the decoder must necessarily also be present in the corresponding encoder in substantially the same functional form. The description of the encoder technology can be omitted, as it is the inverse of the decoder technology, which has been described generically. Only in certain areas is a more detailed description necessary, which is provided below.

[0045] As part of its operation, source coder 403 may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, coding engine 407 codes differences between pixel blocks of the input frame and pixel blocks of reference frame(s) that may be selected as predictive reference(s) for the input frame.

[0046] The local video decoder 406 may decode coded video data of frames that may be designated as reference frames based on symbols created by the source coder 403. The operation of the coding engine 407 may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence may typically be a copy of the source video sequence, with some errors. The local video decoder 406 may replicate the decoding process that may be performed by the video decoder on the reference frames and store the reconstructed reference frames in a reference picture memory 405, which may be, for example, a cache. In this way, the encoder 400 may locally store copies of reconstructed reference frames that have common content as reconstructed reference frames that will be retrieved by a far-end video decoder (without transmission errors).

[0047] The predictor 404 may perform a predictive search for the coding engine 407. That is, for a new frame to be coded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., that may serve as suitable predictive references for the new picture. The predictor 404 may operate on sample block by pixel block to find a suitable predictive reference. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have predictive references drawn from multiple reference pictures stored in the reference picture memory 405.

[0048] The controller 402 may manage the coding operations of the source coder 403, which may be, for example, a video coder, including, for example, setting parameters and subgroup parameters used to encode the video data.

[0049] The output of all the aforementioned functional units may be entropy coded in entropy coder 408. The entropy coder converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0050] The transmitter 409 may buffer the coded video sequence(s) produced by the entropy coder 408 in preparation for transmission over a communication channel 411, which may be a hardware / software link to a storage device that will store the encoded video data. The transmitter 409 may merge the coded video data from the source coder 403 with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0051] The controller 402 may manage the operation of the encoder 400. During coding, the controller 402 may assign a particular coded picture type to each coded picture, which may affect the coding technique that may be applied to the respective picture. For example, pictures may often be assigned as one of the following frame types:

[0052] An intra-picture (I-picture) may be a picture that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow various types of intra-pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0053] A predicted picture (P picture) may be a picture that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict sample values for each block.

[0054] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0055] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be nonpredictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Pixel blocks of a B-picture may be nonpredictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0056] Encoder 400 may be, for example, a video coder and may perform coding operations in accordance with a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, encoder 400 may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0057] In one embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source coder 403 may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures or slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0058] FIG. 5 illustrates a simplified block-style workflow diagram 500 of exemplary viewport-dependent processing in Omnidirectional Media Application Format (OMAF) that may enable 360-degree virtual reality (VR360) streaming described in OMAF.

[0059] In acquisition block 1001, video data A, such as multiple image and audio data of the same time instance, is acquired if the image data can represent a scene in VR360. In processing block 1003, image B of the same time instance is acquired. i are processed by one or more of stitching, mapping onto a projected picture with respect to one or more virtual reality (VR) angles or other angles / viewpoints, and packed by region. Additionally, metadata indicating any of such processed and other information may be created to assist in the distribution and rendering processes.

[0060] For data D, in the image encoding block 1005, the projected picture is converted to data E i In viewport-independent streaming, the video pictures are encoded into a media file in the video encoding block 1004, for example, as a single-layer bitstream. v and data B is encoded as a Regarding the audio data, the audio data is also encoded in the audio encoding block 1002 as data E a can be encoded into

[0061] Data E a , E v , and E i , the coded bitstream F iand / or the entire F may be stored at a (content delivery network (CDN) / cloud) server and may be transmitted in its entirety, typically in a distribution block 1007 or otherwise, to an OMAF player 1020, from which it may be fully decoded by a decoder in a display block 1016, such that at least one area of the decoded picture corresponding to the current viewport with respect to various metadata, file playback, and orientation / viewport metadata, such as the angle at which the user may be looking through the VR imaging device with respect to the device's viewport specifications. A distinctive feature of VR360 is that only the viewport may be displayed at any particular time, and such a feature may be used to improve the performance of omnidirectional video systems by selective delivery according to the user's viewport (or any other criteria, such as recommended viewport-timed metadata). For example, viewport-dependent delivery may be enabled by tile-based video coding according to an exemplary embodiment.

[0062] Similar to the encoding blocks described above, the OMAF player 1020 according to the exemplary embodiment encodes the data F' and / or F' i and one or more file / segment de-encapsulation of the metadata to decode the audio data E' i is converted into video data E' in the video decoding block 1013. v The image decoding block 1014 converts the image data E' i and decodes the image data A′ in VR360 format according to various metadata such as orientation / viewport metadata in a display block 1016. i The speaker / headphone block 1012 outputs the audio data A' s In order to output the data B' in the audio rendering block 1011, aThe audio may proceed to audio rendering and image rendering of the data D' in image rendering block 1015. It should be understood that various metadata may affect the data decoding and rendering processes depending on various tracks, languages, qualities, views that may be selected by or for a user of the OMAF player 1020, and that the order of processing described herein is presented for an exemplary embodiment and may be performed in other orders according to other exemplary embodiments.

[0063] 6 illustrates a simplified block-style content flow process diagram 600 of (coded) point cloud data with viewpoint- and angle-dependent processing of point cloud data (herein "V-PCC") for six-degrees-of-freedom media capture / generation / (de)coding / rendering / display. It should be understood that the described features may be used separately or combined in any order, and that elements such as encoding and decoding, among others, illustrated, may be performed by processing circuitry (e.g., one or more processors or one or more integrated circuits), and that the one or more processors may execute a program stored on a non-transitory computer-readable medium according to example embodiments.

[0064] Diagram 600 illustrates an exemplary embodiment for streaming coded point cloud data with V-PCC.

[0065] In volumetric data acquisition block 1101, a real-world visual scene or a computer-generated visual scene (or a combination thereof) may be captured by a set of camera devices or synthesized by a computer as volumetric data, which may have any format and may be converted into a (quantized) point cloud data format through image processing in point cloud conversion block 1102. For example, data from the volumetric data may be converted area-by-area data into points of a point cloud by pulling one or more of the values described below from the volumetric data and any associated data into a desired point cloud format according to exemplary embodiments. According to exemplary embodiments, the volumetric data may be a 3D dataset of 2D images, such as slices from which 2D projections of the 3D dataset may be projected. According to an example embodiment, a point cloud data format includes a representation of data points in one or more various spaces and may be used to represent volumetric data and may provide improvements with respect to sampling and data compression, such as with respect to temporal redundancy; for example, point cloud data in x, y, z format may represent color values (e.g., RGB, etc.), brightness, intensity, etc. at each of a plurality of points in the cloud data, and may also be used with progressive decoding, polygon meshing, direct rendering, and octree 3D representation of 2D quadtree data.

[0066] In a projection onto image block 1103, the acquired point cloud data may be projected onto a 2D image and encoded as an image / video picture using video-based point cloud coding (V-PCC). The projected point cloud data may consist of attributes, geometry, occupancy maps, and other metadata used for reconstruction of the point cloud data using, for example, Painter's algorithm, ray casting algorithms, (3D) binary space partitioning algorithms, among others.

[0067] Meanwhile, in the scene generator block 1109, the scene generator may generate several metadata to be used for rendering and displaying six degrees of freedom (DoF) media, for example, according to the director's intent or user preferences. Such 6DoF media may include 360VR-like 3D viewing of a scene from rotational changes on 3D axes X, Y, and Z, in addition to additional dimensions enabling forward / backward, up / down, and left / right movement within, or at least in relation to, the virtual experience in accordance with, the point cloud coded data. The scene description metadata defines one or more scenes composed of the coded point cloud data and other media data, including VR360, light field, audio, etc., and may be provided to one or more cloud servers and / or file / segment encapsulation / deencapsulation processes, as indicated in FIG. 6 and the associated description.

[0068] After a video encoding block 1104 and an image encoding block 1105 similar to the video and image encoding described above (it will also be understood that audio encoding may be provided as described above), a file / segment encapsulation block 1106 processes the coded point cloud data so that it is composed into a media file for file playback or a sequence of default segments and media segments for streaming in a particular media container file format, such as one or more video container formats, and in particular as may be used with respect to DASH, described below, of which such description represents an exemplary embodiment. The file container may also include scene description metadata, such as from a scene generator block 1109, in the files or segments.

[0069] According to an exemplary embodiment, files are encapsulated according to scene description metadata to each include at least one viewpoint position and at least one or more angular views at the one or more viewpoint positions within the 6DoF media at one or more times, such that such files may be transmitted on demand according to user or creator input. Further, according to an exemplary embodiment, segments of such files may include one or more portions of such files, such as portions of 6DoF media indicating a single viewpoint and angle at that viewpoint at one or more times, although these are merely exemplary embodiments and may vary according to various conditions, such as network, user, and creator capabilities and input.

[0070] According to an example embodiment, the point cloud data is divided into multiple 2D / 3D regions, which are coded independently, such as in one or more of the video encoding block 1104 and the image encoding block 1105. Each independently coded partition of the point cloud data may then be encapsulated as a track within a file and / or segment in the file / segment encapsulation block 1106. According to an example embodiment, each point cloud track and / or metadata track may include some useful metadata for viewpoint position / angle dependent processing.

[0071] According to an exemplary embodiment, metadata useful for viewpoint position / angle dependent processing, as included in the file and / or segment encapsulated with respect to the file / segment encapsulation block, includes one or more of: layout information of 2D / 3D partitions with indexes; (dynamic) mapping information associating 3D volumetric partitions with one or more 2D partitions (e.g., tiles / tile groups / slices / subpictures); 3D positions of each 3D partition on the 6DoF coordinate system; a representative viewpoint position / angle list; a selected viewpoint position / angle list corresponding to the 3D volumetric partition; an index of the 2D / 3D partition corresponding to the selected viewpoint position / angle list; information on the quality (rank) of each 2D / 3D partition; and rendering information of each 2D / 3D partition according to, for example, each viewpoint position / angle. Seeking such metadata when requested, such as by a user of the V-PCC player or as directed by a content creator on behalf of a user of the V-PCC player, may enable more efficient processing of particular portions of 6DoF media for which such metadata is desired, such that the V-PCC player may deliver a higher quality image of a valued portion of that media than other portions, rather than delivering unused portions of the 6DoF media.

[0072] The file or one or more segments of the file may be delivered from the File / Segment Encapsulation block 1106 using a delivery mechanism (e.g., by Dynamic Adaptive Streaming over HTTP (DASH)) to the V-PCC Player 1125 and directly to a cloud server, such as in the Cloud Server block 1107, where the cloud server can extract one or more tracks and / or one or more specific 2D / 3D partitions from the file and merge multiple coded point cloud data into one data.

[0073] If the current viewpoint position and angle(s) are defined on the 6DoF coordinate system at the client system according to data such as position / viewing angle tracking block 1108, the cloud server may extract the appropriate partition(s) from the store file(s) and merge them (if necessary) according to metadata from, for example, the client system with V-PCC player 1125, and the viewpoint position / angle metadata may be delivered from file / segment encapsulation block 1106 so that the extracted data can be delivered to the client as files or segments, or alternatively may be processed in other ways in cloud server block 1107 from files or segments already on the cloud server.

[0074] For such data, in the file / segment decapsulation block 1115, the file decapsulator processes the file or received segment, extracts the coded bitstream and parses the metadata, and then in the video decode and image decode block, the coded point cloud data is decoded and reconstructed into point cloud data in the point cloud reconstruction block 1112, which can be displayed in the display block 1114 and / or, for scene description data by the scene generator block 1109, can be initially configured according to one or more various scene descriptions in the scene construction block 1113.

[0075] In view of the above, such an exemplary V-PCC flow represents advantages over the V-PCC standard, including one or more of the described partitioning capabilities for multiple 2D / 3D areas, the ability to assemble compressed domains of coded 2D / 3D partitions into a single coherent coded video bitstream, and the bitstream extraction capabilities for assemblies of coded 2D / 3D of a coded picture into a coherent coded bitstream, and such V-PCC system support is further improved by including a container formation for the VVC bitstream to support a mechanism for accommodating metadata that carries one or more of the above-mentioned metadata.

[0076] In that regard, according to an exemplary embodiment described further below, the term "mesh" refers to a configuration of one or more polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in 3D space and information about how the vertices are connected, referred to as connectivity information. Optionally, vertex attributes, such as color, normals, etc., can be associated with mesh vertices. Attributes can also be associated with the surface of a mesh by utilizing mapping information that parameterizes the mesh with a 2D attribute map. Such mapping can be defined by a set of parametric coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. The 2D attribute map is used to store high-resolution attribute information, such as texture, normals, displacement, etc. Such information can also be used for various purposes, such as texture mapping and shading, according to an exemplary embodiment.

[0077] Nevertheless, dynamic mesh sequences can require large amounts of data because they can consist of a significant amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content. Mesh compression standards IC, MESHGRID, and FAMC were previously developed by MPEG to address dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards do not take time-varying attribute maps and connectivity information into account. DCC (Digital Content Creation) tools typically generate such dynamic meshes. Correspondingly, it is difficult for volumetric acquisition techniques to generate constant connectivity dynamic meshes, especially under real-time constraints. This type of content is not supported by existing standards. According to exemplary embodiments herein, aspects of a new mesh compression standard are described for directly processing dynamic meshes with time-varying connectivity information and, optionally, time-varying attribute maps. The standard targets lossy and lossless compression for various applications, such as real-time communication, storage, free-viewpoint video, and AR and VR. Features such as random access and scalable / progressive coding will also be considered.

[0078] FIG. 7 illustrates an exemplary framework 700 for dynamic mesh compression, such as for a 2D atlas sampling-based method. Each frame of an input mesh 1201 can be preprocessed by a series of operations, such as tracking, remeshing, parameterization, and voxelization. Note that these operations can be encoder-only, meaning they may not be part of the decoding process; such a possibility can be signaled in the metadata by a flag, such as 0 indicating encoder-only and 1 otherwise. A mesh with a 2D UV atlas 1202 can then be obtained, with each vertex of the mesh having one or more associated UV coordinates on the 2D atlas. The mesh can then be converted into multiple maps, including a geometry map and an attribute map, by sampling on the 2D atlas. These 2D maps can then be coded by a video / image codec, such as HEVC, VVC, AV1, AVS3, etc. At the decoder 1203 side, a mesh can be reconstructed from the decoded 2D maps. Any post-processing and filtering can also be applied to the reconstructed mesh 1204. Note that other metadata may be signaled to the decoder side for the purpose of 3D mesh reconstruction. Note that chart boundary information, including uv and xyz coordinates of boundary vertices, can be predicted, quantized, and entropy coded in the bitstream. The quantization step size can be configured at the encoder side for a tradeoff between quality and bitrate.

[0079] In some implementations, a 3D mesh can be partitioned into several segments (or patches / charts). Each segment consists of a set of connected vertices associated with their geometry, attributes, and connectivity information. As illustrated in the volumetric data example 800 of FIG. 8, the UV parameterization process 1302 of mapping 3D mesh segments onto a 2D chart, such as the 2D UV atlas 1202 block described above, maps one or more mesh segments 1301 onto a 2D chart 1303 in the 2D UV atlas 1304. Each vertex (v n ) will be assigned 2D UV coordinates in the 2D UV atlas. n ) form a connected component as their 3D counterparts. The geometry, attributes, and connectivity information of each vertex can be inherited from their 3D counterparts as well. For example, if vertex v4 is connected to vertex v 0、 v 5、 Information about the direct connection of v1, v3 may be indicated, and similarly, information for each of the other vertices may be supported. Furthermore, such a 2D textured mesh, according to an exemplary embodiment, would further indicate information such as color information on a per-patch basis for each triangle, e.g., patches of v2, v5, v3, etc.

[0080] For example, see example 900 of FIG. 9, where, in addition to the features of example 800 of FIG. 8, a 3D mesh segment 1301 can also be mapped to multiple separate 2D charts 1401, 1402. In this case, a vertex in 3D can also correspond to multiple vertices in the 2D UV atlas. As shown in FIG. 9, the same 3D mesh segment is mapped to multiple 2D charts in the 2D UV atlas, instead of a single chart as in FIG. 8. For example, 3D vertices v1 and v4 have two 2D correspondences v1, v4, respectively. 1’ and v4, v 4’Thus, a typical 2D UV atlas of a 3D mesh may consist of multiple charts, as shown in Figure 14, where each chart may contain multiple (usually three or more) vertices associated with their 3D geometry, attributes, and connectivity information.

[0081] FIG. 10 shows an example 1000 illustrating a derived triangulation for a chart with boundary vertices B0, B1, B2, B3, B4, B5, B6, and B7. Given such information, any triangulation method can be applied to create connectivity between vertices (including boundary vertices and sampled vertices). For example, for each vertex, find the two closest vertices. Or, for every vertex, successively generate triangles until a minimum number of triangles is achieved after a set number of attempts. As shown in example 1000, there are repeating triangles of various regular shapes and triangles of various irregular shapes, typically closest to the boundary vertex and with their own unique dimensions that may or may not be shared with any other triangles. Connectivity information can also be reconstructed by explicit signaling. If a polygon cannot be reconstructed by implicit rules, the encoder can signal connectivity information in the bitstream, according to an exemplary embodiment.

[0082] Looking at the patches discussed above, the example 1000 may represent one such patch, such as the patch formed from vertices v3, v2, and v5 shown in either FIG. 14 or FIG.

[0083] Boundary vertices B0, B1, B2, B3, B4, B5, B6, B7 are defined in 2D UV space. As shown in Figure 15, a filled vertex is a boundary vertex because it lies on the boundary edge of a connected component (patch / chart). A boundary edge can be determined by checking whether the edge appears in only one triangle. The following information of a boundary vertex is important and needs to be signaled in the bitstream according to an exemplary embodiment: Geometry information, e.g., 3D XYZ coordinates, even if currently in 2D UV parameter format, and 2D UV coordinates.

[0084] For the case where a bounding vertex in 3D corresponds to multiple vertices in the 2D UV atlas, the mapping from 3D XYZ to 2D UV can be one-to-many, as shown in Figure 9. Therefore, a UV to XYZ (or called UV2XYZ) index can be signaled to indicate the mapping function. UV2XYZ can be a 1D array of indices that map each 2D UV vertex to a 3D XYZ vertex.

[0085] According to an example embodiment, to efficiently represent a mesh signal, a subset of mesh vertices may be first coded along with connectivity information between them. In the original mesh, connections between these vertices may not exist because the vertices are subsampled from the original mesh. There are various ways to signal connectivity information between vertices, and therefore such a subset is referred to as a base mesh or base vertices.

[0086] However, other vertices can be predicted by applying interpolation between two or more already decoded mesh vertices. A predictor vertex has its geometric position along the edges of two connected existing vertices, so the geometric information of the predictor can be calculated based on nearby decoded vertices. In some cases, a displacement vector or prediction error from the vertex to be coded to the vertex predictor should also be coded. For example, see example 1100 in Figure 11, which shows an example of such edge-based vertex prediction, more specifically, vertex geometry prediction using extrapolation by expanding a triangle into a parallelogram as shown on the left, and intra-prediction by interpolation using a weighted average of two existing vertices as shown on the right. After decoding the base vertices (i.e., the solid triangles 1801 on the left in Figure 11), interpolation between these base vertices can be performed along the connecting edges. For example, the midpoints of each edge can be generated as predictors. Therefore, the geometric positions of these interpolated points are the (weighted) average of the two nearby decoded vertices (the dashed points 1802 on the left in Figure 11). Having more than one midpoint between two already decoded vertices can also be done in a similar way. Therefore, the actual vertex to be coded can be reconstructed by adding a displacement vector to the predictor (center of Figure 7). After decoding these additional vertices, the connection between the newly decoded vertex and the existing base vertex is still maintained. In addition, further connections between the newly decoded vertices can be established. Along with the base vertex, more intermediate vertex predictors can be generated along new edges (right of Figure 7) by connecting these newly decoded vertices 1803 and the base vertex to each other. Thus, there are more actual vertices to be decoded with associated displacement vectors.

[0087] According to an example embodiment, mesh vertices of mesh frame 1902 can also be predicted from the decoded vertices of the previously coded mesh frame 1901. This prediction mechanism is called inter-prediction. An example of mesh geometry inter-prediction is shown in example 1200 of Figure 12, which illustrates vertex geometry prediction using inter-prediction (vertices of the previous mesh frame become predictors for vertices of the current frame). In some cases, a displacement vector or prediction error from the vertex to be coded to the vertex predictor should also be coded.

[0088] According to an exemplary embodiment, several methods are implemented for dynamic mesh compression, which are part of the edge-based vertex prediction framework described above, where a base mesh is first coded and then more additional vertices are predicted based on connectivity information from the edges of the base mesh. Note that the methods can be applied individually or in any form of combination.

[0089] For example, consider vertex grouping for the example flowchart 1300 of the prediction mode in FIG. 13. In S201, vertices in a mesh may be obtained, and in S202, they can be divided into different groups for prediction purposes (see, e.g., FIG. 10). In one example, the division is performed using patch / chart partitioning in S204 as previously discussed. In another example, the division is performed under each patch / chart S205. The decision S203 to proceed to S204 or S205 may be signaled by a flag or the like. In S205, some vertices in the same patch / chart form a prediction group and share the same prediction mode, while some other vertices in the same patch / chart may use a different prediction mode. Such grouping in S206 can be assigned at different levels by determining the number of respective vertices involved per group. For example, every 64, 32, or 16 vertices in scan order within a patch / chart may be assigned the same prediction mode according to an example embodiment, while other vertices may be assigned differently. For each group, the prediction mode can be an intra prediction mode or an inter prediction mode, which can be signaled or assigned. According to the example flowchart 1300, if a mesh frame or mesh slice is determined to be an intra type in S207, such as by checking whether a flag for the mesh frame or mesh slice indicates an intra type, all vertex groups within the mesh frame or mesh slice shall use the intra prediction mode; otherwise, in S208, either the intra prediction mode or the inter prediction mode may be selected for each group for all vertices therein.

[0090] Furthermore, for a group of mesh vertices using intra prediction mode, the vertices can only be predicted using previously coded vertices within the same subpartition of the current mesh. The subpartition may sometimes be the current mesh itself, according to an exemplary embodiment, and for a group of mesh vertices using inter prediction mode, the vertices can only be predicted using previously coded vertices from other mesh frames, according to an exemplary embodiment. Each of the above information may be determined and signaled by a flag or the like. The above prediction characteristics may be performed in S210, and the results of the above prediction and signaling may occur in S211.

[0091] According to an exemplary embodiment, for each vertex in a vertex group in the exemplary flowchart 1300 and the flowchart 1400 described below, after prediction, the residual becomes a 3D displacement vector indicating the shift from the current vertex to its predictor. The residual of the vertex group needs to be further compressed. In one example, the transform in S211 can be applied to the residual of the vertex group before entropy coding, along with its signaling. To handle the coding of a group of displacement vectors, the following method can be implemented. For example, one method is to appropriately signal the group of displacement vectors, some displacement vectors, or cases where its components have only zero values. In another embodiment, a flag can be signaled for each displacement vector indicating whether this vector has non-zero components, and if not, the coding of all components of this displacement vector can be skipped. Furthermore, in another embodiment, a flag can be signaled for each group of displacement vectors indicating whether this group has non-zero components, and if not, the coding of all displacement vectors of this group can be skipped. Furthermore, in other embodiments, a flag may be signaled for each component of a group of displacement vectors indicating whether this component of the group has any non-zero vectors, and if not, coding of this component of all displacement vectors of this group may be skipped. Furthermore, in other embodiments, there may be signaling if a group of displacement vectors or a component of a group of displacement vectors requires a transform, and if not, the transform may be skipped and quantization / entropy coding may be applied directly to the group or group component. Furthermore, in other embodiments, a flag may be signaled for each group of displacement vectors indicating whether this group needs to undergo a transform, and if not, transform coding of all displacement vectors of this group may be skipped.Furthermore, in other embodiments, a flag may be signaled for each component of a group of displacement vectors whether this component of the group needs to undergo a transform, and if not, the transform coding of this component of all displacement vectors of this group may be skipped. The above-described embodiments in this paragraph regarding the processing of vertex prediction residuals may also be combined and performed in parallel, each on a different patch.

[0092] FIG. 14 shows an example flowchart 1400 in which a mesh frame can be coded as an entire data unit in S221, meaning that all vertices or attributes of the mesh frame may have correlations between them. Alternatively, depending on the determination in S222, the mesh frame can be divided into smaller independent subpartitions in S223, similar in concept to slices or tiles of a 2D video or image. The coded mesh frame or coded mesh subpartition can be assigned a prediction type in S224. Possible prediction types include intra-coding and inter-coding. In the intra-coding type, only prediction from a reconstructed portion of the same frame or slice is allowed in S225. Meanwhile, the inter-prediction type allows prediction from a previously coded mesh frame in addition to intra-mesh frame prediction in S225. Furthermore, the inter-prediction type can be classified into more subtypes, such as P-type and B-type. In the P-type, only one predictor can be used for prediction purposes, while in the B-type, two predictors from two previously coded mesh frames can be used to generate a predictor. A weighted average of two predictors may be an example. If a mesh frame is coded as a whole, the frame may be considered an intra- or inter-coded mesh frame. For inter-mesh frames, the P or B type may be further identified via signaling. Alternatively, if the mesh frame is coded with further division within the frame, a prediction type assignment for each sub-partition is performed in S224. Each of the above information may be determined and signaled by a flag or the like, and similar to S210 and S211 of FIG. 13, the prediction characteristics may be performed in S226, and the results of the prediction and signaling may occur in S227.

[0093] As such, dynamic mesh sequences can require large amounts of data as they can consist of a significant amount of information that changes over time, and efficient compression techniques are needed to store and transmit such content, and the above-described features of Figures 20 and 21 represent such improved efficiency by enabling at least improved mesh vertex 3D position prediction by using previously decoded vertices either within the same mesh frame (intra-prediction) or from a previously coded mesh frame (inter-prediction).

[0094] Further, exemplary embodiments may generate displacement vectors for a third layer 2303 of the mesh based on one or more of the reconstructed vertices of the second layer 2302 and its previous layer(s), such as the first layer 2301. Assuming that the index of the second layer 2302 is T, predictors for vertices in the third layer 2303T+1 are generated based on at least the reconstructed vertices of the current layer or the second layer 2302. An example of such a layer-based prediction structure is shown in example 1600 of FIG. 16, which illustrates reconstruction-based vertex prediction, i.e., progressive vertex prediction using edge-based interpolation, in which predictors are generated based on previously decoded vertices rather than predictor vertices. The first layer 2301 may be a mesh bounded by a first polygon 2340 having as its vertices the decoded vertices at its boundary and an interpolated vertex along one of the lines between the decoded vertices. As the progressive coding proceeds from the first layer 2301 to the second layer 2302, additional polygons 2341 may be formed by displacement vectors from one of the interpolated vertices of the first layer to additional vertices of the second layer 2302, so that the total number of vertices in the second layer 2302 may be greater than the total number of vertices in the first layer 2301. Similarly, proceeding to the third layer 2303, the additional vertices of the second layer 2302, together with the decoded vertices from the first layer 2301, may function in the coding in the same way as the decoded vertices functioned in proceeding from the first layer 2301 to the second layer 2303, i.e., multiple additional polygons may be formed. Note that unlike Figure 16, example 1900 illustrates that when progressing from the first layer 2601 to the second layer 2603 and then to the third layer 2603, each of the additionally formed polygons may be entirely within the polygon formed by the extent of the first layer 2601; see example 1900 in Figure 19 which illustrates such progressive coding.

[0095] For such an example 1600, since the interpolated vertices of the current layer are predicted values, such values need to be reconstructed before being used to generate predictors for the vertices of the next layer. Refer to the exemplary flowchart 1500 of FIG. 15. This is done by coding the base mesh at S231, performing vertex prediction as such at S232, and then adding the decoded displacement vectors of the current layer to the predictors of the vertices such as in layer 2302 at S233. Then, the reconstructed vertices of this layer 2303 can be used to generate and signal the predictor vertices of the next layer 2303 at S235, together with all the decoded vertices of the previous (one or more) layers, such as checking the additional vertex values of such a layer at S234. This process can also be summarized as follows. Let P[t](Vi) represent the predictor of vertex Vi on layer t, R[t](Vi) represent the reconstructed vertex Vi on layer t, D[t](Vi) represent the displacement vector of vertex Vi of layer t, and f(*) represent a predictor generator that can, in particular, be the average of two existing vertices. In that case, according to an exemplary embodiment, for each layer t, the following exists. P[t](Vi)=f(R[s|s<t](Vj), R[m|m<t](Vk)), where Vj and Vk are the reconstructed vertices of the previous layer R[t](Vi)=P[t](Vi)+D[t](Vi) Equation (1)

[0096] Then, for all the vertices of one mesh frame, they are divided into layer 0 (base mesh), layer 1, layer 2, …, etc. In that case, the reconstruction of the vertices on one layer depends on the reconstruction of the vertices on the previous (one or more) layers. Above, each of P, R, and D represents a 3D vector in the context of a 3D mesh representation. D is the decoded displacement vector, and quantization may or may not be applied to this vector.

[0097] According to an exemplary embodiment, vertex prediction using reconstructed vertices may be applied only to certain layers, such as layer 0 and layer 1. For other layers, vertex prediction can still use nearby predictor vertices without adding displacement vectors to them for reconstruction. Therefore, these other layers can be processed simultaneously without waiting for the previous layer to reconstruct. According to an exemplary embodiment, for each layer, the selection of reconstruction-based vertex prediction or predictor-based vertex prediction can be signaled, or the layer (and subsequent layers) that does not use reconstruction-based vertex prediction can be signaled.

[0098] For displacement vectors generated by vertices whose vertex predictors are reconstructed, quantization can be applied to those displacement vectors without further transformation, such as wavelet transform, etc. For displacement vectors whose vertex predictors are generated by other predictor vertices, transformation may be required, and quantization can be applied to the transform coefficients of those displacement vectors.

[0099] As such, dynamic mesh sequences may require large amounts of data because they may consist of a significant amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content. In the framework of the interpolation-based vertex prediction method described above, one key step is compressing displacement vectors, which occupy a major portion of the coded bitstream and are the focus of this disclosure; for example, the features of FIG. 15 alleviate this problem by providing such compression.

[0100] Furthermore, as with the other examples described above, even in those embodiments, dynamic mesh sequences may still require large amounts of data since they may consist of a significant amount of information that changes over time, thus requiring efficient compression techniques to store and transmit such content. Within the framework of the 2D atlas sampling-based method described above, significant advantages may be achieved by inferring connectivity information from sampled vertices and boundary vertices at the decoder side. This is a key part of the decoding process and is the focus of further examples described below.

[0101] According to an exemplary embodiment, the connectivity information of the base mesh can be inferred (derived) from the decoded boundary vertices and sampled vertices for each chart on both the encoder and decoder sides.

[0102] Similar to that described above, any triangulation method can be applied to create connectivity between vertices (including boundary vertices and sampled vertices). For charts without sampling of interior vertices, such as the interior vertices shown in example 1000 of Figure 10 and example 1800 of Figure 18 described further below, a similar method of creating connectivity still applies, but according to an example embodiment, the use of different triangulation methods for boundary vertices and sampled vertices may be signaled.

[0103] For example, according to an exemplary embodiment, for each of four neighboring points at any sampled location, it may be determined whether the number of occupied points is three or greater (examples of occupied or unoccupied points are highlighted in FIG. 18 , which shows an example occupancy map 1800 in which each circle represents an integer pixel), and the triangular connectivity between the four points may be inferred by a specific rule. For example, as illustrated in example 1700 of FIG. 17 , in the case of the illustrated examples (2), (3), (4), and (4), where three of the four points are determined to be occupied, those points may be directly interconnected to form a triangle as in those examples; whereas, if all four points are determined to be occupied, those points may be used to form two triangles as shown in example (1) of FIG. 17 . Note that different rules may be applied to different numbers of neighboring points. This process may be performed across many points, as further illustrated in FIG. 18 . In this embodiment, the reconstructed mesh is a triangular mesh such as that of Figure 17, and also a regular triangular mesh in at least the interior portion of Figure 18, which may not be determined to be signaled according to such regularity, but may instead be coded and decoded as surrounding irregular triangles that should be signaled separately, by inference rather than by separate signaling.

[0104] Also, in an attempt to further reduce complexity and data processing, such a regular interior triangular quadrilateral mesh as shown in FIG. 18 may be inferred as such a quadrilateral mesh of example (1) of FIG. 17, thereby reducing even the amount of complexity from inferring interior regular triangles to instead inferring a reduced number of interior regular quadrilateral meshes.

[0105] According to an exemplary embodiment, a quadrilateral mesh may be reconstructed when all four neighboring points are determined to be occupied, as in example (1) of FIG.

[0106] Inferring from the above description, it is illustrated that the reconstructed mesh in example 1800, as shown in FIG. 18, may be of a hybrid type, i.e., some regions within the mesh frame generate triangular meshes and other regions generate quadrilateral meshes, some of the triangular meshes may be regular compared to other triangular meshes therein, and some may be irregular, such as boundary triangular meshes, but not necessarily all of such meshes are on the boundary.

[0107] According to an example embodiment, such connectivity type can be signaled in a high-level syntax such as a sequence header or a slice header.

[0108] As mentioned above, connectivity information can also be reconstructed by explicit signaling, such as in the case of irregularly shaped triangular meshes. That is, if it is determined that a polygon cannot be reconstructed by implicit rules, the encoder can signal connectivity information in the bitstream. Also, according to exemplary embodiments, the overhead of such explicit signaling can be reduced depending on the polygon boundary. For example, as shown in example 1800 of FIG. 18, triangle connectivity information is signaled to be reconstructed by both implicit rules, such as according to the regular example 2400 of FIG. 17, which can be inferred, and explicit signaling, at least for irregularly shaped polygons shown on the mesh boundary in FIG. 18.

[0109] According to an embodiment, it is decided that only connectivity information between boundary vertices and sampled locations is signaled, connectivity information between the sampled locations themselves is inferred.

[0110] Also, in any of the embodiments, connectivity information can be signaled predictively, so that only the difference between the estimated connectivity (as a prediction) from one mesh to another can be signaled in the bitstream.

[0111] As a note, the inferred triangle orientation (e.g., inferred clockwise or counterclockwise for each triangle) can, according to an example embodiment, be signaled for all charts in a high-level syntax such as a sequence header, slice header, etc., or fixed (assumed) by the encoder and decoder. The inferred triangle orientation can also be signaled differently for each chart.

[0112] As a further note, any reconstructed mesh may have a different connectivity than the original mesh, for example, the original mesh may be a triangle mesh, while the reconstructed mesh may be a polygon mesh (e.g., a quadrilateral mesh).

[0113] According to an exemplary embodiment, connectivity information for any base vertices may not be signaled, and instead, edges between base vertices may be derived using the same algorithm at both the encoder and decoder sides. See, for example, how the bottom vertices in example 1800 are all occupied, and therefore coding may utilize such information, thus determining that such vertices are occupied as bases, thereby requiring connectivity information for any base vertices not to be signaled, and instead, by inferring later, edges between base vertices may be derived using the same algorithm at both the encoder and decoder sides. Also, according to an exemplary embodiment, predicted vertex interpolation for additional mesh vertices may be based on the derived edges of the base mesh.

[0114] According to an exemplary embodiment, a flag may be used to signal whether connectivity information of a base vertex should be signaled or derived, and such a flag can be signaled at different levels of the bitstream, such as the sequence level, the frame level, etc.

[0115] According to an exemplary embodiment, the edges between base vertices are first derived using the same algorithm on both the encoder and decoder sides. Then, the difference between the derived edges and the actual edges is signaled compared with the original connectivity of the base mesh vertices. Therefore, after decoding the difference, the original connectivity of the base vertices can be restored.

[0116] In one example, for a derived edge, if it is determined to be erroneous when compared to the original edge, such information may be signaled in the bitstream (by indicating the pair of vertices that form this edge), and for an original edge, if it is not derived, it may be signaled in the bitstream (by indicating the pair of vertices that form this edge). Furthermore, connectivity on and vertex interpolation involving boundary edges may be performed separately from interior vertices and edges.

[0117] Therefore, by means of the exemplary embodiments described herein, the technical problems described above may be advantageously ameliorated by one or more of these technical solutions. For example, dynamic mesh sequences may consist of a significant amount of information that changes over time and therefore may require a large amount of data, and therefore the exemplary embodiments described herein represent at least an efficient compression technique for storing and transmitting such content.

[0118] The above-described embodiments may further be applied to instance-based mesh coding, where an instance may be a mesh of an object or a portion of an object. For example, the illustrative example 2100 of FIG. 21 shows an example mesh 2801 in which various instances 2802 (representing a mesh of a cup), 2803 (representing a mesh of a spoon), and 2804 (representing a mesh of a plate) exist and may each be separated and coded. Also, while each of the instances 2801, 2802, 2803, and 2804 are illustrated within respective bounding boxes, which will be described further below, it should be noted that the instance 2801 may be considered to be illustrated as being enclosed by a "mesh-based bounding box," while each of the instances 2802, 2803, and 2804 may be considered to be enclosed by a respective "instance-based bounding box."

[0119] According to exemplary embodiments, the proposed methods may be used separately or combined in any order. Although only triangular meshes were used to demonstrate various embodiments, the proposed methods may be used for any polygonal mesh. As mentioned above, it is assumed that an input mesh may contain one or more instances, and a sub-mesh is a portion of an input mesh that has one or more instances, and multiple instances can be grouped to form a sub-mesh.

[0120] In that regard, Figure 20 illustrates an example 2000 in which it is proposed to separately quantize different objects or portions at a given input bit depth (which bit depth may be referred to as "QP"). For example, at 2701, one or more input meshes may be obtained and each separated into multiple sub-meshes. The sub-meshes may be objects, instances of objects, or segmented regions, which, according to an exemplary embodiment, are independently quantized at S2702.

[0121] According to an example embodiment, a mesh M having m points at (x, y, z) coordinates may be quantized by QP bit depth in S2702. The quantization step size in all three dimensions (x, y, z) is determined by the maximum length d of the bounding box in all dimensions. bbox >0. Also, the same quantization step size may be applied in S2704 to all objects in the mesh identified in S2703 as follows:

number

number

number

number

number

[0122] However, in complex scenes, the largest objects are relatively often simple backgrounds that can tolerate higher quantization step sizes, while the main objects are smaller and suffer from large quantization errors that can be accounted for by various embodiments described further below.

[0123] Therefore, as shown in example 2200 of FIG. 22, the bounding box d of the input mesh bbox The maximum length of is always the bounding box of each instance:

number

number

number

[0124] For a given bit depth QP, the quantization step size of each of all instances, instance 2802 (representing a mesh of a cup), instance 2803 (representing a mesh of a spoon), and instance 2804 (representing a mesh of a plate) is always

number

[0125] Therefore, the quantization error per instance is smaller, thus reducing the overall quantization error.

[0126] According to various embodiments, bit depth may be adaptively assigned to each instance / region, referred to as a "submesh," in S2902 and may be determined based on the areal density of that particular instance. Each submesh may be obtained from the volumetric data of the mesh, which may itself individually signal each instance within the mesh, and each submesh is derived from that mesh on an instance-by-instance basis in S2902. For example, each of instances 2802, 2803, 2804 may be assigned its own respective bit depth in S2904 according to its own particular areal density or the number of vertices forming one or more of the aforementioned polygons therein. In general, the more faces each instance has, which may be determined in S2903 by counting the number of such polygons therein, etc., the less quantization should be applied to that instance in S2702. For example, given a mesh M, where the total number of faces is n, the corresponding face of submesh k is:

number

number

number

number

[0127] According to various embodiments, the meshes are represented as a base mesh B and its corresponding displacement D, which are quantized at different bit depths in S2702. For example, for the kth object, the bit depths of the base mesh B and D are

number

number

number

[0128] According to various embodiments, an adaptive bit depth parameter based on minimizing distortion may be used. For example, given an input bit depth QP, the mean squared error (MSE) of the quantization method is ∈_QP and may be as in equation (4). The MSE of each sub-mesh is derived as ∈_QP̂k=ω_k*ε_QP, ∀k∈[1,...,K], where ω_k>0 is a weighting factor. In one example, ω_k=1∀k. A linear search is performed as

number

[0129] Additionally, the best bit depth for displacement is

number

[0130] According to an example embodiment, there may be signaling of the quantization of each object, such as by signaling via the S2907 bitstream. The set of base quantization bit depths in ascending order is

number

number

[0131] [Table 1]

[0132] In the table, u(n) is an unsigned integer using n bits, i(n) is an integer using n bits, mips_quant() is a sequence of signaling data, mips_min_bbox[k] is the minimum bounding box in the i-th dimension, mips_num_instances_minus1 is the number of instances in the mesh minus 1, mips_base_bitdepth_minus1 is the bit depth of the first instance in this order, mips_base_quant[k] is the difference between the quantization of the (k+1)th and kth submesh. Since the quantization set is sorted in ascending order, this number is always non-negative. mips_dist_quant[k] is the kth quantization data of the base mesh bit depth.

[0133] According to various embodiments, to reduce signaling overhead, multiple instances may be grouped into K groups with the same bit depth. The instances are clustered by bounding boxes using a simple clustering method such as K-means clustering.

number

[0134] The techniques described above can be implemented as computer software physically stored on one or more computer-readable media using computer-readable instructions, or by one or more specially configured hardware processors. For example, Figure 23 illustrates a computer system 2300 suitable for implementing certain embodiments of the disclosed subject matter.

[0135] Computer software can be coded using any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, linking, etc. to create code containing instructions that can be executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc. directly, or via interpretation, microcode execution, etc.

[0136] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.

[0137] 23 for computer system 2300 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither should the arrangement of components be interpreted as having a dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system 2300.

[0138] The computer system 2300 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not shown). The human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0139] The input human interface devices may include one or more of a keyboard 2301, a mouse 2302, a trackpad 2303, a touchscreen 2310, a joystick 2305, a microphone 2306, a scanner 2308, and a camera 2307 (only one of each is shown).

[0140] The computer system 2300 may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen 2310 or joystick 2305, although haptic feedback devices that do not function as input devices may also be present), audio output devices (e.g., speakers 2309, headphones (not shown)), visual output devices (e.g., screens 2310 including CRT screens, LCD screens, plasma screens, and OLED screens, each with or without touchscreen input functionality and each with or without haptic feedback capabilities, some of which may be capable of outputting two-dimensional visual output or output in more than three dimensions via means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0141] The computer system 2300 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 2320 with media such as CD / DVD 2311, thumb drive 2322, removable hard drive or solid state drive 2323, legacy magnetic media such as tape or floppy disk (not shown), dedicated ROM / ASIC / PLD based devices such as security dongles (not shown), etc.

[0142] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not encompass transmission media, carrier waves, or other transitory signals.

[0143] The computer system 2300 may also include an interface 2399 to one or more communications networks 2398. The network 2398 may be, for example, wireless, wired, optical, or the like. The network 2398 may further be local, wide area, metropolitan, vehicular, industrial, real-time, delay tolerant, etc. Examples of networks 2398 include local area networks such as Ethernet, WLAN, etc., cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicular and industrial networks including CANBus, etc. Particular networks 2398 generally require an external network interface adapter attached to a particular general-purpose data port or peripheral bus (2350 and 2351) (e.g., a USB port on the computer system 2300), while other networks are generally built into the core of the computer system 2300 by connection to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks 2398, the computer system 2300 can communicate with other entities. Such communications can be unidirectional receive only (e.g., broadcast television), unidirectional transmit only (e.g., a CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Particular protocols and protocol stacks can be used with each of these networks and network interfaces, as described above.

[0144] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be connected to core 2340 of computer system 2300 .

[0145] The core 2340 may include one or more central processing units (CPUs) 2341, graphics processing units (GPUs) 2342, graphics adapters 2317, dedicated programmable processing devices in the form of field programmable gate arrays (FPGAs) 2343, hardware accelerators for particular tasks 2344, etc. These devices may be connected via a system bus 2348, along with read-only memory (ROM) 2345, random access memory 2346, internal mass storage such as an internal non-user-accessible hard drive, SSD, etc. 2347. In some computer systems, the system bus 2348 may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be connected directly to the core's system bus 2348 or via a peripheral bus 2349. Architectures for peripheral buses include PCI, USB, etc.

[0146] The CPU 2341, GPU 2342, FPGA 2343, and accelerator 2344 can execute specific instructions that, in combination, can constitute the aforementioned computer code. That computer code can be stored in ROM 2345 or RAM 2346. Persistent data can be stored in, for example, internal mass storage 2347, while temporary data can also be stored in RAM 2346. Fast storage and retrieval from any of the memory devices can be enabled by the use of cache memory, which can be closely associated with one or more of the CPU 2341, GPU 2342, mass storage 2347, ROM 2345, RAM 2346, etc.

[0147] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0148] By way of example and not limitation, computer system 2300 having the architecture, and specifically core 2340, may provide functionality as a result of processor(s) (including CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be user-accessible mass storage devices as described above, as well as media associated with specific storage of core 2340 that is non-transitory in nature, such as core internal mass storage device 2347 or ROM 2345. Software implementing various embodiments of the present disclosure may be stored in such devices and executed by core 2340. Computer-readable media may include one or more memory devices or chips, depending on particular needs. The software may cause core 2340, and specifically the processor(s) therein (including CPU, GPU, FPGA, etc.), to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM 2346 and modifying such data structures according to the software-defined processes. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator 2344), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software may, where appropriate, encompass logic, and vice versa. References to computer-readable media may, where appropriate, encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0149] While this disclosure describes several exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. Thus, it will be appreciated that those skilled in the art can devise numerous systems and methods that, while not explicitly shown or described herein, embody the principles of the present disclosure and are thus within its spirit and scope. [Explanation of symbols]

[0150] 100 Communication Systems Terminals 101-104 105 Network 201 Video Sources 202 Encoder 203 Capture Subsystem 204 encoded video bitstream 205 Streaming Server 206 Copy of encoded video bitstream 207 Streaming Client 208 Copy of encoded video bitstream 209 Display 210 Output Video Sample Stream 211 Video Decoder 212 Streaming Client 213 uncompressed video sample stream 300 decoder 301 Channel 302 Receiver 303 Buffer Memory 304 Parser 305 Scaler / Descaler Unit 306 Motion Compensation Prediction Unit 307 Intra Prediction Unit 308 Reference Picture Buffer 309 Current Picture 310 Aggregator 311 Loop Filter 312 Display 313 Symbol 400 Encoder 401 Source 402 Controller 403 Source Coder 404 Predictor 405 Reference Picture Memory 406 decoder 407 Coding Engine 408 Entropy Coder 409 Transmitter 410 coded video sequence 411 Channel 500 Workflow Diagrams 600 Content Flow Process Diagram 700 one dynamic mesh compression framework 800 Volumetric Data Examples 900 examples 1001 Acquired Block 1002 Audio Encoding Block 1003 Processing Block 1004 Video Encoding Block 1005 Image Encoding Block 1007 Delivery Block 1008 Head / Eye Tracking Block 1000 examples 1010 Audio Decode Block 1011 Audio Rendering Block 1012 Speaker / Headphone Block 1013 Video Decoding Block 1014 Image Decoding Block 1015 Image Rendering Block 1016 display blocks 1020 OMAF player 1100 examples 1101 Volumetric Data Acquisition Block 1102 Point Cloud Block 1103 Image Block 1104 Video Encoding Block 1105 Image Encoding Block 1106 File / Segment Encapsulation Block 1107 Cloud Server Block 1108 Position / Field of View Tracking Block 1109 Scene Generator Block 1112 Point Cloud Reconstruction Block 1113 Scene Building Blocks 1114 Display Block 1125 V-PCC Player 1200 examples 1201 input mesh 1202 2D UV Atlas 1203 decoder 1204 Reconstructed Mesh 1300 Exemplary Flowchart 1301 3D Mesh Segments 1302 UV Parameterization Process 1303 2D Chart 1304 2D UV Atlas 1400 Exemplary Flowchart 1401, 1402 2D Chart 1500 Exemplary Flowchart 1600~1800 examples 1801 Solid Triangle 1802 dashed points 1803 Newly Decoded Vertices 1900 examples Mesh frames coded before 1901 1902 mesh frame 2100 examples 2301 First Layer 2302 Second Layer 2303 Third Layer 2340 First Polygon 2341 additional polygons 2601 First Layer 2602 Second Layer 2603 Third Layer 2801 mesh example 2802 instances representing the cup mesh 2803 instances representing the spoon mesh 2804 instances representing the dish mesh

Claims

1. A method for video coding performed by at least one processor, the method comprising: obtaining an input mesh comprising volumetric data of at least one three-dimensional (3D) visual content; deriving a plurality of sub-meshes of the input mesh from the frame of volumetric data, each of the sub-meshes including a respective one of the object instances, and at least a first sub-mesh and a second sub-mesh from the plurality of the sub-meshes overlap each other in the frame; setting a first bit depth for the first sub-mesh and a second bit depth for the second sub-mesh, the first bit depth being different from the second bit depth; quantizing the first sub-mesh and the second sub-mesh based on the first bit-depth and the second bit-depth, respectively; signaling the results of quantizing the first sub-mesh and the second sub-mesh; 1. A method for video coding, comprising:

2. the volumetric data includes the first sub-mesh and the second sub-mesh overlapping each other and surrounded by a bounding box; deriving the plurality of sub-meshes includes enclosing each of the instances of the object with a respective second bounding box; At either the first bit depth or the second bit depth, a quantization step size of either the first submesh or the second submesh is smaller than a quantization step size of the first submesh and the second submesh that overlap each other and are enclosed by the bounding box; 2. A method for video coding according to claim 1.

3. the step of setting the first bit depth is based on the step of determining a first areal density of the first sub-mesh; setting the second bit depth is based on determining a second areal density of the second sub-mesh; 3. A method for video coding according to claim 2.

4. quantizing the first sub-mesh and the second sub-mesh includes: applying a first level of quantization to the first sub-mesh based on the first bit depth; applying a second level of quantization to the second sub-mesh based on the second bit depth; the first level is set to be lower than the second level based on determining that the first bit depth indicates that the first sub-mesh has one of the areal densities greater than that indicated by the second bit depth set for the second sub-mesh; A method for video coding according to claim 3.

5. the first sub-mesh is bounded by a first of the second bounding boxes; the second sub-mesh is bounded by a second one of the second bounding boxes; determining the first and second areal densities includes comparing the number of faces of the first sub-mesh and the second sub-mesh with the volume of the first one of the second bounding boxes and the second one of the second bounding boxes, respectively; A method for video coding according to claim 3.

6. signaling the result of quantizing the first submesh and the second submesh includes signaling that each of the first submesh and the second submesh share the same bounding box offset.

3. A method for video coding according to claim 2.

7. signaling the result of quantizing the first sub-mesh and the second sub-mesh includes signaling a total number of the instances of the object in the frame.

2. A method for video coding according to claim 1.

8. signaling the results of quantizing the first submesh and the second submesh includes signaling a difference between the quantizations applied to the first submesh and the second submesh and sorting the difference compared to other quantizations applied to other submeshes of the frame.

2. A method for video coding according to claim 1.

9. the first bit depth is applied to at least one other sub-mesh of the other sub-meshes; and signaling the result of quantizing the first submesh and the second submesh includes grouping the first submesh with the at least one other submesh among the other submeshes based on determining that the first bit depth applies to both the first submesh and the at least one other submesh among the other submeshes. A method for video coding according to claim 8.

10. wherein the step of setting the first bit depth and the second bit depth includes the step of deriving a mean square error for each of the first sub-mesh and the second sub-mesh at each of the first bit depth and the second bit depth.

2. A method for video coding according to claim 1.

11. An apparatus configured to perform the method of any one of claims 1 to 10.

12. A computer program for causing a computer to carry out the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method and apparatus for 3D model compression based on the discovery of repeating structures

    JP2015520886A

  • Apparatus and method for displacement mesh compression

    JP2021149942A

  • Information processing device and method

    WO2019012975A1