Improvement of the Coding of Boundary UV2XYZ Indices for Mesh Compression

An interpolation-based hierarchical prediction scheme addresses the inefficiencies in existing mesh compression standards by predicting and reconstructing boundary vertices, enabling efficient compression and transmission of dynamic meshes for real-time applications.

JP7704979B2Active Publication Date: 2025-07-08TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024527415
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-28
Filing Date
2023-03-30
Publication Date
2025-07-08
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Existing mesh compression standards do not effectively handle dynamic meshes with time-varying connectivity information and attribute maps, especially under real-time conditions, leading to inefficiencies in data storage and transmission.

Method used

A method involving an interpolation-based hierarchical prediction scheme is used to predict and reconstruct boundary vertices in a compressed 2D mesh, deriving prediction residuals to efficiently compress and decompress dynamic meshes with time-varying attributes and connectivity.

Benefits of technology

This approach enables efficient compression and transmission of dynamic meshes, supporting real-time applications like AR and VR by reducing data volume and improving reconstruction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007704979000017
    Figure 0007704979000017
  • Figure 0007704979000018
    Figure 0007704979000018
  • Figure 0007704979000019
    Figure 0007704979000019
Patent Text Reader

Abstract

The method, performed by at least one processor in a decoder, includes receiving a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to a surface of a three-dimensional (3D) volumetric object. The method further includes predicting a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction scheme using at least one sampled vertex included in the compressed 2D mesh. The method further includes deriving a prediction residual associated with the predicted current vertex. The method further includes reconstructing a boundary vertex associated with the 3D volumetric object based on the predicted current vertex and the derived prediction residual.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 331,705, filed Apr. 15, 2022, and U.S. Patent Application No. 18 / 191,600, filed Mar. 28, 2023, the disclosures of which are hereby incorporated by reference in their entirety.

[0002] This disclosure is directed to a set of advanced video coding techniques. More specifically, this disclosure is directed to video - based mesh compression including a method for UV information of boundary vertices for efficient mesh compression.

Background Art

[0003] The world's advanced three - dimensional (3D) representations enable more immersive interactions and communications. To achieve a sense of presence in 3D representations, 3D models have become more sophisticated than ever, and a significant amount of data is associated with the creation and consumption of these 3D models. 3D meshes are widely used in 3D model - immersive content.

[0004] A 3D mesh can be composed of several polygons that describe the surface of a volumetric object. A dynamic mesh sequence can require a large amount of data because it can have a significant amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content.

[0005] Mesh compression standards IC, MESHGRID, FAMC were previously developed to handle dynamic meshes with constant connectivity and time - varying geometry and vertex attributes. However, these standards do not take into account time - varying attribute maps and connectivity information.

[0006] Furthermore, especially under real-time constraints, it is also difficult for volume acquisition techniques to generate a constantly connected dynamic mesh. This type of dynamic mesh content is not supported by existing standards.

Summary of the Invention

Means for Solving the Problems

[0007] According to one or more embodiments, a method performed by at least one processor in a decoder includes receiving a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to a surface of a three-dimensional (3D) volume object. The method includes predicting a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction scheme using at least one sampled vertex included in the compressed 2D mesh. The method includes deriving a prediction residual associated with the predicted current vertex. The method further includes reconstructing boundary vertices associated with the 3D volume object based on the predicted current vertex and the derived prediction residual.

[0008] According to one or more embodiments, a decoder includes at least one memory configured to store program code and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes receive code configured to cause the at least one processor to receive a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to a surface of a three-dimensional (3D) volume object. The program code includes prediction code configured to cause the at least one processor to predict a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction method using at least one sampled vertex included in the compressed 2D mesh. The program code includes derivation code configured to cause the at least one processor to derive a prediction residual associated with the predicted current vertex. The program code includes reconstruction code configured to cause the at least one processor to reconstruct a boundary vertex associated with the 3D volume object based on the predicted current vertex and the derived prediction residual.

[0009] According to one or more embodiments, a non-transitory computer-readable medium storing instructions that, when executed by at least one processor in a decoder, cause the decoder to receive a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to a surface of a three-dimensional (3D) volume object. The instructions further cause the at least one processor to predict a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction scheme that uses at least one sampled vertex included in the compressed 2D mesh. The instructions further cause the at least one processor to derive a prediction residual associated with the predicted current vertex. The instructions further cause the at least one processor to reconstruct a boundary vertex associated with the 3D volume object based on the predicted current vertex and the derived prediction residual.

[0010] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0011]

Fig. 1

Fig. 2

Fig. 3

Fig. 4

Fig. 5

Fig. 6

Fig. 7

Fig. 8

Fig. 9

Fig. 10

Fig. 11

Mode for Carrying Out the Invention

[0012] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0013] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the disclosed embodiments to the exact form disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from the practice of the embodiments. Further, one or more features or components of an embodiment may be incorporated into or combined with those of other embodiments (or one or more features of other embodiments). In addition, in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be interchanged.

[0014] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation form. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.

[0015] Even if specific combinations of features are recited in the claims and / or disclosed herein, these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0016] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described as such. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or similar words are used. Also, as used in this specification, terms such as "has," "have," "having," "include," "including," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.

[0017] Throughout this specification, references to "an embodiment," "one embodiment," or similar expressions mean that the particular feature, structure, or characteristic described in connection with the indicated embodiment is included in at least one embodiment of the solution. Thus, the phrases "in an embodiment," "in one embodiment," and similar phrases do not necessarily all refer to the same embodiment, although they may.

[0018] Furthermore, the features, advantages, and characteristics described in this disclosure may be combined in any suitable manner in one or more embodiments. Those skilled in the art will recognize, in light of the description herein, that the disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages that may not be present in all embodiments of the disclosure may be recognized in a particular embodiment.

[0019] Embodiments of the present disclosure are directed to compressing a mesh. The mesh can be composed of several polygons that describe the surface of a volumetric object. Information about the vertices of the mesh in 3D space and how the vertices are connected defines each polygon and can be referred to as connectivity information. Optionally, vertex attributes such as color, normal, etc. may be associated with the mesh vertices. The attributes may also be associated with the surface of the mesh by utilizing mapping information that parameterizes the mesh with a 2D attribute map. Such a mapping is defined using a set of parametric coordinates called UV coordinates or texture coordinates and can be associated with the mesh vertices. The 2D attribute map can be used to store high-resolution attribute information such as texture, normal, displacement, etc. The high-resolution attribute information can be used for various purposes such as texture mapping and shading.

[0020] As described above, a 3D mesh or dynamic mesh can consist of a significant amount of information that changes over time and thus may require a large amount of data. Existing standards do not take into account time-varying attribute maps and connectivity information. Existing standards also do not support volume acquisition techniques for generating always-connected dynamic meshes, especially under real-time conditions.

[0021] Therefore, there is a need for a new mesh compression standard for directly handling dynamic meshes with time-varying connectivity information and optionally time-varying attribute maps. Embodiments of the present disclosure enable efficient compression techniques for storing and transmitting such dynamic meshes. Embodiments of the present disclosure enable irreversible compression and / or reversible compression for various applications such as real-time communication, storage, free viewpoint video, AR, and VR.

[0022] According to one or more embodiments of the present disclosure, a method, a system, and a non-transitory storage medium for dynamic mesh compression are provided. Embodiments of the present disclosure can also be applied to a static mesh where only one frame of the mesh or the mesh content does not change over time.

[0023] Referring to FIGS. 1-2, one or more embodiments of the present disclosure for implementing the encoding and decoding structures of the present disclosure are described.

[0024] FIG. 1 illustrates a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected via a network 150. In the case of unidirectional data transmission, the first terminal 110 may code video data, which may include mesh data, at a local location for transmission to the other terminal 120 via the network 150. The second terminal 120 may receive the coded video data of the other terminal from the network 150, decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media serving applications and the like.

[0025] FIG. 1 also illustrates a second pair of terminals 130, 140 provided to support bidirectional transmission of coded video that may occur, for example, during a video conference. In the case of bidirectional data transmission, each terminal 130, 140 may code video data captured at a local location for transmission to the other terminal via the network 150. Each terminal 130, 140 may also receive the coded video data transmitted by the other terminal, decode the coded data, and display the restored video data on a local display device.

[0026] In FIG. 1, terminals 110 to 140 may be, for example, servers, personal computers, and smartphones, and / or any other type of terminal. For example, the terminals (110 to 140) may be laptop computers, tablet computers, media players and / or dedicated video conferencing devices. Network 150 represents any number of networks that transmit coded video data among terminals 110 to 140, including, for example, wired and / or wireless communication networks. Communication network 150 can exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network 150 may not be important for the operation of the present disclosure, unless otherwise described herein below.

[0027] FIG. 2 illustrates the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be used in other video-related applications including, for example, video conferencing, digital television, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0028] As illustrated in FIG. 2, the streaming system 200 may include a capture subsystem 213 that includes a video source 201 and an encoder 203. The streaming system 200 may further include at least one streaming server 205 and / or at least one streaming client 206.

[0029] Video source 201 can create a stream 202 that includes, for example, a 3D mesh and metadata associated with the 3D mesh. The 3D mesh can be composed of several polygons that describe the surface of a volumetric object. For example, the 3D mesh may include a plurality of vertices in 3D space where each vertex is associated with 3D coordinates (e.g., x, y, z). The video source 201 may include, for example, a 3D sensor (e.g., a depth sensor) or 3D imaging technology (e.g., one or more digital cameras) and a computing device configured to generate a 3D mesh using data received from the 3D sensor or 3D imaging technology. The sample stream 202 may have a high data volume compared to an encoded video bitstream and can be processed by an encoder 203 coupled to the video source 201. The encoder 203 may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoder 203 may also further generate an encoded video bitstream 204. The encoded video bitstream 204 may have a lower data volume compared to the uncompressed stream 202 and can be stored on a streaming server 205 for later use. One or more streaming clients 206 can access the streaming server 205 and retrieve a video bitstream 209 that can be a copy of the encoded video bitstream 204.

[0030] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may decode a video bitstream 209, which is, for example, an input copy of the encoded video bitstream 204, and create an output video sample stream 211 that can be rendered on the display 212 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 209 may be encoded according to a particular video coding / compression standard.

[0031] FIG. 3 is an exemplary diagram of a framework 300 for dynamic mesh compression and mesh reconstruction using an encoder and a decoder.

[0032] As seen in FIG. 3, the framework 300 may include an encoder 301 and a decoder 351. The encoder 301 may include one or more input meshes 305, a mesh 310 having one or more UV atlases, an occupancy map 315, a geometry map 320, an attribute map 325, and metadata 330. The decoder 351 may include a decoded occupancy map 335, a decoded geometry map 340, a decoded attribute map 345, decoded metadata 350, and a reconstructed mesh 360.

[0033] According to one or more embodiments of the present disclosure, the input mesh 305 may include one or more frames, and each of the one or more frames may be preprocessed by a series of operations and used to generate a mesh 310 having a UV atlas. As an example, the preprocessing operations may include, but are not limited to, tracking, parameterization, remeshing, voxelization, etc. In some embodiments, the preprocessing operations may be performed only on the encoder side and not on the decoder side.

[0034] The mesh 310 with a UV atlas can be a 2D mesh. The 2D mesh can be a chart of vertices each associated with coordinates (e.g., 2D coordinates) in a 2D space. Each vertex in the 2D mesh may be associated with a corresponding vertex in the 3D mesh, and the vertices in the 3D mesh are associated with coordinates in 3D space. The compressed 2D mesh can be a version of the 2D mesh with reduced information compared to the uncompressed 2D mesh. For example, the 2D mesh may be sampled at a sampling rate where the compressed 2D mesh contains sampled points. The 2D mesh with a UV atlas can be a mesh where each vertex of the mesh can be associated with UV coordinates on the 2D atlas. For example, the 2D atlas can be a two-dimensional plane where each 3D coordinate in 3D space can be assigned 2D coordinates in the 2D plane. The connected 2D coordinates can be called a 2D chart or patch. The mesh 310 with a UV atlas can be processed based on sampling and converted into multiple maps. As an example, the UV atlas 310 can be processed based on sampling of the 2D mesh with a UV atlas and converted into an occupancy map, a geometry map, and an attribute map. The generated occupancy map 335, geometry map 340, and attribute map 345 can be encoded using an appropriate codec (e.g., HVEC, VVC, AV1, AVS3, etc.) and sent to a decoder. In some embodiments, metadata (e.g., connectivity information, etc.) can also be sent to the decoder.

[0035] In some embodiments, on the decoder side, a mesh can be reconstructed from the decoded 2D map. Post-processing and filtering can also be applied to the reconstructed mesh. In some examples, the metadata can be signaled to the decoder side for the purpose of 3D mesh reconstruction. The occupancy map can be inferred from the decoder side when the boundary vertices of each patch are signaled.

[0036] According to one aspect, decoder 351 may receive an encoded occupancy map, geometry map, and attribute map from an encoder. In addition to the embodiments described herein, decoder 315 may use appropriate techniques and methods to decode the occupancy map, geometry map, and attribute map. In some embodiments, decoder 351 may generate a decoded occupancy map 335, a decoded geometry map 340, a decoded attribute map 345, and decoded metadata 350. Input mesh 305 may be reconstructed into a reconstructed mesh 360 based on the decoded occupancy map 335, the decoded geometry map 340, the decoded attribute map 345, and the decoded metadata 350 using one or more reconstruction filters and techniques. In some embodiments, metadata 330 may be sent directly to decoder 351, and decoder 351 may use the metadata to generate a reconstructed mesh 360 based on the decoded occupancy map 335, the decoded geometry map 340, and the decoded attribute map 345. Post-filtering techniques, including but not limited to remeshing, parameterization, tracking, voxelization, etc., may also be applied to the reconstructed mesh 360.

[0037] According to some embodiments, a 3D mesh may be divided into several segments (or patches / charts). Each segment may be composed of a set of connected vertices associated with their geometry, attributes, and connectivity information. As illustrated in FIG. 4, the UV parameterization process maps mesh segment 400 onto 2D charts (402, 404) within a 2D UV atlas. 2D UV coordinates within the 2D UV atlas may be assigned to each vertex within the mesh segment. Vertices within the 2D chart may form connection components as their 3D corresponding vertices. The geometry, attributes, and connectivity information of each vertex may also be inherited from their 3D corresponding geometry, attributes, and connectivity information.

[0038] According to some embodiments, a 3D mesh segment can also be mapped to a plurality of separate 2D charts. When a 3D mesh segment is mapped to separate 2D charts, the vertices within the 3D mesh segment may correspond to a plurality of vertices within a 2D UV atlas. As illustrated in FIG. 5, a 3D mesh segment 500 that can correspond to a 3D mesh segment 400 may be mapped to two 2D charts (502A, 502B) in the 2D UV atlas instead of a single chart. As illustrated in FIG. 5, the 3D vertices v1 and v4 each have two 2D corresponding vertices v1' and v4'.

[0039] FIG. 6 illustrates an example of a general 2D UV atlas 600 of a 3D mesh that includes a plurality of charts, where each chart may include a plurality of (e.g., three or more) vertices associated with their 3D geometry, attributes, and connectivity information.

[0040] Boundary vertices can be defined within 2D UV space. As shown in FIG. 7, the filled vertices are boundary vertices because they are on the boundary edges of the connecting components (patches / charts). The boundary edges can be determined by checking whether the edge appears in only one triangle. The geometry information (e.g., 3D xyz coordinates) and 2D UV coordinates can be signaled in a bitstream.

[0041] A dynamic mesh sequence can require a large amount of data because it can be composed of a significant amount of information where the mesh sequence changes over time. In particular, the boundary information represents a significant portion of the entire mesh. Therefore, an efficient compression technique is needed to efficiently compress the boundary information.

[0042] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0043] According to embodiments of the present disclosure, several methods are proposed for coding UV coordinates of patch boundaries in mesh compression. Note that the methods may be applied individually or in any form of combination. Note also that the methods may be applied to a mesh that has only one frame or to a static mesh whose mesh content does not change over time. Further, similar methods may be extended to coding of depth images / attribute images / texture images, etc.

[0044] According to some embodiments, an input mesh having a 2D UV atlas may have vertices, and each vertex of the mesh may have associated UV coordinates on the 2D atlas. The occupancy map, geometry map, and attribute map may be generated by sampling one or more points / positions on the UV atlas. Each sample position may or may not be occupied if the position is inside the polygon defined by the mesh vertices. For each occupied sample, the corresponding 3D geometry coordinates and attributes of the sample may be calculated by interpolating from the associated polygon vertices.

[0045] According to one aspect of the present disclosure, the sampling rate may be consistent across the entire 2D atlas. In some embodiments, the sampling rates of the u-axis and v-axis may be different, enabling anisotropic remeshing. In some embodiments, the entire 2D atlas may be divided into a plurality of regions such as slices or tiles, and each such region may have a different sampling rate.

[0046] According to one or more embodiments of the present disclosure, the sampling rate of each region (or the entire 2D atlas) can be signaled in a high-level syntax including, but not limited to, sequence headers, frame headers, slice headers, etc. In some embodiments, the sampling rate of each region (or the entire 2D atlas) can be selected from a set of pre-established rates assumed by both the encoder and the decoder. Since the set of pre-established rates can be known to both the encoder and the decoder, signaling of one particular sampling rate only needs to signal the index within the pre-established rate set. Examples of such pre-established sets can be, for example, every 2 pixels, every 4 pixels, every 8 pixels, etc. In some embodiments, the sampling rate of each region (or the entire 2D atlas) of a mesh frame can be predicted from a set of pre-established rates, from the previously used sampling rate within other already coded regions of the same frame, or from the previously used sampling rate within other already coded mesh frames.

[0047] In some embodiments, the sampling rate of each region (or the entire 2D atlas) can be based on some characteristic of each region (or the entire 2D atlas). As an example, the sample rate may be based on activity, and in the case of a rich texture region (or the entire 2D atlas), or a high-activity region (or the entire 2D atlas), the sample rate can be set high. As another example, in the case of a smooth region (or the entire 2D atlas), or a low-activity region (or the entire 2D atlas), the sample rate can be set low.

[0048] In some embodiments, the sampling rate of each region of the mesh frame (or the entire 2D atlas) can be signaled in a way that allows a combination of prediction and direct signaling. The syntax can be structured to indicate whether the sampling rate is predicted or directly signaled. If predicted, it can be further signaled which predictor sampling rate should be used. If directly signaled, a syntax representing the value of the rate can be signaled.

[0049] FIG. 8 shows an example of sampling coordinates on the UV plane 800, where the texture points V1, V2, and V3 are the UV coordinates of the boundary vertices. The entire UV plane can be sampled at a constant sample interval. The distance between two adjacent samples in FIG. 8 (e.g., in terms of the original number of samples) can be called the sampling step (or sampling rate). The original UV coordinates of these vertices may not exactly match the sampling positions on the UV plane, such as the original position of V1. The closest sampling coordinate V1' can be used as a predictor in compression.

[0050] The UV coordinates of the boundary vertices can be coded in two parts, for example, including sampling coordinates and an offset. In FIG. 8, the coordinates (u i , v i ) may be the original UV coordinates of the boundary vertices of the patch, where i = 0, 1,..., N - 1 and N is the number of boundary vertices in the chart. The sampling rate of the chart can be represented by S. After sampling the UV plane 800, the sampling coordinates of the boundary vertices

Number

Number

[0051] In one or more examples, various rounding operations, such as floor or ceiling operations, may also be applied after division when calculating sampling coordinates. Offset of the UV coordinates of the boundary vertices

Number

Number

[0052] Therefore, the offset of the boundary vertex V1 in FIG. 8 can be calculated as follows.

Number

[0053] In one or more examples, both the sampling coordinates and the offset can be coded by reversible or irreversible coding. The reconstructed sampling coordinates and offset can be represented by

Number

Number

Number

Number

[0054] According to some embodiments, previous coded sampling coordinates may be used to predict current sampling coordinates. In one or more embodiments, the UV coordinates of vertices on the same boundary loop may be coded together in order. The UV coordinates of one vertex before the current vertex in the decoding order may be used as a predictor for the current UV coordinates. The predicted sampling coordinates of the current boundary vertex are

Number

[0055] In one or more embodiments, previous coded sampling coordinates may be used to predict current coordinates as follows.

Number

[0056] In one or more embodiments, two previously coded sampling coordinates may be used to predict the current coordinates. For example, the current coordinates may be predicted using two previously coded sampling coordinates with linear prediction as follows.

Number

[0057] In one or more examples, prediction by higher-order polynomials may also be employed.

[0058] According to one or more embodiments, an interpolation-based hierarchical prediction method can be used for sampling UV coordinates. FIG. 9 illustrates a set 900 of sampled vertices. In one or more examples, the first prediction type includes coding vertices (e.g., filled vertices) by unidirectional prediction (e.g., thick arrows). For example, as illustrated in FIG. 9, vertices 902A-902D are coded according to the first prediction type. In one or more examples, any combination (e.g., weighted sum) of previous coded filled vertices can be used for predicting the current filled vertex. For example, as illustrated in FIG. 9, vertex 902A is used for predicting vertex 902B. In one or more examples, the spacing between filled vertices may be fixed for both the encoder and the decoder, or may be signaled within the bitstream at different levels. For example, different spacings may be used for different frames or different charts.

[0059] According to some embodiments, the second prediction method includes applying bidirectional prediction, and the prediction can be derived from two coded vertices. For example, as illustrated in FIG. 9, vertices 904A-904H (e.g., dashed vertices) are predicted based on the second prediction method. The prediction of the dashed vertices can be based on filled vertices and / or dashed vertices. In one or more examples, any combination (e.g., weighted sum) of previous coded vertices can be used for predicting the current dashed vertex. In one or more examples, the last group of dashed vertices (e.g., 904G and 904H) can be predicted from the first filled vertex (e.g., 902A). As illustrated in FIG. 9, the dashed vertices can be predicted from the linear interpolation of two adjacent filled vertices in opposite directions. In one or more examples, any interpolation method can be used to derive the prediction, and the distance between two vertices can be defined by the number of vertices between them.

[0060] According to one or more embodiments, after the current coordinates (e.g., the current vertex) are predicted, the prediction residual of the sampled coordinates can be derived as follows.

Number

[0061] The prediction residual may be coded by one or more different methods. In one or more embodiments, the prediction residual can be coded by fixed-length coding. The bit length may be coded in a high-level syntax table for all patches, or may be coded differently for each patch. In one or more embodiments, the prediction residual can be coded by exponential Golomb coding.

[0062] In one or more embodiments, the prediction residual can be coded by unary coding. In one or more embodiments, the prediction residual can be coded by the syntax elements shown in Table 1.

[0063]

Table 1

[0064] In one or more examples, the variable prediction_residual_sign can be coded by bypass coding. In one or more examples, the variable prediction_residual_eq0 can specify whether the prediction residual is equal to 0. In one or more examples, the variable prediction_residual_sign can specify the sign bit of the prediction residual. In one or more examples, the variable prediction_residual_abs_eq1 can specify whether the absolute value of the prediction residual is 1. In one or more examples, the variable prediction_residual_abs_eq2 can specify whether the absolute value of the prediction residual is 2. In one or more examples, the variable prediction_residual_abs_minus3 can specify the absolute value of the prediction residual -3.

[0065] According to one or more embodiments, the prediction residual can be coded by the syntax elements shown in Table 2.

[0066] [Table 2]

[0067] In one or more examples, the variable prediction_residual_sign may be coded by arithmetic coding and different contexts may be used. For example, the sign of the prediction residual of the previous coded boundary vertex may be used as a context. In one or more examples, the variable prediction_residual_eq0 may specify whether the prediction residual is equal to 0. In one or more examples, the variable prediction_residual_sign may specify the sign bit of the prediction residual. In one or more examples, the variable prediction_residual_abs_eq1 may specify whether the absolute value of the prediction residual is 1. In one or more examples, the variable prediction_residual_abs_eq2 may specify whether the absolute value of the prediction residual is 2. In one or more examples, prediction_residual_abs_minus3 may specify the absolute value of the prediction residual -3.

[0068] According to one or more embodiments, on the decoder side, the prediction residual of the sampled UV coordinates can be derived from the above-described syntax elements. For example, the U coordinate can be restored by the following steps.

Number

[0069] As will be understood by those skilled in the art, the V coordinate can be restored using the same or similar steps as described above for restoring the U coordinate.

[0070] According to one or more embodiments, the prediction of the UV boundary coordinates can also be provided from the previously coded mesh frame. In the case of different sampling rates between what is applied to the current mesh frame and what is within the previous mesh frame, the predictor can be inverse quantized to its original value and then quantized according to the current sampling rate.

[0071] According to one or more embodiments, when an interpolation-based hierarchical prediction method is applied as shown in FIG. 9, different residual coding methods may be applied to vertices in different layers, and any combination of the aforementioned residual coding methods may be used.

[0072] In one or more examples, the prediction residual of the UV boundary coordinates may be transformed before entropy coding. For example, various transfer functions such as discrete cosine / sine transform, wavelet transform, (fast) Fourier transform, spline transform, etc. may be employed.

[0073] According to one or more embodiments, the transformation may be applied to the partial prediction residual based on the different prediction methods used. For example, when an interpolation-based hierarchical prediction method is applied as illustrated in FIG. 9, the transformation may be applied to the prediction residual of the vertices of the dashed line in the higher prediction layer having bidirectional prediction.

[0074] According to one or more embodiments, a fixed-point 1-D transformation may be applied to the prediction residual. For example, a 16-point 1-D transformation may be applied to the prediction residual. In this case, the residuals may be grouped into 16 samples so that each group can use the transformation.

[0075] According to one or more embodiments, a fixed-point 2-D transformation may be applied to the prediction residual. For example, a 4×4 2-D transformation may be applied to the prediction residual. In this case, the residuals may be grouped into 16 samples (e.g., in the form of an 8×8 residual block) so that each group can use the transformation.

[0076] According to one or more embodiments, a transform having a plurality of kernels can be applied to a prediction residual. For example, a 4×4 2-D transform, either DCT or DST, may be applied to the prediction residual. In this case, the residual can be grouped into 16 samples (in the form of an 8×8 residual block) such that each group can use the transform. Further, the encoder can send a transform type index selected for each group so that the decoder can decode the correct transform coefficients associated with the transform type.

[0077] According to one or more embodiments, a combination of fixed-point 1-D transforms or 2-D transforms can be applied to a prediction residual. For example, 4×4, 8×4, and 4×8 2-D transforms may be applied to the prediction residual. In this case, the residual can be grouped into either 16 samples or 32 samples (e.g., in the form of a 2-D residual block) such that each group can use the transform. The encoder can send a transform type index selected for each group so that the decoder can decode the correct number of transform coefficients associated with the transform.

[0078] FIG. 10 illustrates a process 1000 for reconstructing boundary vertices based on a compressed 2D mesh according to one or more embodiments. The process can start from operation S1002 where a coded video bitstream including the compressed 2D mesh is received. The compressed 2D mesh can correspond to the compression of the sampled UV plane 800 (FIG. 8). The process proceeds to operation S1004 where the current vertices included in the compressed 2D mesh are predicted based on an interpolation-based hierarchical prediction method. For example, the interpolation-based hierarchical prediction method can predict the current vertices based on the first prediction type or the second prediction type described above.

[0079] The process proceeds to operation S1006, where a prediction residual associated with the predicted current vertex is derived. For example, the prediction residual may be derived according to any one of Equations (3) to (6), or according to the syntax elements specified in any one of Table 1 or Table 2. Further, the prediction residual may be derived based on an N-point 1-D transform or an N×N 2-D transform. The process proceeds to operation S1008, where a boundary vertex associated with the 3D volume object is reconstructed based on the predicted current vertex and the derived prediction residual. For example, referring to FIG. 8, the process of FIG. 10 may predict a sampling 2D coordinate (e.g., the current vertex) V1’(u1’, v1’), derive a corresponding offset, and reconstruct the original 2D coordinate (e.g., the boundary vertex) V1 on the decoder side.

[0080] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 11 shows a computer system 1100 suitable for implementing a particular embodiment of the present disclosure.

[0081] The computer software may be coded using any suitable machine code or computer language that can be subjected to mechanisms such as assembly, compilation, and linking to create code including instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or through interpretation, microcode execution, etc.

[0082] The instructions may be executed on various types of computers or computer components including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, etc.

[0083] The components shown in FIG. 11 for computer system 1100 are examples and are not intended to suggest any limitation as to the use or functionality of the computer software implementing embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components illustrated in the non-limiting embodiments of computer system 1100.

[0084] Computer system 1100 may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, applause, etc.), visual input (gestures, etc.), olfactory input (not shown). The human interface device may also be used to capture specific media that is not necessarily directly related to conscious input by a human, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images, obtained from a still image camera, etc.), video (two-dimensional video, three-dimensional video including stereoscopic video, etc.).

[0085] The input human interface device may include one or more (only one of each is shown) of keyboard 1101, mouse 1102, trackpad 1103, touch screen 1110, data glove, joystick 1105, microphone 1106, scanner 1107, camera 1108.

[0086] Computer system 1100 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by touch screen 1110, data glove, or joystick 1105, although there may be tactile feedback devices that do not function as input devices). For example, such devices may include audio output devices (such as speaker 1109, headphones (not shown)), visual output devices (each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which may be capable of outputting more than three dimensions, such as two-dimensional visual output or three-dimensional output by means such as stereoscopic output, including screens 1110 such as CRT screens, LCD screens, plasma screens, OLED screens, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0087] Computer system 1100 may also include storage devices and their associated media that are accessible to humans, such as optical media including CD / DVD ROM / RW 1120 having a CD / DVD or similar medium 1121, thumb drive 1122, removable hard drive or solid state drive 1123, legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0088] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0089] The computer system 1100 may also include an interface to one or more communication networks. The network may be wireless, wired, or optical. The network may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., wired or wireless wide area digital networks for television including cable television, satellite television, and terrestrial television, vehicular and industrial including CANBus, etc. Certain networks generally require an external network interface adapter connected to a specific general-purpose data port or peripheral bus 1149 (e.g., a USB port of the computer system 1100), and other networks are generally integrated into the core of the computer system 1100 by connection to the system bus as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system 1100 can communicate with other entities. Such communication may be only unidirectional reception (e.g., television broadcast), only unidirectional transmission (e.g., CANbus to a specific CANbus device), or bidirectional, for example, to other computer systems using a local or wide area digital network. Such communication may also include communication to a cloud computing environment 1155. Specific protocols and protocol stacks may be used for each of the networks and network interfaces as described above.

[0090] The aforementioned human interface device, human-accessible storage device, and network interface 1154 may be attached to the core 1140 of the computer system 1100.

[0091] The core 1140 may include one or more central processing units (CPUs) 1141, a graphics processing unit (GPU) 1142, a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) 1143, a hardware accelerator 1144 for specific tasks, etc. These devices can be connected via a system bus 1148 together with a read-only memory (ROM) 1145, a random access memory 1146, an internal mass storage such as a hard drive or SSD that is not accessible to internal users 1147. In some computer systems, the system bus 1148 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus 1148 or may be attached via a peripheral bus 1149. The architecture of the peripheral bus includes PCI, USB, etc. A graphics adapter 1150 may be included in the core 1140.

[0092] The CPU 1141, GPU 1142, FPGA 1143, and accelerator 1144 may be combined to execute specific instructions that can constitute the aforementioned computer code. The computer code may be stored in the ROM 1145 or the RAM 1146. Temporary data may also be stored in the RAM 1146, while persistent data may be stored, for example, in the internal mass storage 1147. Fast storage and retrieval to any of the memory devices may be enabled by the use of a cache memory that may be closely associated with one or more CPUs 1141, GPUs 1142, the mass storage 1147, the ROM 1145, the RAM 1146, etc.

[0093] The computer-readable medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and configured for the purposes of the present disclosure or may be of the types available to those skilled in the art of computer software technology.

[0094] By way of example and not limitation, a computer system 1100 having an architecture, specifically a core 1140, may provide functionality as a result of (one or more) processors (including, for example, a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media may be associated with the user-accessible mass storage as introduced above, as well as media associated with specific storage of the core 1140 that is non-transitory in nature, such as the mass storage 1147 inside the core and the ROM 1145. The software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core 1140. The computer-readable media may include one or more memory devices or chips, depending on specific requirements. The software may cause the core 1140, specifically the processors (including a CPU, GPU, FPGA, etc.) therein, to define data structures stored in the RAM 1146 and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally, or alternatively, the computer system may function as a result of logic wired or otherwise embodied in a circuit (e.g., an accelerator 1144) that operates to execute specific processes or specific portions of specific processes described herein, instead of or in conjunction with software. References to software may, as necessary, include logic, and vice versa. References to computer-readable media may, as necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0095] Although the present disclosure describes several non-limiting embodiments, there are modifications, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Thus, those skilled in the art will understand that although not explicitly illustrated or described herein, numerous systems and methods can be devised that embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

[0096] The above disclosure also encompasses the embodiments listed below.

[0097] (1) A method performed by at least one processor in a decoder, the method comprising receiving a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to the surface of a three-dimensional (3D) volume object; predicting a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction scheme using at least one sampled vertex included in the compressed 2D mesh; deriving a prediction residual associated with the predicted current vertex; and reconstructing a boundary vertex associated with the 3D volume object based on the predicted current vertex and the derived prediction residual.

[0098] (2) The interpolation-based hierarchical prediction scheme of feature (1) includes a first prediction type for predicting a vertex included in the compressed 2D mesh by applying a unidirectional prediction using at least one sampled vertex, and a second prediction type for predicting a vertex included in the compressed 2D mesh by applying a bidirectional prediction using at least one sampled vertex and other sampled vertices included in the compressed 2D mesh.

[0099] (3) The method of feature (2), wherein the current vertex is predicted according to a first prediction type in which at least one sampled vertex is predicted before the current vertex.

[0100] (4) The method according to feature (2), wherein the current vertex is predicted according to a second prediction type.

[0101] (5) The method according to feature (4), wherein at least one sampled vertex is predicted according to a first prediction type, and the other sampled vertices are predicted according to the first prediction type.

[0102] (6) The method according to feature (4), wherein at least one sampled vertex is predicted according to a first prediction type, and the other sampled vertices are predicted according to a second prediction type.

[0103] (7) The method according to feature (4), wherein at least one sampled vertex and the other sampled vertices are in opposite directions.

[0104] (8) The method according to feature (4), wherein the distance between at least one sampled vertex and the other sampled vertices is specified in the syntax elements included in the coded video bitstream.

[0105] (9) The method according to any one of features (1) to (8), wherein the prediction residual is derived according to an N-point 1-D transform.

[0106] (10) The method according to any one of features (1) to (8), wherein the prediction residual is derived according to an N×N 2-D transform.

[0107] (11) At least one memory configured to store program code, and at least one processor configured to read the program code and operate as commanded by the program code, wherein the program code causes the at least one processor to receive a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to the surface of a three-dimensional (3D) volume object (reception code), cause the at least one processor to predict a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction method using at least one sampled vertex included in the compressed 2D mesh (prediction code), cause the at least one processor to derive a prediction residual associated with the predicted current vertex (derivation code), and cause the at least one processor to reconstruct a boundary vertex associated with the 3D volume object based on the predicted current vertex and the derived prediction residual (reconstruction code); and the at least one processor including the at least one processor, a decoder.

[0108] (12) The decoder according to feature (11), wherein the interpolation-based hierarchical prediction method includes: (i) a first prediction type for predicting a vertex included in the compressed 2D mesh by applying unidirectional prediction using at least one sampled vertex; and (ii) a second prediction type for predicting a vertex included in the compressed 2D mesh by applying bidirectional prediction using at least one sampled vertex and other sampled vertices included in the compressed 2D mesh.

[0109] (13) The decoder according to feature (12), wherein the current vertex is predicted according to the first prediction type in which at least one sampled vertex is predicted before the current vertex.

[0110] (14) The decoder according to feature (12), wherein the current vertex is predicted according to the second prediction type.

[0111] (15) The decoder according to feature (14), wherein at least one sampled vertex is predicted according to a first prediction type, and other sampled vertices are predicted according to the first prediction type.

[0112] (16) The decoder according to feature (14), wherein at least one sampled vertex is predicted according to a first prediction type, and other sampled vertices are predicted according to a second prediction type.

[0113] (17) The decoder according to feature (14), wherein at least one sampled vertex and other sampled vertices are in opposite directions.

[0114] (18) The decoder according to feature (14), wherein the distance between at least one sampled vertex and other sampled vertices is specified in a syntax element included in the coded video bitstream.

[0115] (19) The decoder according to any one of features (11) to (18), wherein the prediction residual is derived according to an N-point 1-D transform.

[0116] (20) A non-transitory computer-readable medium storing instructions, which when executed by at least one processor in a decoder, cause the at least one processor to receive a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to the surface of a three-dimensional (3D) volume object, predict a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction method using at least one sampled vertex included in the compressed 2D mesh, derive a prediction residual associated with the predicted current vertex, and reconstruct a boundary vertex associated with the 3D volume object based on the predicted current vertex and the derived prediction residual.

Description of Reference Numerals

[0117] 100 Communication system 110 Terminal 120 Terminal 130 Terminal 140 Terminal 150 Network 200 Streaming system 201 Video source 202 Stream 203 Encoder 204 Encoded video bitstream 205 Streaming server 206 Streaming client 209 Video bitstream 210 Video decoder 211 Output video sample stream 212 Display 213 Capture subsystem 300 Framework for dynamic mesh compression and mesh reconstruction 301 Encoder 305 Input mesh 310 Mesh with UV atlas 315 Occupancy map 320 Geometry map 325 Attribute map 330 Metadata 335 Decoded occupancy map 340 Decoded geometry map 345 Decoded attribute map 350 Decoded metadata 351 Decoder 360 Reconstructed mesh 400 Mesh segment 402 2D chart 404 2D chart 500 3D mesh segment 502A 2D chart 502B 2D chart 600 2D UV Atlas 800 UV Plane 902A Vertex 902B Vertex 902C Vertex 902D Vertex 904A Vertex 904B Vertex 904C Vertex 904D Vertex 904E Vertex 904F Vertex 904G Vertex 904H Vertex 1000 Process for Reconstructing Boundary Vertices Based on Compressed 2D Mesh 1100 Computer System 1101 Keyboard 1102 Mouse 1103 Trackpad 1105 Joystick 1106 Microphone 1107 Scanner 1108 Camera 1109 Speaker 1110 Touch Screen 1120 CD / DVD ROM / RW 1121 CD / DVD or Similar Medium 1122 Thumb Drive 1123 Removable Hard Drive or Solid State Drive 1140 Core of Computer System 1141 Central Processing Unit (CPU) 1142 Graphics Processing Unit (GPU) 1143 Field Programmable Gate Array (FPGA) 1144 Hardware Accelerator 1145 Read Only Memory (ROM) 1146 Random Access Memory 1147 Internal Mass Storage 1148 System Bus 1149 Peripheral Bus 1150 Graphics Adapter 1154 Network Interface 1155 Cloud Computing Environment

Claims

Claim 1 A method executed by at least one processor in a decoder, the method comprising: receiving a coded video bitstream including a compressed two-dimensional (2D) mesh corresponding to a surface of a three-dimensional (3D) volume object; predicting a current vertex included in the compressed 2D mesh based on an interpolation-based hierarchical prediction method using at least one sampled vertex included in the compressed 2D mesh; deriving a prediction residual associated with the predicted current vertex; reconstructing boundary vertices associated with the 3D volume object based on the predicted current vertex and the derived prediction residual. Claim 2 The method according to claim 1, wherein the interpolation-based hierarchical prediction method includes: (i) a first prediction type for predicting a vertex included in the compressed 2D mesh by applying unidirectional prediction using the at least one sampled vertex; and (ii) a second prediction type for predicting the vertex included in the compressed 2D mesh by applying bidirectional prediction using the at least one sampled vertex and other sampled vertices included in the compressed 2D mesh. Claim 3 The method according to claim 2, wherein the current vertex is predicted according to the first prediction type in which the at least one sampled vertex is predicted before the current vertex. Claim 4 The method according to claim 2, wherein the current vertex is predicted according to the second prediction type. Claim 5 The method according to claim 4, wherein the at least one sampled vertex is predicted according to the first prediction type, and the other sampled vertices are predicted according to the first prediction type. Claim 6 The method according to claim 4, wherein the at least one sampled vertex is predicted according to the first prediction type, and the other sampled vertices are predicted according to the second prediction type. Claim 7 The method according to claim 4, wherein the at least one sampled vertex and the other sampled vertices are in opposite directions. Claim 8 The method according to claim 4, wherein an interval between the at least one sampled vertex and the other sampled vertex is specified in a syntax element included in the coded video bitstream.

9. The method according to claim 1, wherein the prediction residual is derived according to an N-point 1-D transform.

10. The method according to claim 1, wherein the prediction residual is derived according to an N×N 2-D transform.

11. A decoder configured to perform the method according to any one of claims 1 to 10.

12. A computer program for causing a computer to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • A scalable compression method for time-consistent 3D mesh sequences.

    JP2010527523A

  • A method, an apparatus and a computer program product for volumetric video encoding and decoding

    WO2021136878A1