Mesh Compression Using Estimated Texture Coordinates

By estimating texture coordinates and using reversible or irreversible codecs, the method optimizes mesh coding efficiency, addressing data resource challenges in 3D mesh encoding and decoding.

JP7698052B2Active Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023553378
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2022-10-26
Publication Date
2025-06-24
Estimated Expiration
2042-10-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently encoding and decoding 3D mesh data, particularly in terms of texture coordinates, leading to high data resource consumption and inefficiencies in data storage and transmission.

Method used

The proposed solution involves a processing circuit that decodes 3D mesh frames by estimating texture coordinates using parameterization based on 3D coordinates and connection information, and employs reversible or irreversible codecs to encode and decode mesh data, optimizing texture coordinate representation.

Benefits of technology

This approach reduces the data required for texture coordinate signaling, enhancing mesh coding efficiency and reducing the amount of data needed for storage and transmission, thereby improving immersive 3D media experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698052000001
    Figure 0007698052000001
  • Figure 0007698052000002
    Figure 0007698052000002
  • Figure 0007698052000003
    Figure 0007698052000003
Patent Text Reader

Abstract

Aspects of the present disclosure provide a method and apparatus for mesh coding (encoding and / or decoding). In some examples, the apparatus for coding a mesh includes a processing circuit. The processing circuit decodes three-dimensional (3D) coordinates of vertices in a first 3D mesh frame and vertex connectivity information from a bitstream carrying the first 3D mesh frame. The first 3D mesh frame represents a surface of an object with polygons. The processing circuit estimates texture coordinates associated with the vertices and decodes a texture map of the first 3D mesh frame from the bitstream. The texture map includes a first one or more 2D charts having 2D vertices with texture coordinates. The processing circuit reconstructs the first 3D mesh frame based on the 3D coordinates of the vertices, the vertex connectivity information, the texture map, and the texture coordinates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority to U.S. Provisional Application No. 63 / 298,106, filed on January 10, 2022, entitled "Mesh Compression with Deduced Texture Coordinates", and to U.S. Patent Application No. 17 / 969,570, filed on October 19, 2022, entitled "MESH COMPRESSION WITH DEDUCED TEXTURE COORDINATES". The disclosure of the prior application is hereby incorporated by reference in its entirety.

[0002] This disclosure generally describes embodiments related to mesh coding.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the context of the present disclosure. The research of the present inventors, to the extent described in this background art section and aspects of the description that may not be recognized as prior art at the time of filing, are not expressly or implicitly recognized as prior art to the present disclosure.

[0004] Various techniques have been developed for capturing and representing the world, such as objects in the three - dimensional (3D) space, the environment of the world, etc. A 3D representation of the world can enable more immersive interactions and communications. In some examples, point clouds and meshes can be used as 3D representations of the world.

Summary of the Invention

Means for Solving the Problems

[0005] Aspects of the present disclosure provide methods and apparatuses for mesh coding (encoding and / or decoding). In some examples, an apparatus for coding a mesh includes a processing circuit. The processing circuit decodes the three-dimensional (3D) coordinates of vertices within a first 3D mesh frame and connection information of the vertices from a bitstream carrying the first 3D mesh frame. The first 3D mesh frame represents the surface of an object with polygons. The processing circuit estimates texture coordinates associated with the vertices and decodes a texture map of the first 3D mesh frame from the bitstream. The texture map includes one or more first 2D charts having 2D vertices with texture coordinates. The processing circuit reconstructs the first 3D mesh frame based on the 3D coordinates of the vertices, the connection information of the vertices, the texture map, and the texture coordinates.

[0006] In some examples, the processing circuit decodes the 3D coordinates of the vertices and the connection information of the vertices using a reversible codec. In some examples, the processing circuit decodes the 3D coordinates of the vertices and the connection information of the vertices using a non-reversible codec.

[0007] In some examples, to estimate texture coordinates associated with the vertices, the processing circuit performs parameterization according to the 3D coordinates and connection information of the vertices to determine texture coordinates associated with the vertices.

[0008] In some examples, to perform parameterization, the processing circuit divides the polygons of a first 3D mesh frame into one or more first 2D charts, and packs the one or more first 2D charts into a 2D map. The processing circuit performs a temporal alignment that aligns the one or more first 2D charts with one or more second 2D charts associated with a second 3D mesh frame. The first 3D mesh frame and the second 3D mesh frame are frames within a 3D mesh sequence. The processing circuit determines texture coordinates from the one or more first 2D charts using the temporal alignment.

[0009] In some examples, to divide the polygons, the processing circuit divides the polygons according to the normal values associated with the polygons.

[0010] In some examples, the processing circuit performs the temporal alignment according to at least one of a scale-invariant metric, a rotation-invariant metric, a translation-invariant metric, and / or an affine transformation-invariant metric.

[0011] In some examples, the processing circuit performs the temporal alignment according to at least one of the center of the chart calculated based on the 3D coordinates associated with the chart, the average depth of the chart, the weighted average texture value of the chart, and / or the weighted average attribute value of the chart.

[0012] In some examples, the processing circuit decodes a flag indicating the enabling of texture coordinate derivation. The flag is at least one of a sequence-level flag, a frame group-level flag, and a frame-level flag.

[0013] In some examples, to estimate texture coordinates associated with vertices, the processing circuit decodes a flag indicating inheritance of texture coordinates, decodes an index indicating a 3D mesh frame selected from a list of decoded 3D mesh frames, and inherits texture coordinates from the selected 3D mesh frame.

[0014] In some examples, the processing circuit decodes a flag associated with a portion of a texture map, the flag indicating whether UV coordinates associated with the portion of the texture map are inherited from the decoded mesh frame or derived by parameterization.

[0015] In some examples, the processing circuit decodes an index of a set of key vertices among the vertices, and parameterization starts from the set of key vertices.

[0016] In some examples, the processing circuit decodes an index indicating a parameterization method selected from a list of parameterization method candidates.

[0017] Aspects of the present disclosure also provide a non-transitory computer-readable medium that stores instructions which, when executed by a computer, cause the computer to execute any one or combination of methods for mesh coding.

[0018] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0020] Aspects of the present disclosure provide techniques in the field of three-dimensional (3D) media processing.

[0021] Technological developments in 3D media processing, such as advancements in 3D capture, 3D modeling, and 3D rendering, have facilitated the widespread presence of 3D media content across several platforms and devices. In one example, a baby's first steps are captured on one continent, and media technology enables grandparents on another continent to view, (and perhaps interact with) and enjoy an immersive experience with the baby. According to one aspect of the present disclosure, to improve the immersive experience, 3D models have become increasingly sophisticated, and the creation and consumption of 3D models occupy a significant amount of data resources such as data storage, data transmission resources, and the like.

[0022] According to some aspects of the present disclosure, point clouds and meshes can be used as 3D models for representing immersive content.

[0023] A point cloud may generally refer to a set of points in 3D space, where each point has associated attributes such as color, material properties, texture information, intensity attributes, reflectivity attributes, motion-related attributes, modality attributes, and various other attributes. The point cloud can be used to reconstruct an object or scene as such a configuration of points.

[0024] A mesh of an object (also called a mesh model) can include polygons that describe the surface of the object. Each polygon can be defined by the vertices of the polygon in 3D space and information on how the vertices are connected to the polygon. The information on how the vertices are connected is called connectivity information. In some examples, the mesh can also include attributes such as color and normal associated with the vertices.

[0025] According to some aspects of the present disclosure, some coding tools for point cloud compression (PCC) can be used for mesh compression. For example, the mesh can be remeshed to generate a new mesh whose connectivity information of the new mesh can be inferred. The vertices of the new mesh, and the attributes associated with the vertices of the new mesh, can be regarded as points in a point cloud and can be compressed using a PCC codec.

[0026] A point cloud can be used to reconstruct an object or scene as a configuration of points. The points can be captured using multiple cameras, depth sensors, or lidars in various settings and can be composed of thousands to up to tens of billions of points to realistically represent the reconstructed scene or object. A patch may generally refer to a continuous subset of the surface described by the point cloud. In one example, a patch includes points having surface normal vectors that deviate from each other by less than a threshold amount.

[0027] PCC can be performed according to various methods, such as a geometry-based method called G-PCC and a video coding-based method called V-PCC. According to some aspects of the present disclosure, G-PCC is a purely geometry-based approach that directly encodes 3D geometry and has few elements in common with video coding, while V-PCC is highly based on video coding. For example, V-PCC can map the points of a 3D cloud to the pixels of a 2D grid (image). The V-PCC method can utilize a general-purpose video codec for point cloud compression. The PCC codec (encoder / decoder) in the present disclosure can be a G-PCC codec (encoder / decoder) or a V-PCC codec.

[0028] According to one aspect of the present disclosure, the V-PCC method can compress the geometry, occupancy, and texture of a point cloud as three separate video sequences using an existing video codec. The additional metadata required to interpret the three video sequences is compressed separately. A small portion of the entire bitstream is metadata, which can be efficiently encoded / decoded using, for example, a software implementation. Most of the information is processed by the video codec.

[0029] FIG. 1 shows a block diagram of a communication system (100) in some examples. The communication system (100) includes a plurality of terminal devices that can communicate with each other via, for example, a network (150). For example, the communication system (100) includes a pair of terminal devices (110) and (120) interconnected via a network (150). In the example of FIG. 1, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of point cloud data. For example, the terminal device (110) can compress a point cloud (e.g., points representing a structure) captured by a sensor (105) connected to the terminal device (110). The compressed point cloud can be transmitted to another terminal device (120) via the network (150) in the form of, for example, a bitstream. The terminal device (120) can receive the compressed point cloud from the network (150), decompress the bitstream to reconstruct the point cloud, and appropriately display the reconstructed point cloud. Unidirectional data transmission can be common in media serving applications and the like.

[0030] In the example of FIG. 1, the terminal devices (110) and (120) may be shown as a server and a personal computer, but the principles of the present disclosure need not be so limited. Embodiments of the present disclosure find use in laptop computers, tablet computers, smartphones, gaming terminals, media players, and / or dedicated three-dimensional (3D) devices. The network (150) represents any number of networks that transmit compressed point clouds between the terminal device (110) and the terminal device (120). The network (150) can include, for example, a wired communication (wired) network and / or a wireless communication network. The network (150) may exchange data over a circuit-switched and / or packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, the Internet, and the like.

[0031] FIG. 2 shows a block diagram of a streaming system (200) in some examples. The streaming system (200) is for the use of point clouds. The disclosed subject matter may equally apply to other point cloud-enabled applications such as 3D telepresence applications, virtual reality applications, and the like.

[0032] The streaming system (200) may include a capture subsystem (213). The capture subsystem (213) may include a point cloud source (201), such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, or a graphics generation component that generates an uncompressed point cloud, such as in software that generates an uncompressed point cloud (202). In one example, the point cloud (202) includes points captured by a 3D camera. The point cloud (202) is shown as a thick line to emphasize its high data volume compared to the compressed point cloud (204) (bitstream of the compressed point cloud). The compressed point cloud (204) may be generated by an electronic device (220) that includes an encoder (203) coupled to the point cloud source (201). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The compressed point cloud (204) (or bitstream of the compressed point cloud (204)), shown as a thin line to emphasize its lower data volume compared to the stream of the point cloud (202), may be stored in the streaming server (205) for future use. One or more streaming client subsystems, such as the client subsystems (206) and (208) of FIG. 2, may access the streaming server (205) to retrieve copies (207) and (209) of the compressed point cloud (204). The client subsystem (206) may include, for example, a decoder (210) within an electronic device (230). The decoder (210) decodes the incoming copy (207) of the compressed point cloud and creates an output stream of the reconstructed point cloud (211) that may be rendered on the rendering device (212).

[0033] Note that the electronic devices (220) and (230) may include other components (not shown). For example, the electronic device (220) may include a decoder (not shown), and the electronic device (230) may also include an encoder (not shown).

[0034] In some streaming systems, the compressed point clouds (204), (207), and (209) (e.g., the bitstream of the compressed point cloud) can be compressed according to a specific standard. In some examples, video coding standards are used for compressing the point cloud. Examples of such standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and the like.

[0035] FIG. 3 shows a block diagram of a V-PCC encoder (300) for encoding a point cloud frame according to some embodiments. In some embodiments, the V-PCC encoder (300) can be used in the communication system (100) and the streaming system (200). For example, the encoder (203) can be configured and operate in the same manner as the V-PCC encoder (300).

[0036] The V-PCC encoder (300) receives a point cloud frame as an uncompressed input and generates a bitstream corresponding to the compressed point cloud frame. In some embodiments, the V-PCC encoder (300) may receive the point cloud frame from a point cloud source such as the point cloud source (201).

[0037] In the example of FIG. 3, the V-PCC encoder (300) includes a patch generation module (306), a patch packing module (308), a geometry image generation module (310), a texture image generation module (312), a patch information module (304), an occupancy map module (314), a smoothing module (336), image padding modules (316) and (318), a group expansion module (320), video compression modules (322), (323) and (332), an auxiliary patch information compression module (338), an entropy compression module (334), and a multiplexer (324).

[0038] According to one aspect of the present disclosure, a V-PCC encoder (300) converts a 3D point cloud frame into an image-based representation, along with some metadata (e.g., occupancy map and patch information) used to convert the compressed point cloud into a decompressed point cloud. In some examples, the V-PCC encoder (300) can convert the 3D point cloud frame into a geometry image, a texture image, and an occupancy map, and then use video coding techniques to encode the geometry image, the texture image, and the occupancy map into a bitstream. Generally, a geometry image is a 2D image having pixels filled with geometry values associated with points projected onto the pixels, and the pixels filled with geometry values can be called geometry samples. A texture image is a 2D image having pixels filled with texture values associated with points projected onto the pixels, and the pixels filled with texture values can be called texture samples. An occupancy map is a 2D image having pixels filled with values indicating whether a patch is occupied or not.

[0039] The patch generation module (306) segments the point cloud into a set of patches (e.g., the patches are defined as continuous subsets of the surface described by the point cloud) that may or may not overlap, such that each patch can be described by a depth field with respect to a plane in 2D space. In some embodiments, the patch generation module (306) aims to decompose the point cloud into the minimum number of patches with smooth boundaries while minimizing the reconstruction error.

[0040] In some examples, the patch information module (304) can collect patch information indicating the size and shape of the patches. In some examples, the patch information is packed into an image frame and then encoded by an auxiliary patch information compression module (338) to generate compressed auxiliary patch information.

[0041] In some examples, the patch packing module (308) maps the extracted patches onto a two-dimensional (2D) grid while minimizing unused space and ensuring that a unique patch is associated with each of the M×M (e.g., 16×16) blocks of the grid. Efficient patch packing can directly affect compression efficiency either by minimizing unused space or by ensuring temporal coherence.

[0042] The geometry image generation module (310) can generate a 2D geometry image associated with the geometry of the point cloud at a given patch location. The texture image generation module (312) can generate a 2D texture image associated with the texture of the point cloud at a given patch location. The geometry image generation module (310) and the texture image generation module (312) use the 3D-to-2D mapping computed during the packing process to store the geometry and texture of the point cloud as images. To better handle the case where multiple points are projected onto the same sample, each patch is projected onto two images called layers. In one example, the geometry image is represented by a WxH monochrome frame in the YUV420-8-bit format. To generate the texture image, the texture generation procedure uses the reconstructed / smoothed geometry to calculate the color associated with the resampled points.

[0043] The occupancy map module (314) can generate an occupancy map that describes padding information for each unit. For example, the occupancy image includes a binary map for each cell of the grid indicating whether the cell belongs to free space or to the point cloud. In one example, the occupancy map uses binary information that describes for each pixel whether the pixel is padded or not. In other examples, the occupancy map uses binary information that describes for each block of pixels whether the block of pixels is padded or not.

[0044] The occupancy map generated by the occupancy map module (314) can be compressed using reversible coding or irreversible coding. When reversible coding is used, the entropy compression module (334) is used to compress the occupancy map. When irreversible coding is used, the video compression module (332) is used to compress the occupancy map.

[0045] Note that the patch packing module (308) may leave some empty space between the 2D patches packed within the image frame. The image padding modules (316) and (318) can fill (referred to as padding) the empty space to generate an image frame suitable for 2D video and image codecs. Image padding is also called background filling which can fill the unused space with redundant information. In some examples, good background filling increases the bitrate minimally but does not introduce significant coding distortion around the patch boundaries.

[0046] The video compression modules (322), (323), and (332) can encode 2D images such as the padded geometry image, the padded texture image, and the occupancy map based on suitable video coding standards such as HEVC, VVC. In one example, the video compression modules (322), (323), and (332) are individual components that operate separately. Note that in other examples, the video compression modules (322), (323), and (332) can be implemented as a single component.

[0047] In some examples, the smoothing module (336) is configured to generate a smoothed image of the reconstructed geometry image. The smoothed image can be provided to texture image generation (312). Next, the texture image generation (312) may adjust the generation of the texture image based on the reconstructed geometry image. For example, if there is some distortion in the patch shape (e.g., geometry) during encoding and decoding, that distortion may be considered when generating the texture image to correct the distortion of the patch shape.

[0048] In some embodiments, the group expansion (320) is configured to pad pixels around the object boundary with redundant low-frequency content to improve the coding gain and the visual quality of the reconstructed point cloud.

[0049] The multiplexer (324) can multiplex the compressed geometry image, the compressed texture image, the compressed occupancy map, and the compressed auxiliary patch information into a compressed bitstream.

[0050] FIG. 4 shows a block diagram of a V-PCC decoder (400) for decoding a compressed bitstream corresponding to a point cloud frame in some examples. In some examples, the V-PCC decoder (400) can be used in the communication system (100) and the streaming system (200). For example, the decoder (210) can be configured to operate in a similar manner to the V-PCC decoder (400). The V-PCC decoder (400) receives the compressed bitstream and generates a reconstructed point cloud based on the compressed bitstream.

[0051] In the example of FIG. 4, the V-PCC decoder (400) includes a demultiplexer (432), video decompression modules (434) and (436), an occupancy map decompression module (438), an auxiliary patch information decompression module (442), a geometry reconstruction module (444), a smoothing module (446), a texture reconstruction module (448), and a color smoothing module (452).

[0052] The demultiplexer (432) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometry image, a compressed occupancy map, and compressed auxiliary patch information.

[0053] The video decompression modules (434) and (436) can decode the compressed images according to appropriate standards (such as HEVC, VVC, etc.) and output the decompressed images. For example, the video decompression module (434) decodes the compressed texture image and outputs the decompressed texture image, and the video decompression module (436) decodes the compressed geometry image and outputs the decompressed geometry image.

[0054] The occupancy map decompression module (438) can decode the compressed occupancy map according to appropriate standards (such as HEVC, VVC, etc.) and output the decompressed occupancy map.

[0055] The auxiliary patch information decompression module (442) can decode the compressed auxiliary patch information according to appropriate standards (such as HEVC, VVC, etc.) and output the decompressed auxiliary patch information.

[0056] The geometry reconstruction module (444) can receive the decompressed geometry image and generate the reconstructed point cloud geometry based on the decompressed occupancy map and the decompressed auxiliary patch information.

[0057] The smoothing module (446) can smooth the mismatches at the edges of the patches. The smoothing procedure aims to mitigate potential discontinuities that may occur at the patch boundaries due to compression artifacts. In some embodiments, a smoothing filter may be applied to the pixels located on the patch boundaries to mitigate distortions that may be caused by compression / decompression.

[0058] The texture reconstruction module (448) can determine the texture information of the points within the point cloud based on the decompressed texture image and the smoothed geometry.

[0059] The color smoothing module (452) can smooth the color mismatches. Non-adjacent patches within the 3D space are often packed adjacent to each other within the 2D video. In some examples, the pixel values from non-adjacent patches may be mixed by a block-based video codec. The purpose of color smoothing is to reduce the visible artifacts that appear at the patch boundaries.

[0060] FIG. 5 shows a block diagram of a video decoder (510) in some examples. The video decoder (510) can be used in the V-PCC decoder (400). For example, the video decompression modules (434) and (436), and the occupancy map decompression module (438) can be configured similarly to the video decoder (510).

[0061] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from a compressed image, such as, for example, a coded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510). The parser (520) may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technology or standard and may follow various principles including variable length coding, Huffman coding, arithmetic coding, etc. with or without context dependency. The parser (520) may extract a set of subgroup parameters of at least one subgroup of pixels within the video decoder from the coded video sequence based on at least one parameter corresponding to a group. The subgroups may include Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (520) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the coded video sequence.

[0062] The parser (520) may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory to create symbols (521).

[0063] The reconstruction of the symbols (521) may include a plurality of different units depending on the type of the coded video picture or a part thereof (such as between pictures and within pictures, between blocks and within blocks, etc.) and other factors. How each unit is involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (520). Such a flow of subgroup control information between the parser (520) and the following plurality of units is not depicted for clarity.

[0064] In addition to the functional blocks already described, the video decoder (510) can be conceptually subdivided into several functional units, as described below. In an actual implementation operating under commercial constraints, many of these units can interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0065] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives, from the parser (520), quantized transform coefficients and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. as (one or more) symbols (521). The scaler / inverse transform unit (551) can output a block including sample values that can be input to the aggregator (555).

[0066] In some cases, the output samples of the scaler / inverse transform (551) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) uses surrounding already-reconstructed information fetched from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0067] In other cases, the output samples of the scaler / inverse transform unit (551) may be inter-coded and potentially related to motion-compensated blocks. In such cases, the motion-compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (521) related to the block, these samples can be added to the output of the scaler / inverse transform unit (551) by the aggregator (555) to generate output sample information (in this case, called residual samples or residual signals). The address in the reference picture memory (557) from which the motion-compensation prediction unit (553) fetches the prediction samples can be controlled by, for example, the motion vectors available to the motion-compensation prediction unit (553) in the form of symbols (521) that can have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (557) when exact sub-sample motion vectors are used, a motion vector prediction mechanism, and the like.

[0068] The output samples of the aggregator (555) can undergo various loop-filtering techniques in the loop filter unit (556). The video compression technology is controlled by the parameters included in the coded video sequence (also called the coded video bitstream), and can include in-loop filter techniques available to the loop filter unit (556) as symbols (521) from the parser (520), but can also respond to meta-information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop-filtered sample values.

[0069] The output of the loop filter unit (556) may be a sample stream that can be output to the rendering device and can also be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0070] When a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and the unused current picture buffer can be reallocated before starting the reconstruction of the next coded picture.

[0071] The video decoder (510) can perform a decoding operation according to a predetermined video compression technology in a standard such as ITU-T Rec.H.265. The coded video sequence may conform to the syntax specified by the video compression technology or standard being used in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile documented in the video compression technology. Specifically, a profile can select specific tools from all the tools available in the video compression technology or standard as the only tools available under that profile. Also, for compliance, it may be necessary that the complexity of the coded video sequence is within the range defined by the level of the video compression technology or standard. In some cases, the level limits, for example, the maximum picture size, the maximum frame rate, the maximum reconstruction sample rate (measured in megasamples per second), the maximum reference picture size, etc. The limits set by the level may, in some cases, be further restricted by the specifications of the hypothetical reference decoder (HRD) and the metadata for HRD buffer management signaled within the coded video sequence.

[0072] FIG. 6 shows a block diagram of a video encoder (603) according to an embodiment of the present disclosure. The video encoder (603) can be used in a V-PCC encoder (300) that compresses a point cloud. In one example, the video compression modules (322) and (323) and the video compression module (332) are configured in the same manner as the encoder (603).

[0073] The video encoder (603) may receive images such as a padded geometry image, a padded texture image, etc., and generate a compressed image.

[0074] According to one embodiment, the video encoder (603) may code pictures (images) of a source video sequence in real time or under any other arbitrary time constraints required by an application, and compress them into a coded video sequence (compressed image). Enforcing an appropriate coding speed is one function of the controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units described below. This coupling is not depicted for clarity. The parameters set by the controller (650) can include rate control related parameters (picture skip, quantizer, lambda value of rate distortion optimization techniques,...), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) may be configured to have other appropriate functions related to the video encoder (603) optimized for a particular system design.

[0075] In some embodiments, the video encoder (603) is configured to operate in a coding loop. As an overly simplified explanation, in one example, the coding loop can include a source coder (630) (which, for example, is responsible for creating symbols such as a symbol stream based on an input picture to be coded and one or more reference pictures) and a (local) decoder (633) incorporated in the video encoder (603). (Since any compression between the symbols and the coded video bitstream in the video compression techniques considered in the disclosed subject matter is reversible) the decoder (633) reconstructs the symbols to create sample data in the same way that a (remote) decoder would also create. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream results in bit-exact results regardless of the position of the decoder (local or remote), the content of the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of the synchronization of the reference pictures (and the resulting drift if the synchronization cannot be maintained, for example, due to channel errors) is also used in some related arts.

[0076] The operation of the "local" decoder (633) can be the same as the operation of a "remote" decoder such as the video decoder (510), which has already been described in detail above in conjunction with FIG. 5. However, referring briefly to FIG. 5 as well, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy encoder (645) and the parser (520) can be reversible, the entropy decoding part of the video decoder (510), including the parser (520), may not be fully implemented in the local decoder (633).

[0077] In some examples, during operation, the source coder (630) may perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from a video sequence designated as a "reference picture". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of a reference picture that may be selected as a predictive reference for the input picture.

[0078] The local video decoder (633) may decode the coded video data of a picture that may be designated as a reference picture based on the symbols created by the source coder (630). The operation of the coding engine (632) may advantageously be an irreversible process. When the coded video data may be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence may typically be a replica of the source video sequence with some errors. The local video decoder (633) may replicate the decoding process that may be performed by the video decoder for the reference picture and store the reconstructed reference picture in the reference picture cache (634). In this way, the video encoder (603) may locally store a copy of the reconstructed reference picture having common content as the reconstructed reference picture obtained by the remote video decoder (in the absence of transmission errors).

[0079] Predictor (635) can perform predictive search of the coding engine (632). That is, in the case of a new picture to be coded, the predictor (635) can obtain sample data (as a candidate reference pixel block) or specific metadata such as reference picture motion vectors and block shapes that can serve as appropriate prediction references for new pixels, and search the reference picture memory (634). The predictor (635) can operate on the sample block for each pixel block to find an appropriate prediction reference. In some cases, the input picture may have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634) as determined by the search result obtained by the predictor (635).

[0080] The controller (650) can manage the coding operation of the source coder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0081] The outputs of all the above functional units may undergo entropy coding by the entropy coder (645). The entropy coder (645) converts the symbols generated by various functional units into a coded video sequence by reversibly compressing the symbols according to techniques such as Huffman coding, variable-length coding, and arithmetic coding.

[0082] The controller (650) can manage the operation of the video encoder (603). During coding, the controller (650) may assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, in many cases, a picture can be assigned as one of the following picture types.

[0083] An intracoded picture (I picture) can be coded and decoded without using other pictures in the sequence as a prediction source. Some video coders enable different types of intracoded pictures, including, for example, an Instantaneous Decoder Refresh (IDR) picture. Those skilled in the art are aware of those variants of I pictures as well as their respective uses and characteristics.

[0084] A predicted picture (P picture) can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0085] A bi - directionally predicted picture (B picture) can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0086] The source picture can generally be spatially subdivided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and coded block by block. The blocks can be coded predictively by referring to other (already coded) blocks determined by the coding assignment applied to each picture of the block. For example, the blocks of an I picture can be coded non-predictively or predictively by referring to already coded blocks of the same picture (spatial prediction or intra prediction). The pixel blocks of a P picture can be coded predictively via spatial prediction or via temporal prediction by referring to one previously coded reference picture. The blocks of a B picture can be coded predictively by referring to one or two previously coded reference pictures by spatial prediction or by temporal prediction.

[0087] The video encoder (603) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.265. In its operation, the video encoder (603) can perform various compression operations including predictive coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to the syntax specified in the video coding technology or standard being used.

[0088] The video may be in the form of a plurality of source pictures (images) in a time series. Intra-picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, and inter-picture prediction utilizes the correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. When a block within the current picture is similar to a reference block within a previously encoded and still buffered reference picture within the video, the block within the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture when multiple reference pictures are being used.

[0089] In some embodiments, a dual-prediction technique may be used in inter-picture prediction. According to the dual-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which are before the decoding order of the current picture in the video (however, the display order may be past and future respectively). A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0090] Furthermore, to improve coding efficiency, a merge mode technique may be used in inter-picture prediction.

[0091] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size such as 64×64 pixels, 32×32 pixels, 16×16 pixels, etc. Generally, a CTU includes three coding tree blocks (CTBs), which are one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type of the CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) according to the temporal predictability and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values) such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0092] FIG. 7 shows a block diagram of a G-PCC encoder (700) in some examples. The G-PCC encoder (700) may be configured to receive point cloud data and compress the point cloud data to generate a bitstream carrying the compressed point cloud data. In one embodiment, the G-PCC encoder (700) may include a position quantization module (710), a duplicate point removal module (712), an octree encoding module (730), an attribute transfer module (720), a level of detail (LOD) generation module (740), an attribute prediction module (750), a residual quantization module (760), an arithmetic coding module (770), an inverse residual quantization module (780), an addition module (781), and a memory (790) for storing the reconstructed attribute values.

[0093] As shown, an input point cloud (701) may be received by the G-PCC encoder (700). The positions (e.g., 3D coordinates) of the point cloud (701) are provided to the quantization module (710). The quantization module (710) is configured to quantize the coordinates to generate quantized positions. The duplicate point removal module (712) receives the quantized positions and is configured to perform a filtering process to identify and remove duplicate points. The octree encoding module (730) receives the filtered positions from the duplicate point removal module (712) and is configured to perform an octree-based encoding process to generate a sequence of occupancy codes that describe a 3D grid of voxels. The occupancy codes are provided to the arithmetic coding module (770).

[0094] The attribute transfer module (720) is configured to receive the attributes of the input point cloud and perform an attribute transfer process for determining the attribute values of each voxel when multiple attribute values are associated with respective voxels. The attribute transfer process can be performed on the sorted points output from the octree encoding module (730). The attributes after the transfer operation are provided to the attribute prediction module (750). The LOD generation module (740) operates on the sorted points output from the octree encoding module (730) and is configured to reorganize the points into different LODs. The LOD information is supplied to the attribute prediction module (750).

[0095] The attribute prediction module (750) processes the points according to the LOD-based order indicated by the LOD information from the LOD generation module (740). The attribute prediction module (750) generates an attribute prediction for the current point based on the reconfigured attributes of the set of neighboring points of the current point stored in the memory (790). Subsequently, a prediction residual can be obtained based on the original attribute values received from the attribute transfer module (720) and the locally generated attribute prediction. When a candidate index is used in each attribute prediction process, the index corresponding to the selected prediction candidate can be provided to the arithmetic coding module (770).

[0096] The residual quantization module (760) is configured to receive the prediction residual from the attribute prediction module (750) and perform quantization to generate a quantized residual. The quantized residual is provided to the arithmetic coding module (770).

[0097] The inverse residual quantization module (780) is configured to receive the quantized residual from the residual quantization module (760) and generate a reconstructed prediction residual by performing the inverse of the quantization operation performed by the residual quantization module (760). The addition module (781) is configured to receive the reconstructed prediction residual from the inverse residual quantization module (780) and each attribute prediction from the attribute prediction module (750). By combining the reconstructed prediction residual and the attribute prediction, a reconstructed attribute value is generated and stored in the memory (790).

[0098] The arithmetic coding module (770) is configured to receive an occupancy code, a candidate index (if used), a quantized residual (if generated), and other information, and perform entropy encoding to further compress the received values or information. Thereby, a compressed bitstream (702) carrying the compressed information can be generated. The bitstream (702) may be transmitted to a decoder that decodes the compressed bitstream, or provided, or stored in a storage device.

[0099] FIG. 8 shows a block diagram of a G-PCC decoder (800) according to an embodiment. The G-PCC decoder (800) may be configured to receive a compressed bitstream, perform point cloud data decompression to decompress the bitstream, and generate decoded point cloud data. In one embodiment, the G-PCC decoder (800) may include an arithmetic decoding module (810), an inverse residual quantization module (820), an octree decoding module (830), a LOD generation module (840), an attribute prediction module (850), and a memory (860) for storing the reconstructed attribute values.

[0100] As shown, a compressed bitstream (801) can be received by an arithmetic decoding module (810). The arithmetic decoding module (810) is configured to decode the compressed bitstream (801) to obtain a quantized residual (if generated) and an occupancy code of a point cloud. An octree decoding module (830) is configured to determine a reconstructed position of a point of the point cloud according to the occupancy code. A LOD generation module (840) is configured to reorganize points into different LODs based on the reconstructed positions and determine an LOD-based order. An inverse residual quantization module (820) is configured to generate a reconstructed residual based on the quantized residual received from the arithmetic decoding module (810).

[0101] An attribute prediction module (850) is configured to perform an attribute prediction process for determining an attribute prediction of a point according to the LOD-based order. For example, an attribute prediction of a current point can be determined based on reconstructed attribute values of adjacent points of the current point stored in a memory (860). In some examples, the attribute prediction can be combined with respective reconstructed residuals to generate a reconstructed attribute of the current point.

[0102] A sequence of reconstructed attributes generated from the attribute prediction module (850), together with the reconstructed positions generated from the octree decoding module (830), corresponds to a decoded point cloud (802) output from a G-PCC decoder (800) in one example. Additionally, the reconstructed attributes are also stored in the memory (860) and can be used later to derive an attribute prediction of a subsequent point.

[0103] In various embodiments, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented in hardware, software, or a combination thereof. For example, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented using a processing circuit such as one or more integrated circuits (ICs) that operate with or without software, such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs). In other examples, the encoder (300), decoder (400), encoder (700), and / or decoder (800) can be implemented as software or firmware including instructions stored on a non-volatile (or non-transitory) computer-readable storage medium. When the instructions are executed by a processing circuit such as one or more processors, the processing circuit is caused to perform the functions of the encoder (300), decoder (400), encoder (700), and / or decoder (800).

[0104] Note that the attribute prediction modules (750) and (850) configured to implement the attribute prediction techniques disclosed herein can be included in other decoders or encoders that may have the same or different structures as those shown in FIGS. 7 and 8. Additionally, the encoder (700) and decoder (800) can be included in the same device or, in various examples, separate devices.

[0105] According to some aspects of the present disclosure, mesh compression can use coding tools different from the PCC coding tools, or can use PCC coding tools such as the above PCC (e.g., G-PCC, V-PCC) encoders, the above PCC (e.g., G-PCC, V-PCC) decoders, and the like.

[0106] The mesh of an object (also called a mesh model or mesh frame) can include polygons that describe the surface of the object. Each polygon can be defined by the vertices of the polygon in 3D space and information about the edges that connect the vertices to the polygon. Information about how the vertices are connected (e.g., edge information) is called connectivity information. In some examples, the mesh of an object is formed by connected triangles that describe the surface of the object. Two triangles that share an edge are called two connected triangles. In some other examples, the mesh of an object is formed by connected quadrilaterals. Two quadrilaterals that share an edge can be called two connected quadrilaterals. Note that the mesh can be formed by other suitable polygons.

[0107] In some examples, the mesh can also include attributes such as colors and normals associated with the vertices. The attributes can be associated with the surface of the mesh by using mapping information that parameterizes the mesh with a 2D attribute map. The mapping information is usually described by a set of parametric coordinates called UV coordinates or texture coordinates associated with the mesh vertices. A 2D attribute map (called a texture map in some examples) is used to store high-resolution attribute information such as textures, normals, and displacements. Such information can be used for various purposes such as texture mapping and shading.

[0108] In some embodiments, the mesh can include components called geometry information, connectivity information, mapping information, vertex attributes, and attribute maps. In some examples, the geometry information is described by a set of 3D positions associated with the vertices of the mesh. In one example, (x, y, z) coordinates can be used to describe the 3D positions of the vertices, also called 3D coordinates. In some examples, the connectivity information includes a set of vertex indices that describe how the vertices are connected to create a 3D surface. In some examples, the mapping information describes how to map the mesh surface to a planar 2D region. In one example, the mapping information is described by a set of UV parametric / texture coordinates (u, v) associated with the mesh vertices along with the connectivity information. The texture coordinates associated with the mesh vertices can define the positions of the mapped mesh vertices within a 2D space such as a 2D map, UV atlas, etc. Texture coordinates are also called UV coordinates or 2D coordinates in some examples. In some examples, the vertex attributes include scalar or vector attribute values associated with the mesh vertices. In some examples, the attribute map includes attributes associated with the mesh surface and stored as a 2D image / video. In one example, the mapping between the video (e.g., 2D image / video) and the mesh surface is defined by the mapping information.

[0109] According to one aspect of the present disclosure, several techniques called UV mapping or mesh parameterization are used to map the surface of a mesh in a 3D domain to a 2D domain. In some examples, the mesh is divided into patches in the 3D domain. A patch is a continuous subset of the mesh having a boundary formed by boundary edges. The boundary edges of a patch belong to only one polygon of the patch and are edges that are not shared by two adjacent polygons within the patch. In some examples, the vertices of the boundary edges within a patch are called the boundary vertices of the patch, and the non-boundary vertices within the patch can be called the interior vertices of the patch.

[0110] In some examples, the mesh of an object is formed by connected triangles, and the mesh can be divided into patches, where each patch is a subset of the connected triangles. The boundary edges of a patch are edges that belong to only one triangle within the patch and are not shared by adjacent triangles within the patch. In some examples, the vertices of the boundary edges within a patch are called the boundary vertices of the patch, and the non-boundary vertices within the patch can be called the interior vertices of the patch. A boundary loop includes a series of boundary vertices, and the boundary edges formed by the series of boundary vertices can form a loop called the boundary loop.

[0111] According to one aspect of the present disclosure, in some examples, each patch is parameterized into a 2D shape (also called a UV patch, 2D patch, or 2D chart). The 2D shape can be packed (e.g., oriented and arranged) into a map, which is also called an atlas in some examples. In some examples, the map can be further processed using 2D image or video processing techniques.

[0112] In one example, the UV mapping technique generates a 2D UV atlas (also called a UV map) corresponding to the patches of a 3D mesh and one or more texture atlases (also called texture maps). The UV atlas includes an assignment of the 3D vertices of the 3D mesh to 2D points within a 2D domain (e.g., a rectangle). The UV atlas is a mapping between the coordinates of the 3D surface and the coordinates of the 2D domain. In one example, a point within the UV atlas at 2D coordinates (u, v) has a value formed by the coordinates (x, y, z) of a vertex within the 3D domain. In one example, the texture atlas includes the color information of the 3D mesh. For example, a point within the texture atlas at 2D coordinates (u, v) (having the 3D value of (x, y, z) within the UV atlas) has a color that specifies the color attribute of the point (x, y, z) of the 3D domain. In some examples, the coordinates (x, y, z) within the 3D region are called 3D coordinates or xyz coordinates, and the 2D coordinates (u, v) are called uv coordinates or UV coordinates.

[0113] According to some aspects of the present disclosure, one or more 2D maps (also referred to as 2D atlases in some examples) can be used to represent a mesh, and then mesh compression can be performed by encoding the 2D map using an image or video codec. Different techniques can be used to generate the 2D map.

[0114] FIG. 9 shows a diagram illustrating the mapping of a 3D mesh (910) to a 2D atlas (920) in some examples. In the example of FIG. 9, the 3D mesh (910) includes four vertices 1 to 4 that form four patches A to D. Each patch has a set of vertices and associated attribute information. For example, patch A is formed by vertices 1, 2, and 3 connected in a triangle, patch B is formed by vertices 1, 3, and 4 connected in a triangle, patch C is formed by vertices 1, 2, and 4 connected in a triangle, and patch D is formed by vertices 2, 3, and 4 connected in a triangle. In some examples, vertices 1, 2, 3, and 4 can each have their own attributes, and the triangles formed by vertices 1, 2, 3, and 4 can each have their own attributes.

[0115] In one example, the 3D patches A, B, C, and D are mapped to a 2D domain such as a 2D atlas (920), also referred to as a UV atlas (920) or map (920). For example, patch A is mapped to a 2D shape (also referred to as a UV patch) A' within the map (920), patch B is mapped to a 2D shape (also referred to as a UV patch) B' within the map (920), patch C is mapped to a 2D shape (also referred to as a UV patch) C' within the map (920), and patch D is mapped to a 2D shape (also referred to as a UV patch) D' within the map (920). In some examples, the coordinates in the 3D domain are referred to as (x, y, z) coordinates, and the coordinates in a 2D domain such as the map (920) are referred to as UV coordinates. The vertices within the 3D mesh can have corresponding UV coordinates within the map (920).

[0116] The map (920) can also be a geometry map having geometry information, a texture map having color, texture style, or other attribute information, or an occupancy map having occupancy information.

[0117] Note that in the example of FIG. 9, each patch is represented by a triangle, but a patch can include any suitable number of vertices connected to form a continuous subset of the mesh. In some examples, the vertices within a patch are connected in a triangle. Note that the vertices within a patch can be connected using other suitable shapes.

[0118] In one example, the geometry information of the vertices can be stored in a 2D geometry map. For example, a 2D geometry map stores the (x, y, z) coordinates of the sampling points at the corresponding points in the 2D geometry map. A point in the 2D geometry map at the (u, v) position has a three-component vector value corresponding to the x, y, and z values of the corresponding sampling point in the 3D mesh, respectively.

[0119] According to one aspect of the present disclosure, the regions within the map may not be fully occupied. For example, in FIG. 9, the regions outside the 2D shapes A', B', C', and D' are undefined. The sample values of the regions outside the 2D shapes A', B', C', and D' after decoding can be discarded. In some cases, the occupancy map is used to store some additional information for each pixel, such as storing a binary value to identify whether a pixel belongs to a patch or is undefined.

[0120] According to one aspect of the present disclosure, a dynamic mesh is a mesh in which at least one of the components (geometry information, connectivity information, mapping information, vertex attributes, and attribute maps) changes over time. A dynamic mesh can be described by a sequence of meshes (also called mesh frames). In some examples, the mesh frames within a dynamic mesh can be representations of the surface of an object at different times, and each mesh frame is a representation of the surface of the object at a specific time (also called a time instance). Since a dynamic mesh can potentially contain a significant amount of information that changes over time, a dynamic mesh may require a large amount of data. Mesh compression techniques can enable efficient storage and transmission of media content in mesh representations.

[0121] In some examples, a dynamic mesh can have constant connectivity information, time-varying geometry, and time-varying vertex attributes. In some examples, a dynamic mesh can have time-varying connectivity information. In one example, digital content creation tools typically generate dynamic meshes that have time-varying attribute maps and time-varying connectivity information. In some examples, volume acquisition techniques are used to generate dynamic meshes. Volume acquisition techniques can generate dynamic meshes that have time-varying connectivity information, particularly under real-time constraints.

[0122] Dynamic meshes may require large amounts of data because they can contain a significant amount of information that changes over time. In particular, the bits spent on UV coordinate signaling are an important part of the bitstream. Some aspects of the present disclosure provide mesh compression techniques in which texture coordinates (also called UV coordinates) are not encoded but are estimated. Estimation of the texture coordinates of the current mesh frame means that the texture coordinates of the current mesh frame are not directly encoded and signaled, but are obtained from other suitable sources, such as derivation from the 3D coordinates and connectivity of the current mesh frame, inheritance of texture coordinates from previously coded mesh frames, etc. This can improve mesh coding efficiency. The mesh compression technique is based on UV coordinate derivation in mesh compression. The mesh compression technique can be applied individually or in any combination of forms. It should also be noted that the mesh compression technique can be applied to both dynamic and static meshes. A static mesh contains one mesh frame.

[0123] In some examples, the mesh compression technique can derive UV coordinates on the decoder side instead of signaling the UV coordinates in the bitstream from the encoder. Thus, a portion of the bitrate for representing the UV coordinates can be saved. In some examples, the mesh codec (encoder / decoder) can be configured to perform reversible mesh coding. In some examples, the mesh codec (encoder / decoding) can be configured to perform irreversible mesh coding. Some aspects of the present disclosure provide mesh compression techniques for use with reversible mesh codecs (encoders / decoders), and some aspects of the present disclosure provide mesh compression techniques for use with irreversible mesh codecs (encoders / decoders).

[0124] FIG. 10 shows a diagram of a framework (1000) of a reversible mesh codec according to some embodiments of the present disclosure. The framework (1000) includes a mesh encoder (1010) and a mesh decoder (1050). The mesh encoder (1010) encodes an input mesh (1005) (mesh frame in the case of a dynamic mesh) into a bitstream (1045), and the mesh decoder (1050) decodes the bitstream (1045) to generate a reconstructed mesh (1095) (mesh frame in the case of a dynamic mesh).

[0125] The mesh encoder (1010) can be any suitable device such as a computer, a server computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The mesh decoder (1050) can be any suitable device such as a computer, a client computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The bitstream (1045) can be transmitted from the mesh encoder (1010) to the mesh decoder (1050) via a communication network (not shown).

[0126] In the example of FIG. 10, the mesh encoder (1010) includes a parameterization module (1020), a texture transformation module (1030), and a plurality of encoders such as a reversible encoder (1040), an image / video encoder (1041), and an auxiliary encoder (1042).

[0127] In one example, the input mesh (1005) can include the 3D coordinates of vertices (represented by XYZ), the original texture coordinates of vertices (also called original UV coordinates, represented by original UV), the connectivity information of vertices (represented by connectivity), and the original texture map. The original texture map is associated with the original texture coordinates of vertices. The 3D coordinates of vertices and the connectivity information of vertices are represented by XYX & connectivity (1021) in FIG. 10 and provided to the parameterization module (1020) and the reversible encoder (1040). The original texture map and original UV represented by (1031) are provided to the texture transformation module (1030).

[0128] The parameterization module (1020) is configured to perform parameterization that utilizes the XYZ and connectivity (1021) of the input mesh (1005) to generate new texture coordinates of vertices (also called new UV coordinates of vertices) represented by new UV (1025).

[0129] The texture transformation module (1030) can receive the original texture map and original UV (1031), can receive the new UV (1025), and can transform the original texture map into a new texture map (1022) associated with the new UV (1025). The new texture map (1022) is provided to the image / video encoder (1041) for encoding.

[0130] XYZ & connectivity (1021) can include the geometric information of vertices, for example, the (x, y, z) coordinates representing the position of vertices in 3D space, the connectivity information of vertices (for example, the definition of a face also called a polygon). In some examples, XYZ & connectivity (1021) also includes vertex attributes such as normals, color reflectance, etc. The reversible encoder (1040) can use reversible coding techniques to encode XYZ & connectivity (1021) into a bitstream (1045).

[0131] A new texture map (1022) (also called a new attribute map in some examples) includes attributes associated with a mesh surface with respect to new texture coordinates of vertices. In some examples of dynamic mesh processing, a new texture map (1022) of a sequence of mesh frames can form a video sequence. The new texture map (1022) can be encoded by an image / video encoder (1041) using appropriate image and / or video coding techniques.

[0132] In the example of FIG. 10, the mesh encoder (1010) can generate auxiliary data (1023) including auxiliary information such as flags, indices, etc. The auxiliary data encoder (1042) receives the auxiliary data (1023) and encodes the auxiliary data (1023) into a bitstream (1045).

[0133] In the example of FIG. 10, the encoded outputs from the reversible encoder (1040), the image / video encoder (1041), and the auxiliary data encoder (1042) are mixed (e.g., multiplexed) into a bitstream (1045) that carries the encoded information of the input mesh (1005).

[0134] In the example of FIG. 10, the mesh decoder (1050) can demultiplex the bitstream (1045) into sections, and the sections can be decoded respectively by a plurality of decoders such as a reversible decoder (1060), an image / video decoder (1061), and an auxiliary data decoder (1062).

[0135] In one example, a reversible decoder (1060) corresponds to a reversible encoder (1040) and can decode a section of a bitstream (1045) encoded by the reversible encoder (1040). The reversible encoder (1040) and the reversible decoder (1060) can perform reversible encoding and decoding. The reversible decoder (1060) can output the decoded 3D coordinates of the vertices and the connection information of the vertices as shown by the decoded XYZ & connectivity (1065). The decoded XYZ & connectivity (1065) can be the same as the XYZ & connectivity (1021).

[0136] In some examples, the parameterization module (1070) includes the same parameterization algorithm as the parameterization module (1020). The parameterization module (1070) is configured to perform parameterization to generate new texture coordinates of the vertices indicated by the new UV (1075) using the decoded XYZ & connectivity (1065). Note that in some examples, the decoded XYZ & connectivity (1065) is identical to the XYZ & connectivity (1021), and thus the new UV (1075) can be the same as the new UV (1025).

[0137] In one example, an image / video decoder (1061) corresponds to an image / video encoder (1041) and can decode a section of a bitstream (1045) encoded by the image / video encoder (1041). The image / video decoder (1061) can generate a decoded new texture map (1066). The decoded new texture map (1066) is associated with the new UV (1025) that can be the same as the new UV (1075) in the mesh decoder (1050). The decoded new texture map (1066) is provided to the mesh reconstruction module (1080).

[0138] In one example, an auxiliary data decoder (1062) corresponds to an auxiliary data encoder (1042) and can decode a section of a bitstream (1045) encoded by the auxiliary data encoder (1042). The auxiliary data decoder (1062) can generate decoded auxiliary data (1067). The decoded auxiliary data (1067) is provided to a mesh reconstruction module (1080).

[0139] The mesh reconstruction module (1080) receives the decoded XYZ & connectivity (1065), new UV (1075), decoded new texture map (1066), and decoded auxiliary data (1067), and generates a reconstructed mesh (1095) accordingly.

[0140] Note that components within the mesh encoder (1010) such as the parameterization module (1020), texture transformation module (1030), reversible encoder (1040), image / video encoder (1041), and auxiliary data encoder (1042) can each be implemented by various techniques. In one example, the components are implemented by an integrated circuit. In other examples, the components are implemented using software that can be executed by one or more processors.

[0141] Note that components within the mesh decoder (1050) such as the reversible decoder (1060), parameterization module (1070), mesh reconstruction module (1080), image / video decoder (1061), and auxiliary data decoder (1062) can each be implemented by various techniques. In one example, the components are implemented by an integrated circuit. In other examples, the components are implemented using software that can be executed by one or more processors.

[0142] According to one aspect of the present disclosure, the parameterization module (1020) and the parameterization module (1070) can be implemented with any suitable parameterization algorithm. In some examples, a parameterization algorithm for generating new texture coordinates aligned across mesh frames within a 3D mesh sequence can be implemented in the parameterization module (1020) and the parameterization module (1070). Thus, the new texture map of the mesh frames within the 3D mesh sequence can have a relatively large correlation across the mesh frames, the new texture map of the mesh frames can be coded using inter prediction, and the coding efficiency can be further improved.

[0143] In some examples, to illustrate the parameterization algorithm, the parameterization module (1020) is used. The parameterization module (1020) can divide the faces of the mesh frame (also called polygons) into different charts (also called 2D patches, 2D shapes, UV patches, etc.) according to the normal values of the faces. The charts can be packed into a texture space (e.g., a UV map, a 2D UV atlas). In some examples, the parameterization module (1020) can process each mesh frame in turn and generate a plurality of UV maps having charts packed into the UV map. The UV maps are the UV maps of the mesh frames within the sequence respectively. The parameterization module (1020) can temporally align the charts across the UV maps of the mesh frames.

[0144] In some examples, alignment can use metrics calculated from the charts. For example, the metrics can be scale-invariant metrics, rotation-invariant metrics, translation-invariant metrics, or affine transformation-invariant metrics, etc. In one example, the metric can include the center of the 3D coordinates of each chart. In another example, the metric can include the average depth of each chart. In another example, the metric can include the weighted average texture value or attribute value of each chart. Note that the metrics used for alignment can include one or more of the metrics listed above, or can include other metrics having similar characteristics.

[0145] In some examples, the parameterization module (1020) can temporally align the charts across the mesh frames in the sequence using one or more metrics calculated from the charts. The UV coordinates obtained after temporal alignment are new UVs.

[0146] Note that the image / video encoder (1041) and the image / video decoder (1061) can be implemented using any suitable image and / or video codec. In the example of a static mesh having a single mesh frame, the image / video encoder (1041) and the image / video decoder (1061) can be implemented by an image codec such as a JPEG or PNG codec to code a new texture map of the single mesh frame. In the example of a dynamic mesh having a sequence of mesh frames, the image / video encoder (1041) and the image / video decoder (1061) can be implemented by a video codec such as H.265 to code a new texture map for the sequence of mesh frames.

[0147] Note that in some examples, the reversible encoder (1040) and the reversible decoder (1060) can be implemented by a reversible mesh codec. The reversible mesh codec is configured to encode the XYZ and connectivity information of the mesh frame and skip the encoding of the UV coordinates. In one embodiment, the reversible encoder (1040) and the reversible decoder (1060) can be implemented to include the SC3DMC codec, which is MPEG reference software for static mesh compression, to reversibly encode the mesh.

[0148] FIG. 11 shows a diagram of a framework (1100) of an irreversible mesh codec according to some embodiments of the present disclosure. The framework (1100) includes a mesh encoder (1110) and a mesh decoder (1150). The mesh encoder (1110) encodes an input mesh (1105) (mesh frame in the case of a dynamic mesh) into a bitstream (1145), and the mesh decoder (1150) decodes the bitstream (1145) to generate a reconstructed mesh (1195) (mesh frame in the case of a dynamic mesh).

[0149] The mesh encoder (1110) can be any suitable device such as a computer, a server computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The mesh decoder (1150) can be any suitable device such as a computer, a client computer, a desktop computer, a laptop computer, a tablet computer, a smartphone, a gaming device, an AR device, a VR device, etc. The bitstream (1145) can be transmitted from the mesh encoder (1110) to the mesh decoder (1150) via a communication network (not shown).

[0150] In the example of FIG. 11, the mesh encoder (1110) includes a parameterization module (1120), a texture conversion module (1130), and a plurality of encoders such as an irreversible encoder (1140), a video encoder (1141), and an auxiliary encoder (1142). Further, the mesh encoder (1110) includes an irreversible decoder (1145) corresponding to the irreversible encoder (1140).

[0151] In one example, the input mesh (1105) can include the 3D coordinates of the vertices (indicated by XYZ), the original texture coordinates of the vertices (also called original UV coordinates and indicated by original UV), the connectivity information of the vertices (indicated by connectivity), and the original texture map. The original texture map is associated with the original texture coordinates of the vertices. The 3D coordinates of the vertices and the connectivity information of the vertices are indicated by the XYX & connectivity (1121) of FIG. 11 and provided to the irreversible encoder (1140). The original texture map and the original UV indicated by (1131) are provided to the texture conversion module (1130).

[0152] XYZ & connectivity (1121) can include vertex geometry information, e.g., (x, y, z) coordinates representing the position of a vertex in 3D space, and vertex connectivity information (e.g., the definition of a face, also called a polygon). In some examples, XYZ & connectivity (1121) also includes vertex attributes such as normals, color reflectance, etc. The irreversible encoder (1140) can use irreversible coding techniques to encode XYZ & connectivity (1121) into a bitstream (1145). Further, the encoded XYZ & connectivity, indicated by (1144), is also provided to the irreversible decoder (1145). In one example, the irreversible decoder (1145) corresponds to the irreversible encoder (1140) and can decode the encoded XYZ & connectivity (1144) encoded by the irreversible encoder (1140). The irreversible encoder (1140) and the irreversible decoder (1145) can perform irreversible encoding and decoding. Note that in some examples, the irreversible encoder (1140) and the irreversible decoder (1145) can be implemented by an irreversible mesh codec. The irreversible mesh codec is configured to encode the XYZ and connectivity information of the mesh frame and skip the coding of the UV coordinates. In one embodiment, the irreversible encoder (1140) and the irreversible decoder (1145) can be implemented to include Draco for irreversible mesh coding. The irreversible decoder (1145) can output the decoded 3D coordinates of the vertices and the vertex connectivity information as indicated by the decoded XYZ & connectivity (1146).

[0153] The parameterization module (1120) is configured to perform parameterization to generate new texture coordinates of the vertices indicated by the new UV (1125) using the decoded XYZ & connectivity (1146).

[0154] The texture conversion module (1130) can receive the original texture map and the original UV (1131), can receive the new UV (1125), and can convert the original texture map into a new texture map (1122) associated with the new UV (1125). The new texture map (1122) is provided to the image / video encoder (1141) for encoding.

[0155] The new texture map (1122) (also called the new attribute map in some examples) includes attributes associated with the mesh surface with respect to the new texture coordinates of the vertices. In some examples of dynamic mesh processing, the new texture maps (1122) of a sequence of mesh frames can form a video sequence. The new texture map (1122) can be encoded by the image / video encoder (1141) using appropriate image and / or video coding techniques.

[0156] In the example of FIG. 11, the mesh encoder (1110) can generate auxiliary data (1123) including auxiliary information such as flags, indices, etc. The auxiliary data encoder (1142) receives the auxiliary data (1123) and encodes the auxiliary data (1123) into a bitstream (1145).

[0157] In the example of FIG. 11, the encoded outputs from the irreversible encoder (1140), the image / video encoder (1141), and the auxiliary data encoder (1142) are mixed (e.g., multiplexed) into a bitstream (1145) that carries the encoded information of the input mesh (1105).

[0158] In the example of FIG. 11, the mesh decoder (1150) can demultiplex the bitstream (1145) into sections, and the sections can be decoded by a plurality of decoders such as the irreversible decoder (1160), the image / video decoder (1161), and the auxiliary data decoder (1162), respectively.

[0159] In one example, the irreversible decoder (1160) corresponds to the irreversible encoder (1140) and can decode a section of the bitstream (1145) encoded by the irreversible encoder (1140). The irreversible encoder (1140) and the irreversible decoder (1160) can perform irreversible encoding and decoding. The irreversible decoder (1160) can output the decoded 3D coordinates of the vertices and the connection information of the vertices as shown by the decoded XYZ & connectivity (1165).

[0160] In some examples, the same decoding technique is used for the irreversible decoder (1145) of the mesh encoder (1110) and the irreversible decoder (1165) of the mesh decoder (1150). Thus, the decoded XYZ & connectivity (1165) can have the same decoded 3D coordinates of the vertices and the connection information of the vertices as the decoded XYZ & connectivity (1146). Note that the decoded XYZ & connectivity (1165) and the decoded XYZ & connectivity (1146) may be different from the XYZ & connectivity (1121) due to irreversible encoding and decoding.

[0161] In some examples, the parameterization module (1170) includes the same parameterization algorithm as the parameterization module (1120). The parameterization module (1170) is configured to perform parameterization to generate new texture coordinates of the vertices indicated by the new UV (1175) using the decoded XYZ & connectivity (1165). Note that in some examples, the decoded XYZ & connectivity (1165) is identical to the XYZ & connectivity (1146), and thus the new UV (1175) can be identical to the new UV (1125).

[0162] In one example, an image / video decoder (1161) corresponds to an image / video encoder (1141) and can decode a section of a bitstream (1145) encoded by the image / video encoder (1141). The image / video decoder (1161) can generate a decoded new texture map (1166). The decoded new texture map (1166) is associated with a new UV (1125) that can be the same as the new UV (1175) in the mesh decoder (1150). The decoded new texture map (1166) is provided to a mesh reconstruction module (1180).

[0163] In one example, an auxiliary data decoder (1162) corresponds to an auxiliary data encoder (1142) and can decode a section of a bitstream (1145) encoded by the auxiliary data encoder (1142). The auxiliary data decoder (1162) can generate decoded auxiliary data (1167). The decoded auxiliary data (1167) is provided to the mesh reconstruction module (1180).

[0164] The mesh reconstruction module (1180) receives the decoded XYZ & connectivity (1165), the new UV (1175), the decoded new texture map (1166), and the decoded auxiliary data (1167), and generates a reconstructed mesh (1195) accordingly.

[0165] Note that components within the mesh encoder (1110), such as the parameterization module (1120), texture conversion module (1130), irreversible encoder (1140), irreversible decoder (1145), image / video encoder (1141), and auxiliary data encoder (1142), can each be implemented by various techniques. In one example, the components are implemented by integrated circuits. In other examples, the components are implemented using software that can be executed by one or more processors.

[0166] Note that components within the mesh decoder (1150), such as the irreversible decoder (1160), parameterization module (1170), mesh reconstruction module (1180), image / video decoder (1161), and auxiliary data decoder (1162), can each be implemented by various techniques. In one example, the components are implemented by integrated circuits. In other examples, the components are implemented using software that can be executed by one or more processors.

[0167] According to one aspect of the present disclosure, the parameterization module (1120) and the parameterization module (1170) can be implemented with any suitable parameterization algorithm. In some examples, a parameterization algorithm for generating new texture coordinates aligned across mesh frames in a 3D mesh sequence can be implemented in the parameterization module (1120) and the parameterization module (1170). Thus, the new texture maps of the mesh frames in the 3D mesh sequence can have a relatively large correlation across the mesh frames, the new texture maps of the mesh frames can be coded using inter-prediction, and the coding efficiency can be further improved.

[0168] In some examples, a parameterization module (1120) is used to show a parameterization algorithm, and the parameterization module (1120) can divide the faces of a mesh frame (also called polygons) into different charts (also called 2D patches, 2D shapes, UV patches, etc.) according to the normal values of the faces. The charts can be packed into a texture space (e.g., a UV map, a 2D UV atlas). In some examples, the parameterization module (1120) can process each mesh frame in turn and generate a plurality of UV maps having charts packed into the UV map. The UV maps are the UV maps of the mesh frames within the sequence respectively. The parameterization module (1120) can temporally align the charts across the UV maps of the mesh frames.

[0169] In some examples, the alignment can use a metric calculated from the charts. For example, the metric can be a scale-invariant metric, a rotation-invariant metric, a translation-invariant metric, or an affine transformation-invariant metric, etc. In one example, the metric can include the center of the 3D coordinates of each chart. In other examples, the metric can include the average depth of each chart. In other examples, the metric can include the weighted average texture value or attribute value of each chart. It should be noted that the metric used for alignment can include one or more of the metrics listed above, or can include other metrics having similar characteristics.

[0170] In some examples, the parameterization module (1120) can temporally align the charts across the mesh frames in the sequence using one or more metrics calculated from the charts. The UV coordinates obtained after the temporal alignment are the new UVs.

[0171] Note that the image / video encoder (1141) and the image / video decoder (1161) can be implemented using any suitable image and / or video codec. In the example of a static mesh with a single mesh frame, the image / video encoder (1141) and the image / video decoder (1161) can be implemented by an image codec such as a JPEG or PNG codec to code a new texture map for the single mesh frame. In the example of a dynamic mesh with a sequence of mesh frames, the image / video encoder (1141) and the image / video decoder (1161) can be implemented by a video codec such as H.265 to code a new texture map for the sequence of mesh frames.

[0172] According to an aspect of the present disclosure, the mesh encoder side can select to signal the texture coordinates of the vertices in the bitstream (so that the mesh decoder can directly decode the texture coordinates from the bitstream) or to derive the texture coordinates on the mesh decoder side. In some examples, a high-level indication flag can be used to signal such a selection and notify the mesh decoder.

[0173] In one embodiment, a sequence-level flag can be used to indicate that the texture coordinates of the vertices should be derived on the mesh decoder side or that the texture coordinates should be signaled in the bitstream. The entire sequence of mesh frames associated with the sequence-level flag follows the indication of the sequence-level flag.

[0174] In other embodiments, a frame-level flag can be used to indicate that the texture coordinates should be derived on the mesh decoder side or that the texture coordinates should be signaled in the bitstream. The entire mesh frame associated with the frame-level flag follows the indication of the frame-level flag.

[0175] In other embodiments, a group-level flag may be used for a group of mesh frames to indicate that texture coordinates should be derived on the mesh decoder side or that texture coordinates should be signaled within the bitstream. The entire group of mesh frames follows the indication of the group-level flag. The concept of a group of mesh frames is similar to that of a Group of Pictures in video coding. For example, a Group of Pictures can include pictures that refer to the same Picture Parameter Set (PPS) or pictures within the same random access point, and a group of mesh frames can include mesh frames that refer to the same Picture Parameter Set (PPS) or mesh frames within the same random access point.

[0176] In other embodiments, as described above, the frame-level flag or the group-level flag is conditionally signaled when a sequence-level flag exists, indicating that texture coordinate derivation may be permitted. Otherwise, the frame-level flag or the group-level flag need not be signaled. Instead, default values (indicating that texture coordinates are signaled in the bitstream) are assigned to the frame-level flag and the group-level flag.

[0177] According to another aspect of the present disclosure, some additional information may be signaled in the bitstream to assist the mesh decoder in deriving texture coordinates. The additional information may be signaled at various levels. In one example, the additional information may be signaled by a high-level syntax such as a sequence header, a frame header, a group header (for a group of mesh frames), etc. In another example, the additional information may be signaled by a lower-level syntax for each patch, each chart, each slice, etc.

[0178] In some embodiments, a flag indicating whether the texture coordinates of the current mesh frame are inherited from a previously decoded mesh frame may be included in the bitstream. Further, if the flag indicates that the texture coordinates of the current mesh frame are inherited from a previously decoded mesh frame, an index may be further signaled to indicate a reference frame from a set of previously decoded frames. Then, by skipping the entire parameterization process at the mesh decoder, the texture coordinates of the current frame may be directly inherited from the indicated mesh frame.

[0179] In some embodiments, key vertices from the vertices of a mesh frame to start parameterization can be determined by the mesh encoder, and the indices of the key vertices can be signaled in relation to the mesh frame. The mesh decoder then decodes the indices of the key vertices and uses the key vertices as starting points to start parameterization to divide the mesh frame into, for example, a plurality of charts / patches / slices.

[0180] In some embodiments, for each chart associated with the mesh frame, each patch associated with the mesh frame, each slice associated with the mesh frame, etc., a flag may be signaled for the current portion of the mesh frame (e.g., the current chart, the current patch, the current slice) to indicate whether the texture coordinates of the current portion are inherited from a previously decoded frame or derived by a parameterization method. If the flag indicates that the texture coordinates of the current portion are inherited from a previously decoded frame, an index indicating the selection of the previously decoded frame and the selection of the decoded portion (e.g., chart, patch, slice) may be signaled within the bitstream.

[0181] In some embodiments, the mesh encoder can select a particular parameterization method from a set of parameterization method candidates and include an index indicating the selection of the particular parameterization method in the bitstream. Note that the index can be signaled at different levels, such as in the sequence header, frame header, Group of Pictures, in relation to patches, in relation to slices (e.g., in the slice header), etc.

[0182] FIG. 12 is a diagram showing an example of a syntax table (1200) in some examples.

[0183] The syntax table (1200) uses a plurality of flags to indicate a choice of whether to signal the texture coordinates of vertices in the bitstream (so that the mesh decoder can directly decode the texture coordinates from the bitstream) or to derive the texture coordinates on the mesh decoder side. For example, the plurality of flags includes a first flag (1210) at the sequence level and a second flag (1220) at the frame level. The syntax table (1200) also includes a third flag (1230) indicating whether the texture coordinates of the current mesh frame are inherited from a previously decoded mesh frame, and an index (1240) indicating a reference frame from a set of previously decoded frames.

[0184] Specifically, in the example of FIG. 12, the first flag (1210) is represented by ps_derive_uv_enabled_flag. When the first flag (1210) is 0, it indicates that the derivation of texture coordinates for the mesh frames within the sequence is not enabled. When the first flag (1210) is 1, it indicates that texture coordinates can be derived for the mesh frames within the sequence. And whether to derive the texture coordinates for the current mesh can be determined by the second flag (1220) associated with the current mesh. In one example, when the first flag (1210) does not exist in the bitstream, the first flag (1210) can be set to 0.

[0185] In the example of FIG. 12, the second flag (1220) is represented by ph_derive_uv_flag. When the second flag (1220) is 0, it indicates that texture coordinates have not been derived for the current mesh frame, and when the second flag (1220) is 1, it indicates that texture coordinates have been derived for the current mesh frame (e.g., by performing parameterization). If it does not exist, the second flag (1220) is set equal to 0. When the second flag (1220) is 0, the texture coordinates of the current mesh frame can be inherited or signaled. Whether the texture coordinates of the current mesh frame are inherited or signaled can depend on the third flag (1230).

[0186] In the example of FIG. 12, the third flag (1230) is represented by ph_inherit_uv_flag. When the third flag (1230) is 0, it indicates that the texture coordinates are not inherited from the previously decoded mesh frame for the current mesh frame and that the texture coordinates are signaled within the bitstream. When the third flag (1230) is 1, it indicates that the texture coordinates are inherited from the previously decoded mesh frame for the current mesh frame. If it does not exist, the third flag (1230) is set equal to 0.

[0187] When the third flag (1230) indicates that the texture coordinates are inherited from a previously decoded mesh frame, an index (1240) represented by inherit_idx is signaled to specify the frame index of the decoded mesh frame from the list of decoded mesh frames. The texture coordinates of the current mesh frame are inherited from the decoded mesh frame. If not present, the index (1240) is set equal to 0 in one example.

[0188] FIG. 13 shows a flowchart illustrating an overview of a process (1300) according to an embodiment of the present disclosure. The process (1300) can be used during an encoding process for one or more mesh frames. In various embodiments, the process (1300) is executed by a processing circuit. In some embodiments, the process (1300) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs the process (1300). The process starts at (S1301) and proceeds to (S1310).

[0189] (S1310), new texture coordinates associated with vertices within the first 3D mesh frame are determined. The first 3D mesh frame represents the surface of an object with polygons.

[0190] (S1320), the original texture map of the first 3D mesh frame is converted to a new texture map associated with the new texture coordinates of the vertices. The original texture map is associated with the original texture coordinates of the vertices.

[0191] (S1330), the 3D coordinates of the vertices, the connectivity information of the vertices, and the new texture map are encoded into a bitstream for carrying the first 3D mesh frame.

[0192] In some examples, the encoding of the 3D coordinates of the vertices and the connectivity information of the vertices is performed according to a reversible codec.

[0193] In some examples, the encoding of the 3D coordinates of the vertices and the connectivity information of the vertices is performed according to an irreversible codec.

[0194] In some examples, to estimate new texture coordinates associated with the vertices, the encoded 3D coordinates and connectivity are decoded to generate the decoded 3D coordinates of the vertices and the decoded connectivity information of the vertices. Parameterization is performed according to the decoded 3D coordinates of the vertices and the decoded connectivity information to determine new texture coordinates associated with the vertices.

[0195] In some examples, to perform parameterization, the polygons of the first 3D mesh frame are divided into one or more first charts in a 2D map according to the decoded 3D coordinates of the vertices and the decoded connectivity information. Temporal alignment is performed to align the one or more first charts with one or more second charts associated with a second 3D mesh frame. The first 3D mesh frame and the second 3D mesh frame are frames within a 3D mesh sequence. New texture coordinates are determined from the one or more first charts using the temporal alignment.

[0196] In some examples, the polygons are divided according to the normal values associated with the polygons.

[0197] In some examples, the temporal alignment is performed according to at least one of a scale-invariant metric, a rotation-invariant metric, a translation-invariant metric, and an affine transformation-invariant metric.

[0198] In one example, the temporal alignment is performed according to the center of the chart calculated based on the 3D coordinates associated with the chart. In other examples, the temporal alignment is performed according to the average depth of the chart. In other examples, the temporal alignment is performed according to the weighted average texture value of the chart. In other examples, the temporal alignment is performed according to the weighted average attribute value of the chart.

[0199] In some examples, a flag indicating the enabling of texture coordinate derivation is encoded in the bitstream. The flag is at least one of a sequence level flag, a group of pictures level flag, and a picture level flag.

[0200] In some examples, a flag indicating the inheritance of texture coordinates is encoded in the bitstream to estimate the texture coordinates associated with the vertices. A particular coded 3D mesh frame is selected from a list of coded 3D mesh frames to inherit texture coordinates from a particular coded 3D mesh frame. An index indicating the selection of a particular coded 3D mesh frame from the list of coded 3D mesh frames is coded in the bitstream.

[0201] In some examples, a flag associated with a part of the texture map is encoded in the bitstream, and the flag indicates whether the texture coordinates associated with a part of the texture map are inherited from the coded mesh frame or derived by parameterization.

[0202] In some examples, a set of key vertices among the vertices is determined, and the parameterization starts from the set of key vertices. An index of the set of key vertices is encoded in the bitstream.

[0203] In some examples, a selected parameterization method is determined from a list of parameterization method candidates, and an index indicating the parameterization method selected from the list of parameterization method candidates is encoded into the bitstream.

[0204] The process then proceeds to (S1399) and ends.

[0205] Process (1300) can be suitably adapted. Steps of process (1300) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0206] FIG. 14 shows a flowchart illustrating an overview of a process (1400) according to an embodiment of the present disclosure. Process (1400) can be used during a decoding process for one or more mesh frames. In various embodiments, process (1400) is executed by a processing circuit. In some embodiments, process (1400) is implemented with software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit performs process (1400). The process starts at (S1401) and proceeds to (S1410).

[0207] In (S1410), the 3D coordinates of the vertices and the connection information of the vertices within the first 3D mesh frame are decoded from the bitstream carrying the first 3D mesh frame. The first 3D mesh frame represents the surface of an object with polygons.

[0208] In (S1420), texture coordinates associated with the vertices are estimated. In some examples, the bitstream does not include encoded texture coordinates.

[0209] In (S1430), the texture map of the first 3D mesh frame is decoded from the bitstream. The texture map includes one or more first 2D charts having 2D vertices with texture coordinates.

[0210] (S1440), the first 3D mesh frame is reconstructed based on the 3D coordinates of vertices, the connection information of vertices, the texture map, and the texture coordinates.

[0211] In some examples, the 3D coordinates of vertices and the connection information of vertices are coded using a reversible codec. In some examples, the 3D coordinates of vertices and the connection information of vertices are coded using an irreversible codec.

[0212] In some examples, to estimate the texture coordinates associated with vertices, parameterization is performed according to the 3D coordinates and connection information of vertices to determine the texture coordinates associated with vertices. For example, the polygons of the first 3D mesh frame are divided into the first one or more charts, and the first one or more charts are packed into a 2D map. Temporal alignment is performed to align the first one or more charts with the second one or more charts associated with the second 3D mesh frame. The first 3D mesh frame and the second 3D mesh frame are frames within a 3D mesh sequence. The texture coordinates are determined from the first one or more charts using temporal alignment.

[0213] In some examples, polygons are divided into the first one or more charts according to the normal values associated with the polygons.

[0214] In some examples, temporal alignment is performed according to at least one of a scale-invariant metric, a rotation-invariant metric, a translation-invariant metric, and an affine transformation-invariant metric.

[0215] In one example, the temporal alignment is performed according to the center of the chart calculated based on the 3D coordinates associated with the chart. In other examples, the temporal alignment is performed according to the average depth of the chart. In other examples, the temporal alignment is performed according to the weighted average texture value of the chart. In other examples, the temporal alignment is performed according to the weighted average attribute value of the chart.

[0216] In some examples, a flag indicating activation of texture coordinate derivation is decoded. The flag is at least one of a sequence level flag, a group of pictures level flag, and a frame level flag.

[0217] In some examples, a flag indicating inheritance of texture coordinates is decoded to estimate the texture coordinates associated with the vertices. Next, an index indicating a selected 3D mesh frame from the list of decoded 3D mesh frames is decoded. And the texture coordinates are inherited from the selected 3D mesh frame.

[0218] In some examples, a flag associated with a part of the texture map is decoded. The flag indicates whether the texture coordinates associated with the part of the texture map are inherited from the decoded mesh frame or derived by parameterization.

[0219] In some examples, the index of a set of key vertices within the vertex is decoded. The parameterization starts from the set of key vertices.

[0220] In some examples, an index indicating a selected parameterization method from the list of parameterization method candidates is decoded. And the parameterization is performed according to the selected parameterization method.

[0221] The process then proceeds to (S1499) and ends.

[0222] The process (1400) can be suitably adapted. The steps of the process (1400) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0223] The techniques disclosed in this disclosure may be used separately or combined in any order. Further, each of the techniques (e.g., methods, embodiments), encoders, and decoders may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In some examples, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0224] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 15 shows a computer system (1500) suitable for implementing certain embodiments of the disclosed subject matter.

[0225] The computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, linking, etc., and create code that includes instructions that can be executed directly, or via interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0226] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, Internet of Things devices, etc.

[0227] The components shown in FIG. 15 for the computer system (1500) are essentially illustrative and are not intended to imply any limitation regarding the use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having dependencies or requirements with respect to any one of the components shown in the exemplary embodiments of the computer system (1500) or combinations of components.

[0228] The computer system (1500) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (keystrokes, swipes, movement of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). The human interface device may also be used to capture specific media that is not necessarily directly related to conscious input by a human, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).

[0229] The input human interface device may include one or more of a keyboard (1501), a mouse (1502), a trackpad (1503), a touch screen (1510), a data glove (not shown), a joystick (1505), a microphone (1506), a scanner (1507), a camera (1508), etc. (only one of each is shown).

[0230] The computer system (1500) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, via tactile output, sound, light, and smell / taste. Such human interface output devices include, for example, tactile output devices (such as tactile feedback by a touch screen (1510), a data glove (not shown), or a joystick (1505), although there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1509), headphones (not shown), etc.), visual output devices (screens (1510) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without a touch screen input function, each with or without a tactile feedback function, some of which may be able to output more than three dimensions, such as two-dimensional visual output or three-dimensional output by means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown), etc.), and may include printers (not shown).

[0231] The computer system (1500) may also include a human-accessible memory device, as well as optical media including a CD / DVD ROM / RW (1520) having a CD / DVD or similar media (1521), a thumb drive (1522), a removable hard drive or solid state drive (1523), legacy magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc., and related media of the memory device can also be included.

[0232] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.

[0233] The computer system (1500) can also include an interface (1554) to one or more communication networks (1555). The network can be, for example, wireless, wired, optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, and vehicle and industrial including CANBus. Certain networks generally require an external network interface adapter attached to a specific general-purpose data port or peripheral bus (1549) (such as a USB port of the computer system (1500)), and other networks are generally integrated into the core of the computer system (1500) by connection to the system bus described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1500) can communicate with other entities. Such communication can be unidirectional, receive only (such as television broadcast), transmit only unidirectional (such as CANbus to a specific CANbus device), or bidirectional to other computer systems using, for example, local area or wide area digital networks. Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0234] The aforementioned human interface device, human-accessible memory device, and network interface can be connected to the core (1540) of the computer system (1500).

[0235] The core (1540) can include one or more central processing units (CPUs) (1541), a graphics processing unit (GPU) (1542), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1543), a hardware accelerator for specific tasks (1544), a graphics adapter (1550), and the like. These devices can be connected via a system bus (1548) together with a read-only memory (ROM) (1545), a random access memory (1546), an internal mass storage such as a built-in hard drive or SSD that is not accessible to the user (1547), and the like. In some computer systems, the system bus (1548) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the system bus of the core (1548) or via a peripheral bus (1549). In one example, a screen (1510) can be connected to the graphics adapter (1550). Architectures for peripheral buses include PCI, USB, and the like.

[0236] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) can execute specific instructions that can, in combination, constitute the aforementioned computer code. That computer code can be stored in the ROM (1545) or RAM (1546). Temporary data can also be stored in the RAM (1546), and persistent data can be stored, for example, in the internal mass storage (1547). The use of cache memory, which can be closely associated with one or more CPUs (1541), GPUs (1542), mass storage (1547), ROM (1545), RAM (1546), etc., can enable fast storage and retrieval to any of the memory devices.

[0237] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or may be of the kind well-known and available to those having skill in the computer software arts.

[0238] By way of example and not limitation, a computer system (1500) having an architecture, specifically a core (1540), can provide functionality as a result of one or more processors (including, e.g., a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as described above, as well as specific storage of the core (1540) that is non-transitory in nature, such as core internal mass storage (1547) and ROM (1545). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1540). The computer-readable media can include one or more memory devices or chips, depending on specific requirements. The software can cause the core (1540), specifically the processors therein (including, e.g., a CPU, GPU, and FPGA), to define data structures stored in RAM (1546) and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally or alternatively, the computer system can provide functionality as a result of logic (e.g., an accelerator (1544)) that is wired or otherwise embodied in circuitry and operates instead of or in conjunction with software to execute specific processes or specific portions of specific processes described herein. References to software can, where appropriate, include logic, and vice versa. References to computer-readable media can, where appropriate, include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0239] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that are included within the scope of the present disclosure. Therefore, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of the present disclosure but are not explicitly shown or described herein and thus fall within the spirit and scope of the present disclosure.

Explanation of Reference Numerals

[0240] 100 Communication system 105 Sensor 110 Terminal device 120 Terminal device 150 Network 200 Streaming system 201 Point cloud source 202 Point cloud 203 Encoder 204 Compressed point cloud 205 Streaming server 206 Client subsystem 207 Copy of compressed point cloud 208 Client subsystem 209 Copy of compressed point cloud 210 Decoder 211 Reconstructed point cloud 212 Rendering device 213 Capture subsystem 220 Electronic device 230 Electronic device 300 V-PCC encoder 304 Patch information module 306 Patch generation module 308 Patch packing module 310 Geometry image generation module 312 Texture image generation module 314 Occupancy map module 316 Image padding module 318 Image Padding Module 320 Group Expansion Module 322 Video Compression Module 323 Video Compression Module 324 Multiplexer 332 Video Compression Module 334 Entropy Compression Module 336 Smoothing Module 338 Auxiliary Patch Information Compression Module 400 V-PCC Decoder 432 Demultiplexer 434 Video Decompression Module 436 Video Decompression Module 438 Occupancy Map Decompression Module 442 Auxiliary Patch Information Decompression Module 444 Geometry Reconstruction Module 446 Smoothing Module 448 Texture Reconstruction Module 452 Color Smoothing Module 510 Video Decoder 520 Parser 521 Symbol 551 Scaler / Inverse Transformation Unit 552 Intra-Picture Prediction Unit 553 Motion Compensation Prediction Unit 555 Aggregator 556 Loop Filter Unit 557 Reference Picture Memory 558 Current Picture Buffer 603 Video Encoder 630 Source Coder 632 Coding Engine 633 Decoder 634 Reference Picture Memory 635 Predictor 645 Entropy Coder 650 Controller 700 G-PCC Encoder 701 Input point cloud 702 Compressed bitstream 710 Position quantization module 712 Duplicate point removal module 720 Attribute transfer module 730 Octree encoding module 740 Level of detail (LOD) generation module 750 Attribute prediction module 760 Residual quantization module 770 Arithmetic coding module 780 Inverse residual quantization module 781 Addition module 790 Memory 800 G-PCC decoder 801 Compressed bitstream 802 Decoded point cloud 810 Arithmetic decoding module 820 Inverse residual quantization module 830 Octree decoding module 840 LOD generation module 850 Attribute prediction module 860 Memory 910 D mesh 920 2D atlas 1000 Framework 1005 Input mesh 1010 Mesh encoder 1020 Parametrization module 1021 XYZ & connectivity 1022 New texture map 1023 Auxiliary data 1025 New UV 1030 Texture conversion module 1031 Original texture map and original UV 1040 Reversible encoder 1041 Image / video encoder 1042 Auxiliary data encoder 1045-bit stream 1050 Mesh decoder 1060 Reversible decoder 1061 Image / Video decoder 1062 Auxiliary data decoder 1065 Decoded XYZ & Connectivity 1066 Decoded new texture map 1067 Decoded auxiliary data 1070 Parameterization module 1075 New UV 1080 Mesh reconstruction module 1095 Reconstructed mesh 1100 Framework 1105 Input mesh 1110 Mesh encoder 1120 Parameterization module 1121 XYZ & Connectivity 1122 New texture map 1123 Auxiliary data 1125 New UV 1130 Texture conversion module 1131 Original texture map and original UV 1140 Irreversible encoder 1141 Image / Video encoder 1142 Auxiliary data encoder 1144 Encoded XYZ & Connectivity 1145 Irreversible decoder 1145-bit stream 1146 Decoded XYZ & Connectivity 1150 Mesh decoder 1160 Irreversible decoder 1161 Image / Video decoder 1162 Auxiliary data decoder 1165 Decoded XYZ & Connectivity 1166 Decoded new texture map 1167 Decoded auxiliary data 1170 Parameterization module 1175 New UV 1180 Mesh reconstruction module 1195 Reconstructed mesh 1200 Syntax table 1210 First flag 1220 Second flag 1230 Third flag 1240 Index 1500 Computer system 1501 Keyboard 1502 Mouse 1503 Trackpad 1510 Touch screen 1505 Joystick 1506 Microphone 1507 Scanner 1508 Camera 1509 Speaker 1510 Touch screen 1520 CD / DVD ROM / RW 1521 Media 1522 Thumb drive 1523 Solid state drive 1540 Core 1541 Central processing unit (CPU) 1542 Graphics processing unit (GPU) 1543 Field programmable gate array (FPGA) 1544 Hardware accelerator 1545 Read-only memory (ROM) 1546 Random access memory 1547 Internal mass storage 1548 System bus 1549 Peripheral bus 1550 Graphics adapter 1554 Interface 1555 Communication network

Claims

1. Decoding the three-dimensional (3D) coordinates of vertices within a first 3D mesh frame and connection information of the vertices from a bitstream carrying the first 3D mesh frame, wherein the first 3D mesh frame represents the surface of an object with polygons; a step; Estimating texture coordinates associated with the vertices; a step; Decoding an index indicating a parameterization method selected from a list of parameterization method candidates; a step; Performing parameterization according to the 3D coordinates and the connection information of the vertices to determine the texture coordinates associated with the vertices; a step; Including; a step; Decoding a texture map of the first 3D mesh frame from the bitstream, wherein the texture map includes one or more first 2D charts having 2D vertices with the texture coordinates; a step; Reconstructing the first 3D mesh frame based on the 3D coordinates of the vertices, the connection information of the vertices, the texture map, and the texture coordinates Including; a method for mesh decompression.

2. The step of decoding the 3D coordinates of the vertices and the connection information of the vertices is performed by at least one of a reversible codec and an irreversible codec. The method according to claim 1.

3. The step of performing the parameterization according to the 3D coordinates and the connection information of the vertices to determine the texture coordinates associated with the vertices includes: Dividing the polygons of the first 3D mesh frame into the one or more first 2D charts and packing them into a 2D map; a step; Performing temporal alignment to align the one or more first 2D charts with one or more second 2D charts associated with a second 3D mesh frame, wherein the first 3D mesh frame and the second 3D mesh frame are frames within a 3D mesh sequence; a step; Further including the step of determining the texture coordinates from the one or more first 2D charts using the temporal alignment. The method according to claim 1.

4. The step of dividing and packing the polygon comprises The method according to claim 3, further comprising the step of dividing the polygon according to a normal value associated with the polygon.

5. The step of performing the temporal alignment comprises a scale-invariant metric, a rotation-invariant metric, a translation-invariant metric, and / or an affine transformation-invariant metric The method according to claim 3, further comprising the step of performing the temporal alignment according to at least one of them.

6. The step of performing the temporal alignment comprises the center of the chart calculated based on the 3D coordinates associated with the chart, the average depth of the chart, the weighted average texture value of the chart, and / or the weighted average attribute value of the chart The method according to claim 3, further comprising the step of performing the temporal alignment according to at least one of them.

7. A step of decoding a flag indicating activation of texture coordinate derivation, wherein the flag is at least one of a sequence level flag, a frame group level flag, and a frame level flag The method according to claim 1, further comprising.

8. The step of estimating the texture coordinates associated with the vertex comprises a step of decoding a flag indicating inheritance of the texture coordinates, a step of decoding an index indicating a 3D mesh frame selected from a list of decoded 3D mesh frames, and a step of inheriting the texture coordinates from the selected 3D mesh frame. The method according to claim 1.

9. A step of decoding a flag associated with a part of the texture map, wherein the flag indicates whether the UV coordinates associated with the part of the texture map are inherited from the decoded mesh frame or derived by parameterization The method according to claim 1, further comprising.

10. A step of decoding an index of a set of key vertices among the vertices, wherein the parameterization starts from the set of key vertices The method according to claim 1, further comprising.

11. An apparatus for mesh decompression configured to perform the method according to any one of claims 1 to 10.

12. A computer program for causing a computer to execute the method according to any one of Claims 1 to 10.

Citation Information

Patent Citations

  • Compression of texture rendered wire mesh models

    US20130106834A1

  • Encoding and decoding of texture mapping data in textured 3D mesh models

    US20180253867A1