Learning-based mesh compression using latent mesh
Patent Information
- Application Number
- EP2024799758
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-10-16
- Publication Date
- 2026-09-09
AI Technical Summary
Current methods for compressing 3D meshes are inefficient, either sacrificing reconstruction quality for high bitrate consumption or vice versa, and do not effectively balance rate and distortion.
The proposed learning-based mesh compression system transforms the base mesh into a latent mesh, which is then encoded and decoded using a novel encoder and decoder architecture. This approach allows for optimized coding that balances bitrate and reconstruction quality.
The latent mesh compression method achieves improved bitrate-distortion tradeoff, providing better reconstruction quality at lower bitrates compared to traditional methods.
Smart Images

Figure US2024051561_08052025_PF_FP_ABST
Abstract
Description
LEARNING-BASED MESH COMPRESSION USING LATENT MESHTECHNICAL FIELD[1] The present embodiments generally relate to a method and an apparatus for 3D mesh compression.BACKGROUND[2] 3D point cloud is a universal data format across several business domains from autonomous driving, robotics, AR / VR, civil engineering, computer graphics, to the animation / movie industry. 3D LiDAR sensors have been widely deployed in self-driving cars. Affordable LiDAR sensors are released from Velodyne Velabit, Apple iPad Pro 2020 and Intel RealSense LiDAR camera L515. With great advances in sensing technologies, 3D point cloud data has become more practical than ever and is expected to be an ultimate enabler in the applications mentioned.[3] Point cloud data is also believed to consume a large portion of network traffic, e.g., among connected cars over 5G network, and immersive communications (VR / AR). Point cloud understanding and communication would essentially lead to efficient representation formats. In particular, raw point cloud data need to be properly organized and processed for the purposes of world modeling and sensing.[4] Furthermore, point clouds may represent a sequential scan of the same scene, which contains multiple moving objects. They are called dynamic point clouds as compared to static point clouds captured from a static scene or static objects. Dynamic point clouds are typically organized into frames, with different frames being captured at different time.[5] Mesh is a data format that is composed of both points (also known as vertex) and faces. Comparing to point clouds, mesh is assumed a complete modeling of 3D surfaces, while the points in point clouds are only a set of discretized samples of the surface.[6] For Computer Graphics (CG), mesh is typically available as the data is generated by the content providers. For data acquired by sensors, including LiDAR, the raw data is typically in point cloud format. For those point cloud data, mesh can be created via a meshing procedure.SUMMARY i[7] According to an embodiment, a method for decoding a mesh is presented, comprising: decoding a feature map; decoding a latent mesh, including: decoding a face list for a latent mesh, obtaining a predicted point list for a point list of said latent mesh, decoding residual information associated with said point list of said latent mesh, and reconstructing said point list for said latent mesh based on said residual information and said predicted point list for said point list of said latent mesh; and decoding said mesh based on said decoded feature map and said decoded latent mesh based on a first neural network.[8] According to another embodiment, a method for encoding a mesh is presented, comprising: obtaining a feature map and a base mesh corresponding to said mesh based on a fourth neural network; encoding said feature map; transforming said base mesh to a latent mesh; and encoding said latent mesh, including: encoding a face list of said latent mesh, obtaining a predicted point list for a point list of said latent mesh, obtaining residual information associated with said point list, and encoding said residual information for said point list.[9] According to another embodiment, an apparatus for decoding a mesh is presented, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to: decode a feature map; decode a latent mesh, wherein said one or more processors are configured to: decode a face list for a latent mesh, obtain a predicted point list for a point list of said latent mesh, decode residual information associated with said point list of said latent mesh, and reconstruct said point list for said latent mesh based on said residual information and said predicted point list for said point list of said latent mesh; and decode said mesh based on said decoded feature map and said decoded latent mesh based on a first neural network.
[0010] According to another embodiment, an apparatus for encoding a mesh is presented, comprising one or more processors and at least one memory coupled to the one or more processors, wherein said one or more processors are configured to: obtain a feature map and a base mesh corresponding to said mesh based on a fourth neural network; encode said feature map; transform said base mesh to a latent mesh; and encode said latent mesh, wherein said one or more processors are configured to: encode a face list of said latent mesh, obtain a predicted point list for a point list of said latent mesh, obtain residual information associated with said point list, and encode said residual information for said point list.
[0011] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the encoding method or decoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding a mesh according to the methods described herein.
[0012] One or more embodiments also provide a computer readable storage medium having stored thereon data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the data generated according to the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented.
[0014] FIG. 2 illustrates a diagram of a learning-based mesh compression system.
[0015] FIG. 3 illustrates a diagram of a learning-based mesh compression system using base graph.
[0016] FIG. 4 illustrates a diagram of the wrapping block.
[0017] FIG. 5 illustrates face convolution kernels.
[0018] FIG. 6 illustrates a residual network diagram for point position update in latent mesh.
[0019] FIG. 7 illustrates a diagram of a learning-based mesh compression system using latent mesh, according to an embodiment.
[0020] FIG. 8 illustrates a diagram of the wrapping block.
[0021] FIG. 9 illustrates a diagram of the inverse wrapping block.
[0022] FIG. 10 illustrates a diagram to encode the latent mesh, according to an embodiment.
[0023] FIG. 11 illustrates parallelogram-based predictor generation.
[0024] FIG. 12 illustrates a diagram to decode the latent mesh, according to an embodiment.
[0025] FIG. 13 illustrates a diagram to encode the latent mesh with hyper prior model, according to an embodiment.
[0026] FIG. 14 illustrates a diagram to decode the latent mesh with hyper prior model, according to an embodiment.
[0027] FIG. 15 illustrates a diagram to encode the latent mesh with auto-regressive hyper prior model, according to an embodiment.
[0028] FIG. 16 illustrates a diagram to decode the latent mesh with auto-regressive hyper prior model, according to an embodiment.
[0029] FIG. 17 illustrates a diagram of a feature encoder.
[0030] FIG. 18 illustrates subdivision - downsampling.
[0031] FIG. 19 illustrates a diagram of a feature decoder.
[0032] FIG. 20 illustrates subdivision - upsampling.
[0033] FIG. 21 illustrates a non-manifold mesh.DETAILED DESCRIPTION
[0034] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.
[0035] The system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitriesas known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device, and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.
[0036] System 100 includes an encoder / decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.
[0037] Program code to be loaded onto processor 110 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0038] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fastexternal dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, JPEG Pleno, MPEG-I, V-DMC, ETEVC, or VVC.
[0039] The input to the elements of system 100 may be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0040] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, bandlimiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog- to-digital converter. In various embodiments, the RF portion includes an antenna.
[0041] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or withinprocessor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device.
[0042] Various elements of system 100 may be provided within an integrated housing, Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.
[0043] The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.
[0044] Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802. 11. The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for Wi-Fi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.
[0045] The system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.
[0046] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0047] The automotive industry and autonomous car are domains in which point clouds may be used. Autonomous cars should be able to “probe” their environment to make good driving decisions based on the reality of their immediate surroundings. Typical sensors like LiDARs produce (dynamic) point clouds that are used by the perception engine. These point clouds are not intended to be viewed by human eyes and they are typically sparse, not necessarily colored, and dynamic with a high frequency of capture. They may have other attributes like the reflectance ratio provided by the LiDAR as this attribute is indicative of the material of the sensed object and may help in making a decision.
[0048] Virtual Reality (VR) and immersive worlds are foreseen by many as the future of 2D video. The basic idea is to immerse the viewers in an environment all around them as opposed to standard TV where they can only look at the virtual world in the front. There are several gradations in the immersivity depending on the freedom of the viewer in the environment. Point cloud as well as mesh are good format candidates to distribute VR contents. They may be static or dynamic and are typically of moderate size, no more than millions of points at a time.
[0049] Point clouds and meshes may be also used for various purposes such as culture heritage / buildings in which objects like statues or buildings are scanned in 3D in order to share the spatial configuration of the object without sending or visiting them. Also, it is a way to ensurepreserving the knowledge of the object in case it may be destroyed; for instance, a temple by an earthquake. Such point clouds and meshes are typically static, colored, and huge.
[0050] World modeling and sensing via point clouds / meshes could be an essential technology to allow machines to gain knowledge about the 3D world around them, which is crucial for the applications discussed above.
[0051] 3D point cloud data are essentially discrete samples on the surfaces of objects or scenes. To fully represent the real world with point samples, in practice it requires a huge number of points. For instance, a typical VR immersive scene contains millions of points, while point clouds typically contain hundreds of millions of points. Therefore, the processing of such large-scale point clouds is computationally expensive, especially for consumer devices, e.g., smartphone, tablet, and automotive navigation system, that have limited computational power. Additionally, the discrete samples comprising the 3D point cloud data still contain incomplete information about the underlying surfaces of objects and scenes. Hence, recently efforts are being made to also explore mesh representation for 3D scene / surface representation.
[0052] Meshes can be considered as a 3D point cloud along with the connectivity information between the points. Thus, the mesh representation bridges the gap between point clouds and the underlying continuous surfaces through local 2D polygonal patches (called face) approximation of the underlying surface. In another view, mesh can simplify the point cloud by reducing the number of points. For example, a face in a mesh can be appropriate to represent a flat area while point clouds may require a dense block of points to represent the same flat area.
[0053] The first step for any processing or inference on the mesh data is to have efficient storage methodologies. To store and process the mesh with affordable computational cost, one solution is to down-sample it first, where the down-sampled mesh summarizes the geometry of the input mesh while having much fewer (but bigger) faces. The down-sampled mesh is then fed to the subsequent machine task for further consumption. However, further reduction in storage space can be achieved by converting the raw mesh data (original or downsampled) into a fixed length codeword or a feature map living on a very low-resolution mesh. This codeword or the feature map can be converted to a bitstream through existing entropy coding techniques. Moreover, the codeword or feature map is also useful by itself as it represents global or local surface information (respectively) of the underlying scene / object and can be paired with subsequent downstream (machine vision)tasks.
[0054] FIG. 2 illustrates a typical diagram of a learning-based mesh compression system that is composed of two branches. The main branch is for feature coding. The main branch encoder is to extract and aggregate features from the input mesh using a feature extractor (210). The extracted feature map F is sent to an entropy coder (230, ENCF) to form bitstream BSF. The feature map F decoded by an entropy decoder (250, DECF) is sent to a feature decoder (260) to reconstruct the mesh. With the feature extraction / aggregation, the input mesh is typically downsampled to a base mesh M with a lower resolution for the benefit of compression. The feature map F is associated with the base mesh M.
[0055] The side branch of the mesh compression in FIG. 2 is to code the base mesh M. It is composed of an ENCM block (220) for encoding into to form bitstream BSM, and a DECM block (240) for decoding. The encoding and decoding of the base mesh involve coding a list of faces. With triangle mesh case, each face is represented by a list of three connected points. Each point is described by a 3D coordinate. In one example, we let M = (X, T), where X is the list of points, and T is the list of faces. Each point in a face is represented by its index in X. To code a base mesh, one needs to code the point list X and the face list T.
[0056] FIG. 3 shows a high-level diagram of another related work for learning-based mesh compression, which is a graph-based mesh compression approach as described in a commonly owned patent application entitled “Base Graph based Mesh Compression” (U.S. Provisional Application No. 63 / 533,303, Attorney Docket No. 2023PF00502). Comparing to the diagram in FIG. 2, the base mesh is transformed to a new representation known as base graph before being coded. On the encoder side, the transform is done via a wrapping block WNG (310, WrappingNet). It attempts to disentangle the topology information from the geometry information in the base mesh. Hence the base graph intends to convey the topology information only. On the decoder side, an inverse wrapping IWNG (320) is conducted to transform the base graph back to a base mesh. The entropy encoding block ENCG (330) and entropy decoding block DECG (340) are updated accordingly to handle the base graph and the corresponding bitstream BSG rather than the base mesh. We let G = (A, T), where N is the list of graph nodes, and T is the face list. The face list is the same as defined in the base mesh. The graph node list is computed via a wrapping block WNG.
[0057] The wrapping block WNG is composed of two steps as shown in FIG. 4. It first deforms the input mesh M to a sphere mesh Msusing a “deforming” block (410). In one embodiment, the shape of the sphere mesh is set to be deformed close to a sphere. Then the points in the sphere mesh are matched to sphere grid points (5) via a sphere matching block (420). The sphere grid points are pre-sampled over the sphere and remain fixed. Finally, the wrapping block outputs a graph node list G.
[0058] The “deforming” block in FIG. 4 is further composed of two structures to generate the sphere mesh Ms. One is a CNN network (411, 413) to aggregate the features on faces. Then the point positions of the input mesh are deformed by a position update block (PU, 412, 414).
[0059] The first CNN structure takes mesh M as its input. In one embodiment as shown in FIG. 4, the CNN+PU structure is repeated twice to enhance the deformation. After a first CNN+PU structure (411, 412), the mesh M is downsampled to Mr, while a feature map F associated with the downsampled mesh M is computed. In one example, the feature aggregation is performed on faces rather than points. Hence, the CNN is selected to run convolutions over neighboring faces, also known as face convolution layers. CNN can be exploited in case of manifold mesh with a regular spatial neighborhood. Regular face convolutional (CNN) kernels shown in FIG. 5 are in use. For example, a 3-neighbor, a 6-neighbor, or 9-neighbor kernel.
[0060] The position update PU block is designed as shown in FIG. 6 that is mainly composed of a Residual Update block (RU, 610) and an addition operation (620) between the input position and a displacement. The RU block (610) first performs a face to node feature conversion via a learnable network, and then computes a displacement to update the point position.
[0061] A sphere matching loss is used to supervise the training of the deforming block. In one embodiment, the sphere matching is defined as the Chamfer distance between the points of the deformed mesh and randomly sampled point clouds from the unit sphere.
[0062] In the following, we propose a method to handle the coding of base mesh in a learningbased mesh compression system.
[0063] FIG. 7 illustrates the overall diagram of the proposed learning-based mesh compression, according to an embodiment. The proposed learning-based mesh codec is composed of two branches. In the main branch, features are extracted / aggregated based on the input mesh Mesh.A feature map F is produced (710), and entropy encoded by ENCF (760) to form bitstream BSF and decoded by DECF (770). The feature map is associated with a base mesh M . The processing of the base mesh is the second branch of the proposed codec.
[0064] In particular, in the second branch, the base mesh M is transformed to a latent mesh MLvia a wrapping block WNi. (720, WrappingNet). The latent mesh MLis entropy encoded by ENCi. (730) to form bitstream BSL and decoded by DECL (740). On the decoder side, an inverse wrapping IWNL (750) is conducted to transform the decoded latent mesh MLback to a reconstructed base mesh M. The decoded feature map F and reconstructed base mesh M are sent to a feature decoder (780) to reconstruct the mesh Mesh.
[0065] Comparing to the typical learning-based mesh compression shown in FIG. 2, the proposed approach won’t code the base mesh directly. It is additionally different from the system shown in FIG. 3, as the base mesh is not to be transformed to a base graph. Note that from the base mesh to the base graph, there is an intention to fully disentangle the topology information. Instead, here it is proposed to transform the base mesh to a novel format denoted as latent mesh in this document.
[0066] On the encoder side, the base mesh is deformed via a novel WNL block (720). From a high-level perspective, the latent mesh could be viewed as a tradeoff format between base mesh and base graph. We want to address the following problems:• With direct coding of the base mesh, it favors the reconstruction quality, but can result in undesired high bitrate consumption.• With coding of the base graph, it can favor the coding bitrate of the base mesh, however, it can sacrifice too much on the reconstruction quality.
[0067] With the proposed latent mesh ML, we suggest a novel way to control the tradeoff and use the latent mesh for an optimized coding with respect to both rate and distortion.
[0068] Latent Mesh Generation
[0069] As described before, the base mesh is represented as M = (X, T), where X is the point list, and T is the face list. The base graph is represented as G = (N, T), where N is the list of graph nodes, and T is the same face list in base mesh. To disentangle topology from the base mesh, the point list is replaced by the graph node list. The graph nodes N is denoted by the indices of the grid points.
[0070] Similar to base mesh, the proposed latent mesh MLis composed of a point list and a face list, ML= (XL, T)" . The face list T in latent mesh MLis the same face list as defined in the base mesh M. However, the point list XLis novelly deformed from the point list X.
[0071] In one embodiment, the deformation block WNL (720) can be designed as shown in FIG. 8 that is similar to WNG in FIG. 4. Note that in FIG. 8, the sphere matching block (420) in FIG. 4 is removed. It (420) was designed to match the points in the sphere mesh to sphere grid points. The output of the deformation block WNL is a deformed mesh, referenced as a (novel) latent mesh ML. The deformation for the latent mesh MLis conducted differently than for the sphere mesh Ms. The deformation to a sphere mesh (410) targets a shape close to a sphere. It (410) was implemented via a training based on a sphere loss. In case of the latent mesh, a trade-off on the deformation is proposed.
[0072] To make the latent mesh MLto be in-between the base mesh and the base graph, a specific training loss is proposed instead of the Chamfer distance in sphere loss. In particular, a proposed loss function is composed of two items. A first item reflects the distortion measured between the base mesh and the latent mesh. A second item is the rate required to code the latent mesh. The rate can be estimated using an entropy bottleneck layer as proposed in a commonly owned patent application entitled “Entropy Bottleneck Layer for Learning-Based Mesh Compression” (U.S. Provisional Application No. 63 / 595,559, our Attorney Docket No. 2023PF00974). The overall loss function is a weighted combination of the distortion term and the rate term:Loss = Distortion + ■ Rate.
[0073] The deformation degree is determined by the tradeoff between the rate and distortion. In other words, the hyperparameter A used to hit a target rate point and distortion can implicitly guide the deformation degree.
[0074] Deformation Network Structure
[0075] Like the deforming block in FIG. 4, the proposed wrapping (deformation) block WNL in FIG. 8 is composed of two structures. One is a CNN network (810, 830) to aggregate the features on faces. After feature aggregation, the point positions of the input mesh are updated by a PU block (820, 840).
[0076] In addition to mesh M, the first CNN structure (810) takes an initial feature F as its input.In one embodiment as shown in FIG. 8, the two structures (CNN and PU) are repeated twice to enhance the wrapping (deformation). In another embodiment, CNN and PU may be repeated more than twice. In one example, the feature aggregation is performed for faces rather than points. Hence, the CNN is selected to run convolution over neighboring faces, and this layer is also known as a face convolution layer. CNN can be exploited in case of manifold mesh with a regular spatial neighbor relationship. In other words, the regular face convolutional (CNN) kernels as shown in FIG. 5, with a 3-neighbor, a 6-neighbors, or 9-neighbor kernel can be used. The very first feature input is the 7-dimensional features containing the area (1-dim), normal vector (3-dim), and curvature vector (3-dim) of each face. The first face convolution layer maps to a hidden dimension of 128 channels, and the second face convolution layer maintains 128 dimensions. The PU block uses a face-wise 2-layer MLP (MultiLayer Perceptron) with input-output dimensions (128, 128, 128) to generate the displacement vector outputs.[771 The position update PU block can be designed as shown in FIG. 6 that is mainly composed of a RU block and an addition operation between the input position and a displacement. The RU block first performs a face to node feature conversion via a learnable network, and then computes a displacement to update the point position. For example, the RU block applies a face-wise 2- layer MLP with dimension (128, 128, 131), to an output size of 131. First 3 entries are used as the displacement vector, and the last 128 entries are outputted as the updated features. The 3- dimensional displacement is added to the positions of the point list X of the input mesh.
[0078] The corresponding inverse wrapping (deformation, 750) is shown in FIG. 9. The diagram is simply an inverse diagram of FIG. 8. The dimensions are identical to the modules described in FIG. 8, except that at the input to the module, we concatenate the 7-dimensional fixed features, i.e., area (1-dim), normal vector (3-dim), and curvature vector (3-dim) with the learned feature map of dimension 128. Hence the first face convolution layer (910) is (135, 128). The remaining face convolution layers (930) and 2-layer MLPs in the PU blocks (920, 940) have input, hidden, and output dimensions of 128.
[0079] It should be noted that the size of latent mesh (the number of faces) plays an important role in determining the rate point of the learning-based mesh codec besides the hyper parameter A presented above. The size of latent mesh is typically determined by the number of faces. In addition, the latent mesh size can be configured as an input parameter to run the codec. With asmaller latent mesh size, the codec tends to operate on a lower bitrate. If higher quality of reconstruction is desired, the latent mesh size (and coarse faces) should be increased. With a larger latent mesh size (and fine faces), more geometric (shape) details are losslessly maintained.
[0080] Encoding of Latent Mesh
[0081] The latent mesh is proposed to be coded using a novel encoder (730) as shown in FIG. 10, according to an embodiment.
[0082] The latent mesh MLis first quantized using a Q block (1010) based on an input quantization parameter QP. The quantization is conducted on the list of points. The reconstructed point position from the quantization / dequantization is represented by X . The face list T is losslessly coded (1020) into bitstream. The quantized point list X is to be coded following the procedure below. Note that we don’t differentiate the quantized and dequantized point in the document. When being referenced for coding, X represents quantized values. When being referenced for their positions, X represents the dequantized values.
[0083] First the input latent mesh is being serialized (1040). The serialization has the mesh organized into Ms, a ID sequence of points, representing a coding order of the points. The serialization is implemented by a process known as edgebreaker. This edgebreaker serialization process operates as a state machine. In each state, it moves from one triangle to an adjacent one in a spiral-like order. Once passed through the adjacent, the next triangle would be either left or right since the remaining triangle is the one passed on the current state. During this traversal, all visited triangle and bounding vertices are marked sequentially. Assuming that T is coded losslessly, the decoder can reproduce the exact same traversal order. The organized sequence Msalso maintains a prediction table indicating which points are used as references to code a next point. For a mesh, each point in the mesh is associated with a position and an index, and each face is defined by three points (e.g., using their indices). In one example, with serialization (e.g., edgebreaker), the point indices are stored into an array indicating their coding order, and at the same time, a prediction relationship is stored.
[0084] It is worthy to note that the “prediction table” can be virtual and be generated on the fly when the edgebreaker runs over the fact list T. The exact same prediction table can also be reproduced by the decoder on the fly.[851 For each point in the organized sequence Ms(when reference points are available), a predictive coding (1030) is applied to code a next point position. The three reference points are first accessed as per the prediction table. In one embodiment, a predictor XPis computed (1030) as per the following parallelogram equation:XP= X0+ (X2- X1)
[0086] This is further illustrated in FIG. 11. Xo, X and X2are reference points, and Xp is the predicted point, where we want to compute the residual from. Once a predictor is computed, a residual EN— XN— XPis computed (1060), where XNis the current point, and the current triangle moves from X0X1X2to the next triangle XQX2Xp.
[0087] Lastly in FIG. 10, an arithmetic encoding (1050) is conducted on the residual EN. The arithmetic encoding is guided by a probability distribution parameter set P . The distribution parameter is obtained from the training stage. They are assumed to be available when the encoding procedure is invoked.
[0088] During the training stage, the distribution parameter set P is initialized with random values. When performing backward propagation in training, the distribution parameter set P is updated in the same way as other neural network parameters.
[0089] Decoding of Latent Mesh
[0090] The latent mesh decoding (740) is shown in FIG. 12, according to an embodiment. From the bitstream BSL, we first run an arithmetic decoder (1240) to reconstruct the face list T. From the face list, a serialization (1250) is performed in a similar manner as in the encoder. Through the serialization, a prediction table is created and the points to be decoded are assigned with indices.
[0091] Next is to decode the point list. The distribution parameter set P estimated from the training stage is also available at the decoder. It will guide the arithmetic decoding (1210) of the residual EN. Then, according to the prediction table, three reference points (0,1,X2)areretrieved to compute (1230) a predictor for the current point XP= Xo+ (X2—^i) The current point is computed (1220) XN= XP+ ENto generate quantized point list X. Based on the decoded face list T and quantized point list X , inverse quantization (1260) is performed to obtain dequantized point list X (note that we use the same notation X for quantized and dequantized pointlists) for reconstructed mesh ML.
[0092] Latent Mesh Coding with Hyper Prior
[0093] In FIGs. 13 and 14, we propose a latent mesh encoder and decoder with a hyper prior, according to an embodiment. Similar to previous implementations, the latent mesh codec uses an arithmetic coding guided by a probability distribution parameter set. However, the parameter of the distribution used in previously presented codec is estimated offline during the training stage. In this embodiment, we propose an advanced approach to estimate the distribution parameter on the fly during the inference stage - that is, during the encoding / decoding time.
[0094] An EDP block (1310, 1410) is inserted in FIGs. 13 and 14 to estimate the distribution parameter set on the fly. On the encoder side, the EDP (1310) takes the residual information as its input, and outputs an estimated distribution parameter set. In one embodiment, the residual information is up to the current point for a low latency configuration. In another embodiment, the residual information is for the entire point list for a high coding efficiency configuration that allows a longer latency. At the same time, a bitstream is encoded to represent the distribution parameter set, also known as hyper prior parameters. A dedicated bitstream is to encode the hyper prior parameters. The estimated distribution parameters are used to guide the arithmetic coding of the residual information.
[0095] On the decoder side, the hyper prior parameters are decoded from the bitstream via a DP block (1410). Then the hyper prior parameter (distribution parameter set) is used to help the arithmetic decoding of the residual information.
[0096] Latent Mesh Coding with Auto-regressive Hyper Prior
[0097] In FIGs. 15 and 16, we propose a latent mesh encoder and decoder using a predicted hyper prior, according to an embodiment. Comparing to the codec shown in FIGs. 13 and 14, in this embodiment, we don’t use a bitstream to transmit the hyper priors. Instead, both the encoder and decoder use an EDP block (1510, 1610) to predict the distribution parameter set. The prediction done by the EDP here is to use previously coded points instead of the point to be coded next as in FIG. 13.
[0098] In one embodiment, we use the residual information from the 3 reference points identified from the prediction table to predict the distribution parameters. The EDP design is shared in bothencoder (FIG. 15) and decoder (FIG. 16).
[0099] Feature Encoder and Feature Decoder
[0100] In this document, we provide some example designs of the feature encoder and decoder to make the overall codec description complete.
[0101] FIG. 17 shows an example feature encoder design. The feature encoder is basically a feature extraction and aggregation network. It would generate a feature map that is associated with a base mesh. In this example, we assume the input mesh is semi-regular mesh, that is, every face in the input mesh has three neighboring faces.
[0102] The feature encoder is composed of two basic blocks. One is a face convolution block (1710, 1730), and the other is a subdivision-based pooling (downsampling, 1720, 1740). The face convolution is defined by a selected face convolutional kernel, e.g., 3-neighbor kernel, 6-neighbor kernel or 9-neighbor as shown in FIG. 5. The subdivision pooling is to merge the 3 faces connected to a current face from input as one face, as illustrated in FIG. 18. For example, the first face convolution layer has dimensions (7, 128), where the input features are the area, normal vector, and curvature vectors of M. The following face convolution layers are all of size (128, 128).
[0103] FIG. 19 shows an example feature decoder design. The feature decoder is an inverse procedure of the feature encoder. It takes a feature map and corresponding base mesh as input and output a reconstructed mesh. In one embodiment, it is composed of two basic blocks. One is face convolution block (1910, 1930), and the other is a subdivision-based upsampling (1920, 1940). The subdivision upsampling is to divide a single face from input into 4 faces as an output by breaking every existing edge into two and placing a new node at the midpoint of the edge, as illustrated in FIG. 20. For example, the input feature F is again concatenated with the fixed 7- dimensional features (area, normal, curvature) of M. The first face convolution layer is size (135, 128), with following ones of size (128, 128).
[0104] Graph Convolutional Network (GCN)
[0105] In the selection of the CNN for feature encoding / decoding, and latent mesh encoding / decoding, we presented how regular CNNs are in use for face convolutions. In another embodiment, the CNNs are implemented using graph convolutional network (GCN).
[0106] The implementation of CNNs requires the mesh to maintain a regular spatial neighborhood.That is, mesh has to follow a regular pattern how faces are connected to each other. For example, in case of triangle meshes, each face typically has always three immediate neighboring faces unless they are along the boundary of the surfaces.
[0107] Moreover, it is easy to run regular CNNs on a manifold mesh, but not easy to run over a non-manifold mesh. With a manifold mesh, one edge can only appear in two faces at most. In a non-manifold mesh, one edge can however appear in more than two faces. For example, an edge from the intersection of two manifolds can be in at least three faces, as illustrated in FIG. 21. The mesh has two manifolds. One manifold has faces (%0^i^2) and X2,X3,X0'), and the other manifold has face (Ao, X2, V4) . Then the edge (Xo, X2) is being used in all the three faces.
[0108] When regular CNNs have difficulties in such cases, we propose using graph convolutional network (GCN) to perform the convolutional computation. Such GCN can support more flexible spatial relationship between the neighboring faces (or points). It first builds a graph structure describing the spatial topology over the faces. Then the convolution is computed along the graph edges. GCN is a known replacement for regular CNN.
[0109] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.[HO] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.[Hl] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (forexample, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.
[0112] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.
[0113] Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0114] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0115] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0116] It is to be appreciated that the use of any of the following “ / ”, “and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.
[0117] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted to carry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.
Claims
CLAIMS1. A method for decoding a mesh, comprising: decoding a feature map; decoding a latent mesh, including: decoding a face list for a latent mesh, obtaining a predicted point list for a point list of said latent mesh, decoding residual information associated with said point list of said latent mesh, and reconstructing said point list for said latent mesh based on said residual information and said predicted point list for said point list of said latent mesh; and decoding said mesh based on said decoded feature map and said decoded latent mesh based on a first neural network.
2. The method of claim 1, further comprising: transforming said decoded latent mesh to a reconstructed base mesh, wherein said mesh is decoded based on said decoded feature map and said reconstructed base mesh.
3. The method of claim 2, wherein said face list of said latent mesh is used as a face list of said reconstructed base mesh.
4. The method of claim 2 or 3, wherein said transforming said decoded latent mesh comprises: aggregating features on faces of said decoded latent mesh based on a second neural network; and updating point positions of said decoded latent mesh using a third neural network.
5. The method of claim 4, wherein said third neural network includes a residual update block that performs: converting a face to a node feature via a learnable network; and obtaining displacement information to update point positions of said decoded latent mesh.
6. The method of claim 5, wherein said residual update block includes a face-wise MLP (MultiLayer Perceptron).
7. The method of any one of claims 4-6, wherein said second neural network includes a Convolutional Neural Network or a Graph Convolutional Network.
8. The method of any one of claims 1-7, wherein said obtaining a predicted point list comprises: performing a serialization on said decoded face list for said latent mesh to obtain an ordered list of points and generate a prediction table, wherein a point is predicted is based on previous points in said ordered list of points indicated by said prediction table.
9. The method of any one of claims 1-8, wherein said decoding residual information associated with said point list of said latent mesh comprises: decoding a probability distribution parameter, wherein said residual information associated with said point list of said latent mesh is decoded based on said probability distribution parameter.
10. The method of any one of claims 1-8, wherein said decoding residual information associated with said point list of said latent mesh comprises: obtaining a probability distribution parameter based on previously decoded residual information, wherein said residual information for said point list of said latent mesh is decoded based on said probability distribution parameter.
11. The method of any one of claims 1-10, wherein said reconstructing said point list for said latent mesh comprises: inverse quantizing said point list for said point list of said latent mesh.
12. A method for encoding a mesh, comprising: obtaining a feature map and a base mesh corresponding to said mesh based on a fourth neural network; encoding said feature map; transforming said base mesh to a latent mesh; and encoding said latent mesh, including:encoding a face list of said latent mesh, obtaining a predicted point list for a point list of said latent mesh, obtaining residual information associated with said point list, and encoding said residual information for said point list.
13. The method of claim 12, wherein a face list of said base mesh is used as said face list of said latent mesh.
14. The method of claim 12 or 13, wherein said transforming said base mesh comprises: aggregating features on faces of said base mesh based on a fifth neural network; and updating point positions of said base mesh to form said latent mesh based on a sixth neural network.
15. The method of claim 14, wherein said sixth neural network includes a residual update block that performs: converting a face to a node feature via a learnable network; and obtaining displacement information to update point positions of said base mesh.
16. The method of claim 15, wherein said residual update block includes a face-wise MLP (MultiLayer Perceptron).
17. The method of any one of claims 14-16, wherein said fifth neural network includes a Convolutional Neural Network or a Graph Convolutional Network.
18. The method of any one of claims 12-17, further comprising obtaining a size of said latent mesh.
19. The method of any one of claims 12-18, wherein said obtaining a predicted point list comprises: quantizing said point list of said latent mesh; and performing a serialization on said face list of said latent mesh to obtain an ordered list of points and generate a prediction table, wherein a point is predicted is based on previous points in said ordered list of points indicated by said prediction table.
20. The method of any one of claims 12-19, wherein said encoding residual information comprises: encoding a probability distribution parameter, wherein said residual information for said point list is encoded based on said probability distribution parameter.
21. The method of claim 20, wherein said probability distribution parameter is obtained based on said residual information.
22. The method of any one of claims 12-19, wherein said encoding said residual information associated with said point list comprises: obtaining a probability distribution parameter based on previously encoded residual information, wherein said residual information for said point list is encoded based on said probability distribution parameter, and wherein said probability distribution parameter is not transmitted to a decoder.
23. The method of any one of claims 12-22, further comprising: obtaining a parameter controlling a tradeoff between a rate to encode said latent mesh and a distortion between said latent mesh and said base mesh; estimating said rate based on an entropy bottleneck layer; and training a wrapping network based on said parameter controlling said tradeoff, said rate and said distortion, wherein said wrapping network is used to transform said base mesh to said latent mesh.
24. The method of claim 23, further comprising: obtaining a probability distribution parameter during said training, wherein said residual information for said point list is encoded based on said probability distribution parameter.
25. An apparatus, comprising one or more processors and at least one memory coupled to said one or more processors, wherein said one or more processors are configured to perform the method of any of claims 1-24.
26. A non-transitory computer readable medium comprising instructions which, when theinstructions are executed by a computer, cause the computer to perform the method of any of claims 1-24.