Heterogeneous grid automatic encoder
The heterogeneous grid automatic encoder solves the problem of inefficient point cloud data processing by generating and matching grid features, and improves data utilization efficiency and quality, especially its application effect in autonomous driving and computer graphics.
Patent Information
- Application Number
- CN202380087229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-03
- Filing Date
- 2023-11-09
- Publication Date
- 2025-07-25
AI Technical Summary
Existing point cloud data processing methods have problems such as inefficiency and insufficient data processing in the fields of autonomous driving, robotics, augmented reality/virtual reality, civil engineering and computer graphics, especially in the utilization of 3D lidar sensor data.
Using heterogeneous grid automatic encoder, a fixed-length codeword is generated by generating initial mesh surface features, basic mesh features and matching index sets, and a reconstructed mesh is generated by matching predefined template mesh with basic mesh.
It improves the processing efficiency and quality of point cloud data, and enhances application capabilities in different fields, especially in autonomous driving and computer graphics.
Smart Images

Figure CN120380504A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is an international application that claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 424,421, titled "HETEROGENEOUS MESH AUTOENCODERS", filed on November 10, 2022, and U.S. Provisional Patent Application Serial No. 63 / 463,747, titled "LEARNING BASED HETEROGENEOUS MESH AUTOENCODERS", filed on May 3, 2023, under 35 U.S.C.§119(e), and each patent application is hereby incorporated by reference in its entirety.
[0003] Incorporation by reference
[0004] This application also incorporates by reference in its entirety the following applications: International Application No. PCT / US2021 / 034400, titled "METHODS, APPARATUS AND SYSTEMS FOR GRAPH - CONDITIONED AUTOENCODER (GCAE) USING TOPOLOGY - FRIENDLY REPRESENTATIONS" (the "400 application"), filed on May 27, 2021, which claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 047,446, titled "METHODS, APPARATUS AND SYSTEMS FOR GRAPH - CONDITIONED AUTOENCODER (GCAE) USING TOPOLOGY - FRIENDLY REPRESENTATIONS", filed on July 2, 2020, under 35 U.S.C.§119(e); these applications are hereby incorporated by reference in their entirety. Background of the Invention
[0005] The point cloud (PC) data format is a common data format across multiple business domains (e.g., autonomous driving, robotics, augmented reality / virtual reality (AR / VR), civil engineering, computer graphics, and the animation / movie industry). 3D lidar (light detection and ranging) sensors have been deployed in autonomous vehicles, and affordable lidar sensors are available. With the advancement of sensing technology, 3D point cloud data has become more practical than before. Summary of the Invention
[0006] The embodiments described herein include methods used in video encoding and decoding (collectively referred to as "code processing").
[0007] According to a first example method of some embodiments, it may include: accessing a semi-regular input grid to generate initial grid face features for each grid face, where the semi-regular input grid includes a list of faces and a plurality of vertex positions; generating a base grid including vertex positions and information indicating basic connectivity and a set of face features on the base grid through a learning-based feature aggregation module; generating a fixed-length codeword based on the base face features using a feature pooling module; accessing a predefined template grid and the base grid to generate information indicating matching vertices between the predefined template grid and the base grid; and outputting the generated fixed-length codeword and the information indicating the basic connectivity.
[0008] According to a second example method of some embodiments, it may include: accessing an input remeshed grid to generate initial grid face features, where the input remeshed grid includes a list of faces and vertex positions; generating a base grid and an atlas of face features on the base grid; generating a fixed-length codeword from the base face features; accessing a predefined number of vertices and a predefined spherical grid of base grid vertices to generate a match between the spherical grid vertices and the base network vertices; and outputting the generated fixed-length codeword and the base grid connectivity information.
[0009] According to a third example method of some embodiments, it may include: accessing an input grid, where the input grid includes a list of faces and a plurality of vertex positions; generating at least two initial grid face features for at least one face listed in the list of faces of the input grid; generating a base grid and at least two base grid face features on the base grid, where the base grid includes vertex positions and information indicating base grid connectivity; generating a fixed-length codeword from the at least two base grid face features; accessing a predefined template grid; generating information indicating matching vertices between the predefined template grid and the base grid; and outputting the fixed-length codeword and the information indicating the base grid connectivity.
[0010] For some embodiments of the third example method, the input grid is a semi-regular grid.
[0011] For some embodiments of the third example method, generating the base grid may include: generating the vertex positions; and generating the information indicating the base grid connectivity.
[0012] For some embodiments of the third example method, generating the at least two base grid face features on the base grid is performed by performing learning-based aggregation on the at least two initial grid face features.
[0013] For some embodiments of the third example method, generating the fixed - length codeword is performed by pooling the at least two base mesh face features.
[0014] For some embodiments of the third example method, the predefined template mesh is a mesh corresponding to a unit sphere.
[0015] For some embodiments of the third example method, the information indicating the base connectivity includes a triangle list having information with an indication index, and the index corresponds to a matching vertex indicated by a set of matching indices.
[0016] For some embodiments of the third example method, generating the base mesh and the at least two base mesh face features on the base mesh is performed by a learning - based heterogeneous mesh encoder, and the heterogeneous mesh encoder includes at least one down - sampling face convolutional layer.
[0017] For some embodiments of the third example method, generating the fixed - length codeword from the at least two base mesh face features includes using a learning - based AdaptMaxPool process.
[0018] For some embodiments of the third example method, the matching indices are generated by a learning - based SphereNet process.
[0019] Some embodiments of the third example method may further include: outputting information indicating the matching vertices, where the information indicating the matching vertices includes a set of matching indices, and the set of matching indices indicates the matching vertices between the predefined template mesh and the base mesh.
[0020] A first example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access a semi - regular input mesh to generate an initial mesh face feature for each mesh face, where the semi - regular input mesh includes a face list and a plurality of vertex positions; generate a base mesh including vertex positions and information indicating base connectivity and a set of face features on the base mesh by a learning - based feature aggregation module; generate a fixed - length codeword based on the base face features using a feature pooling module; access a predefined template mesh and the base mesh to generate information indicating the matching vertices between the predefined template mesh and the base mesh; and output the generated fixed - length codeword and the information indicating the base connectivity.
[0021] A second example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input remeshed grid to generate initial mesh face features, where the input remeshed grid includes a face list and vertex positions; generate a base grid and an atlas of face features on the base grid; generate a fixed-length codeword from the base face features; access a predefined number of vertices and a predefined sphere grid of base grid vertices to generate a match between the sphere grid vertices and the base network vertices; and output the generated fixed-length codeword and base grid connectivity information.
[0022] A third example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access an input grid, where the input grid includes a face list and a plurality of vertex positions; generate at least two initial mesh face features for at least one face listed on the face list of the input grid; generate a base grid and at least two base grid face features on the base grid, where the base grid includes vertex positions and information indicating base grid connectivity; generate a fixed-length codeword from the at least two base grid face features; access a predefined template grid; generate information indicating a match of vertices between the predefined template grid and the base grid; and output the fixed-length codeword and information indicating the base grid connectivity.
[0023] A fourth example method according to some embodiments may include: accessing base connectivity information and a predefined sphere grid to generate a reconstructed base grid and a base face feature map through a learning-based module DeSphereNet; and generating K reconstructed meshes at K hierarchical resolutions through a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
[0024] A fifth example method according to some embodiments may include: accessing base grid connectivity information, a fixed-length codeword, and a predefined sphere grid to generate a reconstructed base grid and a base face feature map; and generating K reconstructed meshes at K hierarchical resolutions.
[0025] A sixth example method according to some embodiments may include: receiving a fixed-length codeword, information indicating base grid connectivity, and a predefined grid to generate a reconstructed base grid and at least two base face features; generating the reconstructed base grid and at least two base face features; and generating at least one reconstructed mesh for at least two hierarchical resolutions.
[0026] For some embodiments of the sixth example method, generating the at least one reconstructed mesh generates K reconstructed meshes for K hierarchical resolutions.
[0027] For some embodiments of the sixth example method, a heterogeneous mesh decoder is used to generate K reconstructed meshes.
[0028] For some embodiments of the sixth example method, the heterogeneous mesh decoder performs at least one upsampling surface convolution process and at least one Face2Node process.
[0029] For some embodiments of the sixth example method, generating the at least one reconstructed mesh generates at least two reconstructed meshes for at least two corresponding hierarchical resolutions.
[0030] For some embodiments of the sixth example method, generating the reconstructed base mesh is performed by a learning-based DeSphereNet process.
[0031] For some embodiments of the sixth example method, generating the at least one reconstructed mesh for at least two hierarchical resolutions includes: determining input surface features from the base surface feature map; generating updated surface features corresponding to the input surface features; determining updated differential positions of one or more nodes of the reconstructed mesh; and using the corresponding updated differential positions to update the positions of one or more nodes of the reconstructed base mesh.
[0032] A fourth example device according to some embodiments may include: a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh decoder to: access base connectivity information and a predefined spherical mesh to generate a reconstructed base mesh and a base surface feature map through a learning-based module, DeSphereNet; and generate K reconstructed meshes at K hierarchical resolutions through a learning-based module, HetMeshDec, consisting of a series of K pairs of UpFaceConv and Face2Node modules.
[0033] A fifth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access base mesh connectivity information, fixed-length codewords, and a predefined spherical mesh to generate a reconstructed base mesh and a base surface feature map; and generate K reconstructed meshes at K hierarchical resolutions.
[0034] A sixth example device according to some embodiments may include: a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh decoder to: receive a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; generate the reconstructed base mesh and the at least two base surface features; and generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0035] A mesh decoder configured to obtain a fixed-length codeword, base connectivity information, and a set of sphere matching indices and generate a reconstructed mesh according to some embodiments may be configured to: access the base connectivity information and a predefined sphere mesh to generate a reconstructed base mesh and a base surface feature map via a learning-based module DeSphereNet; and generate K reconstructed meshes at K hierarchical resolutions via a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
[0036] A seventh example method according to some embodiments may include: determining initial mesh surface features from an input mesh; determining a base mesh including a set of surface features based on a first learning-based module including a series of mesh feature extraction layers; generating a fixed-length codeword from the base mesh on the mesh surface using a second learning-based pooling module; and generating a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
[0037] A seventh example device according to some embodiments may include a memory and a processor, the processor being configured to execute: determining initial mesh surface features from an input mesh; determining a base mesh including a set of surface features based on a first learning-based module including a series of mesh feature extraction layers; generating a fixed-length codeword from the base mesh on the mesh surface using a second learning-based pooling module; and generating a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
[0038] An eighth example method according to some embodiments may include: in the presence of a predefined template mesh, determining a reconstructed base mesh and a base surface feature map via a first learning-based module using a fixed codeword and a base graph; generating at least one reconstructed mesh at multiple hierarchical resolutions via a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
[0039] An eighth example device according to some embodiments may include a memory and a processor configured to perform: in the presence of a predefined template grid, determining a reconstructed base grid and a base surface feature map via a first learning-based module using a fixed codeword and a base graph; generating at least one reconstructed grid at multiple hierarchical resolutions via a second learning-based module including a series of layers, the series of layers including a series of grid feature extraction and node generation layers.
[0040] A ninth example method according to some embodiments may include: a heterogeneous grid encoder including a series of layers including paired grid feature extraction modules and grid downsampling modules; and a heterogeneous grid decoder including a learning-based module including a series of layers including paired grid node generation modules and grid upsampling modules.
[0041] For some embodiments of the ninth example device, a base grid is transmitted from the heterogeneous grid encoder to the heterogeneous grid decoder.
[0042] For some embodiments of the ninth example device, in addition to the directly consumed grid, multiple input features are used.
[0043] For some embodiments of the ninth example method, the loop subdivision-based upsampling module includes: constructing an enhanced node-specific face feature set; using a shared module to update the enhanced node-specific face feature set; averaging the updated node-specific faces; and performing neighborhood averaging on node positions.
[0044] Some embodiments of the ninth example method may further include: converting a codeword into a face-specific codeword set; and transforming the face-specific codeword into base grid features and geometry.
[0045] Some embodiments of the ninth example method may further include: converting an original grid into partitions; shifting the origin of the partitions; and encoding or decoding each partition grid separately.
[0046] For some embodiments of the ninth example method, the grids have different sizes and connectivities.
[0047] A tenth example device according to some embodiments may include a non-transitory computer-readable medium containing data content generated according to any of the methods listed above for playback using a processor.
[0048] A first example signal according to some embodiments may include: video data generated according to any of the methods listed above for playback using a processor.
[0049] An example computer program product according to some embodiments can include instructions that, when the program is executed by a computer, cause the computer to perform any of the methods listed above.
[0050] A first non-transitory computer-readable medium according to some embodiments can include data content that includes instructions for performing any of the methods listed above.
[0051] For some embodiments of the seventh example device, the third module is a learning-based module.
[0052] For some embodiments of the seventh example device, the third module is a conventional non-learning-based module.
[0053] An eleventh example method according to some embodiments can include: accessing an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; generating at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; generating a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; accessing a predefined template mesh; generating information indicating matching vertices between the predefined template mesh and the base mesh; and outputting information indicating the connectivity of the base mesh.
[0054] An eleventh example device according to some embodiments can include a processor; a memory storing instructions that, when executed by the processor, are operable to cause the processor: access an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; generate at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; access a predefined template mesh; generate information indicating matching vertices between the predefined template mesh and the base mesh; and output information indicating the connectivity of the base mesh.
[0055] A twelfth example method according to some embodiments can include: receiving information indicating the connectivity of a base mesh and a predefined mesh to generate a reconstructed base mesh and at least two base face features; generating a reconstructed base mesh and at least two base face features; and generating at least one reconstructed mesh for at least two hierarchical resolutions.
[0056] The twelfth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: receive information indicating base mesh connectivity and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; generate a reconstructed base mesh and at least two base surface features; and generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0057] The thirteenth example method according to some embodiments may include: accessing an input mesh, where the input mesh includes a face list and a plurality of vertex positions; performing a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed in the face list of the input mesh; performing an AdaptMaxPool process to: generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating base mesh connectivity; and generate a fixed-length codeword from the at least two base mesh face features; output the fixed-length codeword and the information indicating the base mesh connectivity.
[0058] The thirteenth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input mesh, where the input mesh includes a face list and a plurality of vertex positions; perform a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed in the face list of the input mesh; perform an AdaptMaxPool process to: generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating base mesh connectivity; and generate a fixed-length codeword from the at least two base mesh face features; output the fixed-length codeword and the information indicating the base mesh connectivity.
[0059] The fourteenth example method according to some embodiments may include: receiving a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; performing a Base Convolutional Graph Neural Network (BaseConGNN) process to generate a reconstructed base mesh and at least two base surface features; and performing a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0060] A fourteenth example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: receive a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; perform a Base Mesh Reconstruction Graph Neural Network (BaseConGNN) process to generate a reconstructed base mesh and at least two base surface features; and perform a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0061] A fifteenth example method according to some embodiments may include: accessing an input mesh; partitioning the input mesh into a first input mesh and a second input mesh, where the first input mesh includes a first face list and a first plurality of vertex positions, and where the second input mesh includes a second face list and a second plurality of vertex positions; generating at least two first initial mesh surface features for at least one first face listed on the first face list of the first input mesh; generating a first base mesh and at least two first base mesh surface features on the first base mesh, where the first base mesh includes first vertex positions and first information indicating first base mesh connectivity; generating a first fixed-length codeword from the at least two first base mesh surface features; accessing a first predefined template mesh; outputting the first fixed-length codeword and the first information indicating the first base mesh connectivity; generating at least two second initial mesh surface features for at least one second face listed on the second face list of the second input mesh; generating a second base mesh and at least two second base mesh surface features on the second base mesh, where the second base mesh includes second vertex positions and second information indicating first base mesh connectivity; generating a second fixed-length codeword from the at least two second base mesh surface features; accessing a second predefined template mesh; and outputting the second fixed-length codeword and the second information indicating the second base mesh connectivity.
[0062] Some embodiments of the fifteenth example method may further include: generating a first matching index set, where the first matching index set indicates first matching vertices between the first predefined template mesh and the first base mesh; outputting the first matching index set; generating a second matching index set, where the second matching index set indicates second matching vertices between the second predefined template mesh and the second base mesh; and outputting the second matching index set.
[0063] The fifteenth example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input mesh; partition the input mesh into a first input mesh and a second input mesh, wherein the first input mesh includes a first list of faces and a first plurality of vertex positions, and wherein the second input mesh includes a second list of faces and a second plurality of vertex positions; for at least one first face listed on the first list of faces of the first input mesh, generate at least two first initial mesh face features; generate a first base mesh and at least two first base mesh face features on the first base mesh, wherein the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; generate a first fixed-length codeword from the at least two first base mesh face features; access a first predefined template mesh; output the first fixed-length codeword and the first information indicating the connectivity of the first base mesh; for at least one second face listed on the second list of faces of the second input mesh, generate at least two second initial mesh face features; generate a second base mesh and at least two second base mesh face features on the second base mesh, wherein the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; generate a second fixed-length codeword from the at least two second base mesh face features; access a second predefined template mesh; and output the second fixed-length codeword and the second information indicating the connectivity of the second base mesh.
[0064] The sixteenth example device according to some embodiments may include: at least one processor configured to perform any of the methods listed above.
[0065] The seventeenth example device according to some embodiments may include a computer-readable storage medium storing instructions for causing one or more processors to perform any of the methods listed above.
[0066] The eighteenth example device according to some embodiments may include: at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.
[0067] The second example signal according to some embodiments may include: a bitstream generated according to any of the methods listed above.
[0068] In additional embodiments, an encoder and a decoder device are provided to perform the methods described herein. The encoder or decoder device may include a processor configured to perform the methods described herein. The device may include a computer-readable medium (e.g., a non-transitory medium) that stores instructions for performing the methods described herein. In some embodiments, the computer-readable medium (e.g., a non-transitory medium) stores video encoded using any of the methods described herein.
[0069] One or more of the present embodiments also provide a computer-readable storage medium having instructions stored thereon for performing bidirectional optical flow, encoding or decoding video data according to any of the methods described above. The present embodiment also provides a computer-readable storage medium having a bitstream stored thereon generated according to the methods described above. The present embodiment also provides a method and a device for transmitting a bitstream generated according to the methods described above. The present embodiment also provides a computer program product that includes instructions for performing any of the described methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1A is a system diagram showing an example communication system according to some embodiments.
[0071] Figure 1B is a system diagram showing an example wireless transmit / receive unit (WTRU) that may be used within the Figure 1A illustrated communication system according to some embodiments.
[0072] Figure 1C is a system diagram showing an example set of interfaces according to some embodiments.
[0073] Figure 2A is a functional block diagram of a block-based video encoder (such as a video compression encoder) according to some embodiments.
[0074] Figure 2B is a functional block diagram of a block-based video decoder (such as a video decompression decoder) according to some embodiments.
[0075] Figure 3A is a schematic diagram showing an example FoldingNet encoder-decoder architecture.
[0076] Figure 3B is a schematic diagram showing an example encoder-decoder architecture according to some embodiments.
[0077] Figure 4Ais a functional block diagram showing an example fixed-length codeword autoencoder with soft disentanglement according to some embodiments.
[0078] Figure 4B is a schematic diagram showing an example fixed-length codeword autoencoder with soft disentanglement according to some embodiments.
[0079] Figure 5A is a functional block diagram showing an example feature map autoencoder process according to some embodiments.
[0080] Figure 5B is a schematic diagram showing an example feature map autoencoder process according to some embodiments.
[0081] Figure 6A is a functional block diagram showing an example heterogeneous grid encoder process according to some embodiments.
[0082] Figure 6B is a schematic diagram showing an example heterogeneous grid encoder process according to some embodiments.
[0083] Figure 7A is a functional block diagram showing an example heterogeneous grid decoder process according to some embodiments.
[0084] Figure 7B is a schematic diagram showing an example heterogeneous grid decoder process according to some embodiments.
[0085] Figure 8A is a functional block diagram showing an example face convolutional downsampling process according to some embodiments.
[0086] Figure 8B is a schematic diagram showing an example face convolutional downsampling process according to some embodiments.
[0087] Figure 9A is a functional block diagram showing an example upsampling face convolutional process according to some embodiments.
[0088] Figure 9B is a schematic diagram showing an example upsampling face convolutional process according to some embodiments.
[0089] Figure 10 is a schematic diagram showing an example aggregation of adjacent faces around a node according to some embodiments.
[0090] Figure 11A is a functional block diagram showing an example process for converting face features into differential position updates according to some embodiments.
[0091] Figure 11Bis a schematic diagram showing an example process for converting surface features into differential position updates according to some embodiments.
[0092] Figure 12A is a functional block diagram showing an example process for deforming a base mesh into a canonical sphere shape according to some embodiments.
[0093] Figure 12B is a schematic diagram showing an example index matching process according to some embodiments.
[0094] Figure 13 is a functional block diagram showing an example fixed-length codeword autoencoder with hard decoupling according to some embodiments.
[0095] Figure 14 is a functional block diagram showing an example residual surface convolution process according to some embodiments.
[0096] Figure 15 is a functional block diagram showing an example inception-residual surface convolution according to some embodiments.
[0097] Figure 16 is a functional block diagram showing an example partition-based encoding process according to some embodiments.
[0098] Figure 17 is a functional block diagram showing an example partition-based decoding process according to some embodiments.
[0099] Figure 18 is a functional block diagram showing an example mesh classification architecture based on a fixed-length codeword autoencoder according to some embodiments.
[0100] Figure 19 is a functional block diagram showing an example fixed-length codeword autoencoder with soft decoupling and SphereNet according to some embodiments.
[0101] Figure 20 is a flowchart showing an example encoding method according to some embodiments.
[0102] Figure 21 is a flowchart showing an example decoding method according to some embodiments.
[0103] Figure 22 is a flowchart showing an example encoding process according to some embodiments.
[0104] Figure 23 is a flowchart showing an example decoding process according to some embodiments.
[0105] The entities, connections, arrangements, etc. depicted and described in connection with the various figures are presented by way of example and not limitation. Accordingly, any and all statements or other indications as to what a particular figure “depicts,” what a particular element or entity “is” or “has” in a particular figure, and any and all similar statements that might be construed in isolation and out of context as absolute and thus restrictive can only be properly construed as constructively beginning with a clause such as “in at least one embodiment, ….” For the sake of simplicity and clarity of presentation, this implicit leading clause will not be repeated ad nauseam in the detailed description. Detailed Description
[0106] Figure 1A is a system diagram showing an example communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multi-access system that provides content such as voice, data, video, messaging, broadcast, etc. to a plurality of wireless users. The communication system 100 may enable the plurality of wireless users to access such content through shared system resources including wireless broadband. For example, the communication system 100 may employ one or more channel access methods such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single carrier FDMA (SC-FDMA), zero-tail unique word DFT spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block filtered OFDM, filter bank multicarrier (FBMC), etc.
[0107] As Figure 1AAs shown, the communication system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a radio access network (RAN) 104 / 113, a core network (CN) 106, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, but it will be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, the WTRUs 102a, 102b, 102c, 102d (any one of which may be referred to as a "station" and / or "STA") may be configured to transmit and / or receive wireless signals and may include a user equipment (UE), a mobile station, a fixed or mobile subscriber unit, a subscription-based unit, a pager, a cellular phone, a personal digital assistant (PDA), a smartphone, a laptop computer, a netbook, a personal computer, a wireless sensor, a hotspot or Mi-Fi device, an Internet of Things (IoT) device, a watch or other wearable device, a head-mounted display (HMD), a vehicle, a drone, a medical device and application (e.g., remote surgery), an industrial device and application (e.g., a robot and / or other wireless devices operating in the context of an industrial and / or automation processing chain), a consumer electronic device, a device operating on a commercial and / or industrial wireless network, etc. Any one of the WTRUs 102a, 102b, 102c, 102d may be interchangeably referred to as a UE.
[0108] The communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks such as the CN 106, the Internet 110, and / or other networks 112. For example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node B, an eNode B, a master Node B, a master eNode B, a gNB, an NR NodeB, a site controller, an access point (AP), a wireless router, etc. Although the base stations 114a, 114b are each depicted as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.
[0109] Base station 114a may be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as cells (not shown). These frequencies may be in licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide wireless service coverage to a specific geographical area that may be relatively fixed or may change over time. A cell may also be divided into cell sectors. For example, the cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In an embodiment, base station 114a may employ multiple-input multiple-output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.
[0110] Base stations 114a, 114b may communicate with one or more of WTRUs 102a, 102b, 102c, 102d via air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, millimeter wave, infrared (IR), ultraviolet (UV), visible light, etc.). Any suitable radio access technology (RAT) may be used to establish air interface 116.
[0111] More specifically, as described above, communication system 100 may be a multiple access system and may employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base station 114a in RAN 104 / 113 and WTRUs 102a, 102b, 102c may implement radio technologies, such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA) that may use Wideband CDMA (WCDMA) to establish air interface 116. WCDMA may include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA may include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0112] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies, such as the evolved UMTS terrestrial radio access (E-UTRA) that may use Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-A Pro (LTE-A Pro) to establish the air interface 116.
[0113] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies, such as the New Radio (NR) air interface access that may use the New Radio (NR) to establish the air interface 116.
[0114] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple air interface access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may jointly implement LTE radio access and NR radio access using, for example, the dual connectivity (DC) principle. Thus, the air interface used by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).
[0115] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c may implement radio technologies, such as IEEE 802.11 (i.e., Wi-Fi), IEEE 802.16 (i.e., WiMAX), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rate for GSM Evolution (EDGE), GSM EDGE (GERAN), etc.
[0116] Figure 1AThe base station 114b therein can be, for example, a wireless router, a master Node B, a master eNode-B, or an access point, and can utilize any suitable RAT to facilitate wireless connectivity in a local area such as a commercial premise, a home, a vehicle, a campus, an industrial facility, an aviation corridor (e.g., for drones), a road, etc. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a pico cell or a femto cell. As Figure 1A shown, the base station 114b can be directly connected to the Internet 110. Thus, the base station 114b may not need to access the Internet 110 via the CN 106.
[0117] The RAN 104 / 113 can communicate with the CN 106, which can be any type of network configured to provide voice, data, applications, and / or voice over Internet Protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have different quality of service (QoS) requirements, such as different throughput requirements, latency requirements, fault tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 can provide call control, billing services, location-based services, prepaid calls, Internet connectivity, video distribution, etc., and / or perform advanced security functions such as user authentication. Although Figure 1A not shown in the figure, it will be understood that the RAN 104 / 113 and / or the CN 106 can communicate directly or indirectly with other RANs that employ the same RAT or a different RAT as the RAN 104 / 113. For example, in addition to being connected to the RAN 104 / 113 that may be utilizing the NR radio technology, the CN 106 can also communicate with another RAN (not shown) that employs a GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.
[0118] CN 106 can also serve as a gateway for the WTRU 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 can include a circuit-switched telephone network that provides plain old telephone service (POTS). The Internet 110 can include a global system of interconnected computer networks and devices that use common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and / or Internet Protocol (IP) in the TCP / IP Internet protocol suite. The network 112 can include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 can include another CN connected to one or more RANs, and the other CN can employ the same RAT or a different RAT as the RAN 104 / 113.
[0119] Some or all of the WTRU 102a, 102b, 102c, 102d in the communication system 100 can include multi-mode capabilities (e.g., the WTRU 102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links). For example, Figure 1A the illustrated WTRU 102c can be configured to communicate with a base station 114a that can employ a cellular-based radio technology and with a base station 114b that can employ IEEE 802 radio technology.
[0120] Figure 1B is a system diagram showing an exemplary WTRU 102. As Figure 1B shown, the WTRU 102 can include a processor 118, a transceiver 120, transmit / receive elements 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and / or other peripheral devices 138, etc. It will be appreciated that the WTRU 102 can include any sub-combination of the foregoing elements while remaining consistent with the embodiments.
[0121] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors associated with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 118 can perform signal encoding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to a transceiver 120, which can be coupled to a transmit / receive element 122. Although Figure 1B the processor 118 and the transceiver 120 are described as separate components, it should be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0122] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) via an air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be a transmitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In another embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF signals and optical signals. It will be appreciated that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0123] Although the transmit / receive element 122 is depicted as a single element in Figure 1B the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals via the air interface 116.
[0124] The transceiver 120 can be configured to modulate the signals to be transmitted by the transmit / receive element 122 and demodulate the signals received by the transmit / receive element 122. As described above, the WTRU 102 can have multi-mode capabilities. Thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs (e.g., such as NR and IEEE 802.11).
[0125] The processor 118 of the WTRU 102 may be coupled to and may receive user input data from: a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 118 may also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information from and store data in any type of suitable memory such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random access memory (RAM), read only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 118 may access information from and store data in a memory that is not actually located on the WTRU 102, such as on a server or a home computer (not shown).
[0126] The processor 118 may receive power from a power source 134 and may be configured to distribute power to and / or control the power of other components in the WTRU 102. The power source 134 may be any suitable device for powering the WTRU 102. For example, the power source 134 may include one or more dry cells (e.g., nickel cadmium (NiCd), nickel zinc (NiZn), nickel metal hydride (NiMH), lithium ion (Li-ion), etc.), a solar cell, a fuel cell, etc.
[0127] The processor 118 may also be coupled to a GPS chipset 136, which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) via an air interface 116 and / or may determine its location based on the timing of signals received from two or more nearby base stations. It will be appreciated that the WTRU 102 may obtain location information by any suitable location determination method while remaining consistent with the embodiments.
[0128] The processor 118 may also be coupled to other peripheral devices 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripheral devices 138 may include an accelerometer, an electronic compass, a satellite transceiver, a digital camera (for photos and / or videos), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands-free headset, Modules, FM radio units, digital music players, media players, video game player modules, Internet browsers, virtual reality and / or augmented reality (VR / AR) devices, activity trackers, etc. Peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, an orientation sensor, a proximity sensor, a temperature sensor, a time sensor; a geographical location sensor; an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0129] The WTRU 102 may include a full-duplex radio, where some or all of the transmission and reception of signals (e.g., associated with a particular subframe for both UL (e.g., for transmission) and downlink (e.g., for reception)) may be concurrent and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference via hardware (e.g., chokes) or signal processing via a processor (e.g., a separate processor (not shown) or via processor 118). In an embodiment, the WTRU 102 may include a half-duplex radio, where some or all of the transmission and reception of signals (e.g., associated with a particular subframe for either UL (e.g., for transmission) or downlink (e.g., for reception)).
[0130] Although the WTRU is depicted as a wireless terminal in Figures 1A to 1B it is envisioned that in certain representative embodiments, such a terminal may (e.g., temporarily or permanently) use a wired communication interface to a communication network.
[0131] In a representative embodiment, other network 112 may be a WLAN.
[0132] In view of Figures 1A to 1B and the corresponding description, one or more or all of the functions described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or simulate network and / or WTRU functions.
[0133] The simulation device can be designed to implement one or more tests on other devices in a laboratory environment and / or an operator network environment. For example, one or more simulation devices can perform one or more or all functions when fully or partially implemented and / or deployed as part of a wired and / or wireless communication network, so as to test other devices within the communication network. One or more simulation devices can perform one or more or all functions when temporarily implemented / deployed as part of a wired and / or wireless communication network. The simulation device can be directly coupled to another device for testing purposes and / or can perform tests using over-the-air wireless communication.
[0134] One or more simulation devices can perform one or more (including all) functions when not implemented / deployed as part of a wired and / or wireless communication network. For example, the simulation device can be used to test test scenarios in a laboratory and / or an undeployed (e.g., for testing) wired and / or wireless communication network, so as to implement testing of one or more components. One or more simulation devices can be test equipment. The simulation device can transmit and / or receive data using direct RF coupling and / or wireless communication via an RF circuit system (e.g., which can include one or more antennas).
[0135] Figure 1C is a system diagram showing a set of example interfaces according to some embodiments. For some embodiments, an extended reality display device and its control electronics can be implemented. System 150 can be embodied as a device including various components described below and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of System 150 can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components individually or in combination. For example, in at least one embodiment, the processing element and encoder / decoder element of System 150 are distributed over multiple ICs and / or discrete components. In various embodiments, System 150 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, System 150 is configured to implement one or more aspects described in this document.
[0136] System 150 includes at least one processor 152 configured to execute instructions loaded therein to implement various aspects described, for example, in this document. The processor 152 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 150 includes at least one memory 154 (e.g., volatile memory devices and / or non-volatile memory devices). System 150 may include a storage device 158, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 158 may include internal storage devices, attached storage devices (including detachable and non-detachable storage devices), and / or network-accessible storage devices.
[0137] System 150 includes an encoder / decoder module 156 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 156 may include its own processor and memory. The encoder / decoder module 156 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 156 may be implemented as a separate element of System 150 or may be incorporated within the processor 152 as a combination of hardware and software known to those skilled in the art.
[0138] The program code to be loaded onto the processor 152 or the encoder / decoder 156 to execute the various aspects described in this document may be stored in the storage device 158 and subsequently loaded onto the memory 154 for execution by the processor 152. According to various embodiments, one or more of the processor 152, the memory 154, the storage device 158, and the encoder / decoder module 156 may store one or more of various items during the execution of the processes described in this document. Such stored items may include but are not limited to input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.
[0139] In some embodiments, the memory internal to the processor 152 and / or the encoder / decoder module 156 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 152 or the encoder / decoder module 152) is used for one or more of these functions. The external memory can be the memory 154 and / or the storage device 158, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, i.e., the new standard developed by JVET (Joint Video Exploration Team)).
[0140] Input to the elements of the system 150 can be provided by a variety of input devices, as indicated in block 172. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives, for example, an RF signal transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a set of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1C Other examples not shown include composite video.
[0141] In various embodiments, the input device of block 172 has corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also known as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and filtering again to a desired band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0142] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 150 to other electronic devices via a USB and / or HDMI connection. It should be understood that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, within a separate input processing IC or within processor 152 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 152 as needed. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 152 and an encoder / decoder 156 operating in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0143] The various elements of system 150 may be provided in an integrated housing. In the integrated housing, the various elements may be interconnected using suitable connection means 174 (e.g., internal buses known in the art, including inter-integrated circuit (I2C) buses, wiring, and printed circuit boards) and data may be transmitted between them.
[0144] System 150 includes a communication interface 160 that enables communication with other devices via a communication channel 162. The communication interface 160 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 162. The communication interface 160 may include, but is not limited to, a modem or a network card, and the communication channel 162 may be implemented, for example, within a wired and / or wireless medium.
[0145] In various embodiments, data is streamed to or otherwise provided to the system 150 using a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via a communication channel 162 and a communication interface 160 suitable for Wi-Fi communication. The communication channel 162 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other top communication. Other embodiments use a set-top box to provide streaming data to the system 150, and the set-top box delivers data via an HDMI connection of the input block 172. Still other embodiments use an RF connection of the input block 172 to provide streaming data to the system 150. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use a wireless network other than Wi-Fi, e.g., a cellular network or a Bluetooth network.
[0146] The system 150 may provide output signals to various output devices including a display 176, a speaker 178, and other peripheral devices 180. The display 176 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 176 may be used for a television, a tablet, a laptop computer, a mobile phone (cellular phone), or other devices. The display 176 may also be integrated with other components (e.g., as in a smartphone) or be separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, the other peripheral devices 180 include one or more of a standalone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more of the peripheral devices 180 that perform functions based on the output of the system 150. For example, a disc player performs the function of playing the output of the system 150.
[0147] In various embodiments, control signals are transmitted between system 150 and display 176, speaker 178, or other peripheral devices 180 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 150 via dedicated connections through respective interfaces 164, 166, and 168. Alternatively, the output devices may be connected to system 150 via communication interface 160 using communication channel 162. Display 176 and speaker 178 may be integrated in a single unit with other components of system 150 in an electronic device (e.g., such as a television). In various embodiments, display interface 164 includes a display driver, e.g., such as a timing controller (T Con) chip.
[0148] For example, if RF input section 172 is part of a separate set-top box, display 176 and speaker 178 may alternatively be separate from one or more of the other components. In various embodiments where display 176 and speaker 178 are external components, output signals may be provided via dedicated output connections including, for example, HDMI ports, USB ports, or COMP outputs.
[0149] System 150 may include one or more sensor devices 168. Examples of sensor devices that may be used include one or more GPS sensors, gyroscopic sensors, accelerometers, light sensors, cameras, depth cameras, microphones, and / or magnetometers. Such sensors may be used to determine information such as the location and orientation of a user. In cases where system 150 serves as a control module (such as control modules 124, 132) for an extended reality display, the location and orientation of the user may be used to determine how to render image data so that the user perceives the correct portion of a virtual object or virtual scene from the correct perspective. In the case of a head-mounted display device, the location and orientation of the device itself may be used to determine the location and orientation of the user for rendering virtual content. In the case of other display devices (such as a phone, tablet, computer monitor, or television), other inputs may be used to determine the location and orientation of the user for rendering content. For example, a user may select and / or adjust a desired perspective and / or viewing direction by using a touch screen, keypad or keyboard, trackball, joystick, or other input. In cases where the display device has sensors such as accelerometers and / or gyroscopes, the perspective and orientation may be selected and / or adjusted based on the movement of the display device for rendering content.
[0150] The embodiments may be performed by computer software implemented by the processor 152 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, the memory 154 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, the processor 152 may be of any type suitable for the technical environment and may cover one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.
[0151] Embodiments described herein include various aspects, including tools, features, embodiments, models, methods, etc. Many aspects in these aspects are described as having specificity, and at least in order to show each feature, are usually described in a manner that may sound restrictive. However, this is for the purpose of describing clarity, and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide other aspects. In addition, these aspects can also be combined and interchanged with the aspects described in previous files.
[0152] The aspects described and contemplated in this application can be implemented in many different forms. Figure 1C , Figure 2A and Figure 2B Some embodiments are provided, but other embodiments are contemplated and are not intended to be construed as Figure 1C , Figure 2A and Figure 2B The discussion does not limit the breadth of implementations. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, devices, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or computer-readable storage media having stored thereon bitstreams generated according to any of the methods described.
[0153] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture" and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" or "reconstruction" is used on the decoder side.
[0154] This document describes various methods, and each of these methods includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined. Additionally, in various embodiments, terms such as "first", "second", etc. can be used to modify elements, components, steps, operations, etc., such as, for example, "first decoding" and "second decoding". Unless specifically required, the use of such terms does not imply an ordering of the modified operations. Thus, in this example, the first decoding does not need to be performed before the second decoding, but can occur, for example, before, during, or in a time period overlapping with the second decoding.
[0155] The various methods and other aspects described in this application can be used to modify the blocks of the video encoder 200 and decoder 250 as shown in Figure 2A and Figure 2B e.g., intra prediction 220, 262, entropy encoding 212, and / or entropy decoding 252. Additionally, this aspect is not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations (whether pre-existing or future developments), as well as extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically precluded, these aspects described in this application can be used alone or in combination.
[0156] Various numerical values are used in this application. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0157] Figure 2A is a functional block diagram of a block-based video encoder (such as a video compression encoder) according to some embodiments. Figure 2A The encoder 200 is shown. Variations of such an encoder 200 are contemplated, but for clarity, the encoder 200 is described below without describing all the expected variations. Before being encoded, a video sequence can undergo pre-encoding processing 202, e.g., applying a color transformation to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the preprocessing and appended to the bitstream.
[0158] In encoder 200, pictures are encoded by encoder elements as described below. The picture to be encoded is partitioned 204 and processed in units such as coding units (CUs). Each unit is encoded using, for example, an intra or inter mode. When a unit is encoded in the intra mode, the encoder performs intra prediction 220. In the inter mode, motion estimation 226 and compensation 228 are performed. The encoder decides 230 which of the intra or inter modes to use to encode the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting 206 the predicted block from the original image block.
[0159] The prediction residual is then transformed 208 and quantized 210. The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded 212 to output a bitstream. The encoder may skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass both the transformation and quantization, where the residual is directly encoded without applying the transformation or quantization process.
[0160] The encoder decodes the encoded blocks for use as a reference for further prediction. The quantized transform coefficients are dequantized 214 and inverse transformed (216) to decode the prediction residual. The decoded prediction residual and the predicted block are combined 218 to reconstruct the image block. An in-loop filter 222 is applied to the reconstructed picture to perform, for example, deblocking / sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer 224.
[0161] Figure 2B is a functional block diagram of a block-based video decoder (such as a video decompression decoder) according to some embodiments. Figure 2B A block diagram of video decoder 250 is shown. In decoder 250, the bitstream is decoded by decoder elements as described below. Video decoder 250 generally performs a decoding process that is the reverse of the encoding process as described in Figure 2A above. Encoder 200 generally also performs video decoding as part of encoding video data.
[0162] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 200. First, entropy decoding 252 is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. The picture partitioning information indicates how to partition the picture. Thus, the decoder can partition 254 the picture according to the decoded picture partitioning information. The transform coefficients are dequantized 256 and inverse-transformed 258 to decode the prediction residuals. The decoded prediction residuals and the prediction blocks are combined 260 to reconstruct the image block. The prediction block can be obtained 272 from intra prediction 262 or motion-compensated prediction (inter prediction) 270. The in-loop filter 264 is applied to the reconstructed image. The filtered image is stored in the reference picture buffer 268.
[0163] The decoded picture may also undergo post-decoding processing 266, such as an inverse color transformation (e.g., conversion from YcbCr 4:2:0 to RGB 4:4:4) or inverse remapping, which performs the inverse process of the remapping process performed in the precoding process 202. The post-decoding processing can use the metadata derived in the precoding process and signaled in the bitstream.
[0164] According to some embodiments, the present application discloses mesh and point cloud processing, including the analysis, interpolated representation, understanding, and processing of mesh and point cloud signals.
[0165] Point cloud data may consume most of the network traffic, e.g., between cars connected via a 5G network and in immersive (such as AR / VR / MR) communications. An efficient representation format can be used for point clouds and communications. In particular, the raw point cloud data can be organized and processed to model and sense, for example, the world, environment, or scene. Compression of the raw point cloud can be used for data storage and transmission.
[0166] Furthermore, a point cloud can represent successive scans of the same scene that may contain multiple moving objects. Dynamic point clouds capture moving objects, while static point clouds capture static scenes and / or static objects. Dynamic point clouds can typically be organized into frames, where different frames are captured at different times. The processing and compression of dynamic point clouds can be performed in real time or with a low latency amount.
[0167] The automotive industry and autonomous vehicles are some of the fields where point clouds can be used. Autonomous cars "detect" and sense their environment to make good driving decisions based on the reality of their surroundings. Sensors such as lidar generate (dynamic) point clouds for the perception engine. These point clouds are generally not intended to be seen by the human eye, and these point clouds may or may not have color and are generally sparse and dynamic, with a high capture frequency. Such point clouds may have other attributes, such as reflectivity provided by lidar, as this attribute indicates the material of the sensed object and may help in making decisions.
[0168] Virtual reality (VR) and immersive worlds have become a hot topic and are seen by many as the future of 2D flat video. Instead of a standard TV where viewers only see the virtual world in front of them, viewers can be immersed in an omnidirectional environment. Depending on the degree of freedom of the viewer in the environment, there are several levels of immersion. Point cloud formats can be used to distribute VR world and environmental data. Such point clouds can be static or dynamic and are typically of average size, such as less than millions of points at a time.
[0169] Point clouds can also be used for a variety of other purposes, such as scanning cultural heritage objects and / or buildings, where an object such as a statue or building is 3D scanned. The spatial configuration data of the object can be shared without sending or accessing the actual object or building. Additionally, this data can be used to preserve knowledge of the object in case the object or building is damaged (such as a temple damaged by an earthquake). Such point clouds are typically static, colored, and of large size.
[0170] Another use case is terrain and cartography using 3D representations, where the map is not limited to a flat surface and can include bumps and depressions. For example, some map websites and applications use meshes instead of point clouds for their 3D maps. However, point clouds can be a data format suitable for 3D maps, and such point clouds are also typically static, colored, and of large size.
[0171] World modeling and sensing via point clouds can allow machines to record and use spatial configuration data about the 3D world around them, which can be used in the applications discussed above.
[0172] 3D point cloud data consists of discrete samples of the surface of an object or scene. To adequately represent the real world with point samples, a large number of points can be used. For example, a typical VR immersive scene includes millions of points, while point clouds can typically include hundreds of millions of points. Therefore, processing such large-scale point clouds is computationally expensive, especially for consumer devices that may have limited computing power, such as smartphones, tablets, and car navigation systems.
[0173] Additionally, the discrete samples comprising the 3D point cloud data may still contain incomplete information about the underlying surface of the object and scene. Therefore, there have also been recent efforts to explore mesh representations of 3D scene / surface representations. A mesh can be considered a 3D point cloud together with connectivity information between points. Thus, the mesh representation bridges the gap between the point cloud and the underlying continuous surface by approximating local 2D polygon patches (called faces) of the underlying surface.
[0174] The first step in performing any kind of processing or inference on mesh data is to adopt an efficient storage method. To store and process the input point cloud at an affordable computational cost, the input point cloud can be downsampled, where the downsampled point cloud summarizes the geometry of the input point cloud while having fewer (but larger) faces. The downsampled point cloud is fed into subsequent machine tasks for further processing. However, the storage space can be further reduced by converting the original mesh data (original or downsampled) into a fixed-length codeword or a feature map on an extremely low-resolution mesh. This codeword or feature map can be converted into a bitstream through entropy coding techniques. Additionally, the codeword or feature map can be used to represent the global or local surface information of the underlying scene / object respectively, and can be paired with subsequent downstream (machine vision) blocks.
[0175] Raw data from a sensing modality can produce a mesh representation that includes hundreds of thousands of faces to be efficiently stored. Compared to point clouds, meshes provide more information about the underlying 3D shape of the mesh representation. Meshes provide this additional information through connectivity information. This connectivity information poses challenges in designing efficient learning-based architectures for mesh processing and compression. According to some embodiments, the present application describes a mesh autoencoder framework for generating and "learning" a representation of a heterogeneous 3D triangular mesh in parallel with convolutional-based autoencoders in 2D vision.
[0176] In recent years, various attempts have been made to design autoencoders on meshes, such as the Litany, the article by Or et al., Deformable Shape Completion with Graph Convolutional Autoencoders, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018) (the "Litany"), and the article by Ranjan, Anurag et al., Generating 3D Faces Using Convolutional Mesh Autoencoders, European Conference on Computer Vision (ECCV) 704-720 (2018) (the "CoMA"). Litany is understood to treat the mesh purely as a graph and apply variational graph autoencoders using mesh geometry as input features. See Variational Graph Auto-Encoders by Kipf, Thomas, and Max Welling, arXiv preprint arXiv:1611.07308 (2016). This method has no hierarchical pooling and does not apply any mesh-specific operations. Ranjan defined fixed upsampling and downsampling operations in a hierarchical manner based on quadratic error simplification combined with spectral convolutional layers, which are understood to require operations on meshes of the same size and connectivity. This is because the pooling and up-pooling operations are predefined and depend on the connectivity represented as an adjacency matrix.
[0177] The articles by Bouritsas, Giorgos et al., "Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and Generation" (Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and Generation), Proceedings of the IEEE / CVF International Conference on Computer Vision 7213-7222 (2019); the article by Yuan, Yu-Jie et al., "Mesh Variational Autoencoders with Edge Contraction Pooling" (Mesh Variational Autoencoders with Edge Contraction Pooling), Proceedings of the IEEE / CVF Workshop on Computer Vision and Pattern Recognition 274-275 (2020); the article by Zhou, Yi et al., "Fully Convolutional Mesh Autoencoder Using Efficient Spatially Varying Kernels" (Fully Convolutional Mesh Autoencoder Using Efficient Spatially Varying Kernels), 33 Advances in Neural Information Processing Systems 9251-9262 (2020) improved the convolutional layer, but are still limited by fixed-size and connectivity constraints.
[0178] The article by Hanocka, Rana et al., "MeshCNN: A Network with an Edge" (MeshCNN: A Network with an Edge), ACM Transactions on Graphics (TOG) 38:4 1-12 (2019) defined learnable upsampling and downsampling modules suitable for different meshes of different sizes. These layers should be understood as not being proven to build good autoencoders, but rather for mesh classification and segmentation.
[0179] The article "Subdivision-Based Mesh Convolution Networks" by Hu, Shi-Min et al., 41:3 ACM Transactions on Graphics (TOG) 1-16 (2022) ("Hu") studied subdivision-based network processing, in which the original mesh is transformed into a new mesh that well approximates the original mesh but exhibits subdivision connectivity (semi-regular mesh), called the remeshed mesh. Broadly speaking, as an example, the semi-regular mesh is filled with hierarchical structures, where each face has three adjacent faces (corresponding to its three edges), and one face and its three neighbors can be combined to form a single face. This property makes the semi-regular mesh easy to perform fixed upsampling and downsampling operations, which are the "cornerstones" of convolution-based architectures. In addition, Hu defined a learning-based module on the subdivision mesh, which is a general framework for defining face-based convolutional layers (treating faces as pixels in an image). The article "Neural Subdivision" by Liu, Hsueh-Ti Derek et al., arXiv IV arXiv preprint arXiv IV :2005.01819 (2020) achieved a coarse-to-fine mesh super-resolution function.
[0180] For autoencoders, "Mesh Convolutional Autoencoder for Semi-Regular Meshes of Different Sizes" by Hahner, Sara and Jochen Garcke, Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision 885-894 (2022) ("Garcke") attempted to use the autoencoding capabilities of subdivision meshes and meshes of different sizes to implement an autoencoder. However, the method they described in Garcke is understood to be unable to generate a fixed-length latent representation from meshes of different sizes, and can only generate a latent feature map on the base mesh, and the latent feature map can only be compared between meshes with the same base mesh connectivity and face ordering. This detail precludes meaningful latent space comparisons between heterogeneous meshes where the size, connectivity, or ordering may be different. According to some embodiments, a fixed-length latent representation may be preferred because a fixed-length latent representation enables subsequent analysis / understanding of the input mesh geometry.
[0181] According to some embodiments, the present application discloses heterogeneous semi-regular meshes and how, for example, an efficient fixed-length codeword or an autoencoder based on feature map generation learning can be used for these heterogeneous meshes.
[0182] In an image autoencoder system, the encoder and decoder typically alternate between convolutional and upsampling / downsampling operations. Due to the fixed grid support of images, these downsampling and upsampling layers can be set with a fixed ratio (e.g., 2x pooling). Additionally, since images can be resized to the same size via interpolation techniques, hard-coded processing layer sizes that map the image to a fixed-size latent representation and back to the original image size can be used. In contrast, the size of triangular mesh data is variable and has highly irregular support, which includes geometry (a list of points) and connectivity (a list of triangles with indices corresponding to the points). This triangular mesh data construction may prevent the use of convolutional neighborhood structures, the use of upsampling and downsampling structures, and the extraction of a fixed-length latent representation from a variable-size grid. While other mesh autoencoders may have attempted to address some of these issues, it should be understood that no other autoencoder method can handle heterogeneous meshes and extract meaningful fixed-length latent representations that generalize across meshes of different sizes and connectivities in a manner similar to image autoencoders.
[0183] It can be compared with autoencoders for point cloud data, as point clouds typically have an irregular structure and variable size. Although meshes already contain connectivity information that carries more topological information about the underlying surface compared to point clouds, the connectivity information can pose additional challenges. The articles "Foldingnet: Point Cloud Auto-Encoder via Deep Grid Deformation" by Yang, Yaoqing, et al., Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2018) and "Tearingnet: Point Cloud Autoencoder to Learn Topology-Friendly Representations" by Pang, Jiahao, Duanshun Li, and Dong Tian, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (2021) discuss point cloud autoencoders that are able to extract fixed-length latent representations on point clouds of different sizes and reconstruct the point cloud in a certain canonical ordering that may be different from the original ordering of the input point cloud. This detail prevents the reconstruction of meshes with the original connectivity, as the connectivity may no longer align with the output point ordering. Additionally, there is the issue of how to integrate the connectivity information into such learning pipelines.
[0184] Figure 3AFIG. is a schematic diagram showing an example FoldingNet encoder-decoder architecture. The input mesh 302 is input into the encoder 304, and a codeword c 306 is generated at the output. The decoder 310 takes the codeword c 306 and the surface 308 as inputs and reconstructs the mesh 312.
[0185] Figure 3B FIG. is a schematic diagram showing an example encoder-decoder architecture according to some embodiments. For some embodiments, the heterogeneous encoder-decoder architecture can be Figure 3B the example HetMeshNet encoder-decoder architecture shown in FIG. The heterogeneous mesh encoder 354 receives the input mesh 352, which can include a list of features and a list of faces (or triangles for some embodiments), and encodes the input mesh to generate the codewords c 356 of triangles and vertices and the output mesh 358. The decoder reverses the process. The decoder 360 also takes as input a uniform sphere 362 with uniformly spaced vertices to reconstruct the mesh 364.
[0186] According to some embodiments, the present application discusses an end-to-end learning-based mesh autoencoder framework that can operate on meshes of different sizes and handle connectivity while producing a fixed-length latent representation, thus mimicking those frameworks in the image domain. In some embodiments, unsupervised transfer classification can be performed on heterogeneous meshes, and interpolation can be performed in the latent space. When classified by an SVM, such extracted latent representations perform similarly or better than those extracted by point cloud autoencoders.
[0187] Generally, as an example, a subdivided mesh of level L has a hierarchical face structure, where each face has three adjacent faces (corresponding to its three edges), and one face and its three neighbors can be combined to form a single face, which reverses the loop subdivision operation. This process can be repeated L times, where each iteration reduces the number of faces by a factor of 4 until the base mesh is reached (this occurs when further reduction is not possible). The operations on the subdivided mesh set up a hierarchical pooling and up-pooling scheme that operates globally on the mesh.
[0188] Figure 4A FIG. is a functional block diagram showing an example fixed-length codeword autoencoder with soft decoupling according to some embodiments. For some embodiments, the fixed-length codeword autoencoder with a soft decoupling encoder-decoder architecture can be Figure 4AThe exemplary HetMeshNet encoder-decoder architecture shown in. A mesh autoencoder system (e.g., the HetMeshNet encoder-decoder architecture) extracts fixed-length codewords from heterogeneous meshes of different sizes. To perform convolutions on data with an irregular structure, for some embodiments, the mesh input can be remeshed into a subdivided or semi-regular structure. This remeshing can alleviate the irregularity of the data and enable more image-like convolutions. Doing so can also eliminate the need to explicitly construct or transmit connectivity information at each upsampling step in the decoder. Additionally, for some embodiments, this method is capable of outputting a latent feature map on the base mesh or learning a fixed-length latent representation. For some embodiments, learning a fixed-length latent representation can be achieved by applying global pooling at the end of the encoder and a new module that decouples the latent representation from the base mesh.
[0189] In Figure 4A several items are shown as inputs and / or outputs of various process blocks:
[0190] Item represents a list of features in the input subdivided mesh.
[0191] Item represents a list of features in the intermediate mesh.
[0192] Item represents a list of triangles in the input subdivided mesh.
[0193] Item represents a list of triangles in the base mesh.
[0194] Item represents a list of triangles in the reconstructed output mesh.
[0195] Item represents a list of positions in the input subdivided mesh.
[0196] Item X b represents a list of positions in the base mesh.
[0197] Item represents a list of positions in the intermediate mesh.
[0198] Item represents a list of positions in the reconstructed output mesh.
[0199] Item represents a list of positions on the unit sphere.
[0200] Item represents a list of matching indices on the unit sphere.
[0201] Item Represents a codeword.
[0202] For some embodiments of the autoencoder, such as Figure 4A the autoencoder shown, a single input subdivision mesh can be represented as: (1) a list of positions and (2) a list of triangles which contains indices of corresponding points. Due to the structure of the subdivision mesh, the base mesh is immediately known, with corresponding positions (X b ) and triangles (T b ). The heterogeneous mesh encoder (e.g., HetMeshEnc 404) consumes the input subdivision mesh 402 through a series of DownFaceConv layers ( Figure 4A not shown in ) and outputs an initial feature map on the faces of the base mesh T b . For example, the initial feature map is in w×m b , and the faces of the base mesh are in m b ×3. The DownFaceConv process can include a face convolutional layer followed by a reverse cycle of subdivision pooling. The AdaptMaxPool process 408 is applied on the faces to generate a single latent vector 412, while also deforming the base mesh into a canonical sphere shape 414 using a learnable process (SphereNet 410) and a list of positions 406 on the unit sphere. In one embodiment, AdaptMaxPool simply performs max pooling on the feature map w×m b over the list of m b faces and generates a w-dim latent vector. In another embodiment, a series of face-oriented fully-connected multi-layer perceptron (MLP) layers are applied before the max pooling. The introduced MLP takes each face-oriented feature as input and performs feature aggregation to enhance representability.
[0203] For decoding, according to some embodiments, another learnable process (e.g., DeSphereNet418, which can have the same architecture as SphereNet 410 in some embodiments) is first used to deform the sphere shape and the latent vector back to the base mesh 420. For some embodiments, DeSphereNet 418 can use the list of positions on the unit sphere 416 as input. DeSphereNet 418 can include a series of face convolutional and mesh processing layers Face2Node. By estimating the base mesh and the codeword, the heterogeneous mesh decoder (e.g., HetMeshDec 422) can perform decoding using UpFaceConv layers (a cycle of subdivision upsampling, face convolution, and Face2Node layers) and produce a final reconstructed mesh 424 at the same resolution as the input subdivision mesh 402.Figure 4A Presents a diagram of this overall example pipeline.
[0204] For some embodiments, the Face2Node block is used to transform features from the face domain to the node domain. For some embodiments, the Face2Node block can be used in, for example, HetMeshEncoder blocks, HetMeshDecoder blocks, SpereNet blocks, and / or DeSphereNet blocks. For some embodiments, the focus of the autoencoder is to generate a codeword that is passed through the interface between the encoder and the decoder.
[0205] For some embodiments, the AdaptMaxPool block is architecturally similar to the PointNet block, applying a face-oriented multi-layer perceptron (MLP) process first, followed by a max pooling process, followed by another MLP process. The AdaptMaxPool block treats the face feature map output by the heterogeneous mesh encoder (e.g., HetMeshEnc) as a "point cloud".
[0206] In Figure 4A A complete end-to-end architecture is shown. The SphereNet process can be pre-trained with 3D positions sampled from the unit sphere using the Chamfer loss, and the weights can be fixed when training the rest of the model. During the SphereNet process, the loss for each level of subdivision is implemented at the decoder. For some embodiments, the decoder outputs a total of K + 1 lists of positions. The first is the base mesh reconstruction from the output of the DeSphereNet process, and the remaining K lists of positions are generated by the HetMeshDecoder blocks that output a list of positions for each level of subdivision. Due to the subdivision mesh structure, the correspondence between the input mesh and the output lists of positions is maintained. Thus, the squared L2 loss between each output list of positions and the input mesh geometry is supervised.
[0207] In some embodiments, it is ensured that the face features propagated throughout the model are local face features of the region of the mesh where the face is located. Additionally, the face features have "knowledge" of their global position. Furthermore, the model is invariant to the ordering of the faces or nodes. In this sense, SphereNet locally deforms the regions on the base mesh into a sphere, and the decoder locally deforms the sphere mesh back to the original shape. The global orientation of the shape is maintained within the sphere. In other words, although the model is not guaranteed to be equivariant to 3D rotations, using local feature processing helps to achieve this ability.
[0208] Figure 4BA schematic diagram showing an example fixed - length codeword auto - encoder with soft decoupling is presented. For some embodiments, a fixed - length codeword auto - encoder with a soft - decoupled encoder - decoder architecture can be Figure 4B the example HetMeshNet encoder - decoder architecture shown in Figure 4B The schematic diagram of Figure 4B shows how an example mesh object (table) 452 is transformed at each stage of Figure 4B shows the same example process as Figure 4A the same.
[0209] For some embodiments, the heterogeneous mesh encoder 454 encodes the input mesh object 452 to output an initial feature map 456 on the faces of the base mesh. The AdaptMaxPool process 458 is applied on the faces to generate a latent vector codeword c462. A learnable process (e.g., SphereNet 460) deforms the base mesh into an output sphere shape 464 using a list of input positions on the unit sphere 466. In this application, the sphere shape 464 is also referred to as the base graph or base connectivity. For some embodiments, another learnable process (e.g., DeSphereNet 468) can deform the sphere shape 464 back to the base mesh 470 based on the latent vector codeword c462. By estimating the base mesh 470 and the codeword, the heterogeneous mesh decoder 472 can perform decoding and produce a final reconstructed mesh 474 at the same resolution as the input subdivision mesh 452.
[0210] Figure 5A A functional block diagram showing an example feature - map auto - encoder process according to some embodiments is presented. To generate a feature map 506 instead of a codeword, the AdaptMaxPool block can be skipped, and thus SphereNet and DeSphereNet are not used because the base mesh itself is transmitted to the decoder. Figure 5A A diagram presenting this example process is shown. For some embodiments, the heterogeneous mesh encoder 504 encodes the input mesh 502 into a base mesh 506, and the heterogeneous mesh decoder 508 decodes the input base mesh 506 into an output mesh 510. Two examples of the auto - encoder ( Figure 4A and Figure 5A ) are end - to - end trained with the mesh re - meshed with the ground truth of that stage at each reconstruction stage using the MSE loss.
[0211] Face-centered features can propagate throughout the model. It can be sought that the input features are invariant to the ordering of nodes and faces and the global position or orientation of the faces. Thus, the input face features can be chosen as the normal vector of the face, the face area, and a vector containing curvature information of the face. For face i, let j0, j1, and j2 denote the face indices of its three neighbors. The curvature vector is given by Equation 1:
[0212]
[0213] where c i 、 are the centroid of the face respectively. Thus, for any process block that directly consumes the mesh but uses some input face features, a total of 7 input features are used.
[0214] For some embodiments, a latent feature map 506 on the base mesh can be used, as Figure 5A shown, instead of a single fixed-length latent code. Such a model can be used for comparison with other recent mesh autoencoders that perform latent space comparison between meshes with the same connectivity.
[0215] Figure 5B is a schematic diagram showing an example feature map autoencoder process according to some embodiments. Figure 5B The schematic diagram of Figure 5B shows how the example mesh object (table) 552 is transformed at each stage of Figure 5B In some embodiments, Figure 5A shows the same example process as
[0216] Figure 6A is a functional block diagram showing an example heterogeneous mesh encoder process according to some embodiments. For some embodiments, the heterogeneous mesh encoder process can be the example HetMeshEncoder process shown in Figure 6A The encoding process shown in Figure 6A and named HetMeshEnc includes DownFaceConv layers 604, 606, 608 (in Figure 8Ak repetitions (shown in [[ ]]) to encode the input mesh 602 into the base mesh 610. Each DownFaceConv layer is a pair of FaceConv and SubDivPool. The FaceConv layer (see Hu) is a mesh face feature propagation process for a given subdivided mesh. The FaceConv layer works similarly to a traditional 2D convolution; a learnable kernel defined on the faces of the mesh accesses each face of the mesh and aggregates local features from adjacent faces to produce an updated feature for the current face. Loop, Charles' article "Smooth Subdivision Surfaces Based on Triangles" (1987) ("Loop") discusses subdivision-based pooling / downsampling (SubdivPool). For some embodiments, the SubdivPool block (or layer for some embodiments) combines a set of four adjacent mesh faces into one larger face, and thereby reduces the total number of faces. For some embodiments, the faces can be triangles. Additionally, the features of the combined face (if any) are averaged to obtain the final face feature.
[0217] The two ends of the end-to-end autoencoder architecture can be an encoder block labeled HetMeshEncoder and a decoder block labeled HetMeshDecoder. These blocks perform multi-scale feature processing. HetMeshEncoder extracts features and pools them onto a feature map supported by the faces of the base mesh. At the decoder, HetMeshDecoder receives an approximate version of the base mesh as input and super-resolves the base mesh input back to the original-sized mesh. For some embodiments, Figure 6A the HetMeshEncoder block shown in [[ ]] can be a series of K repetitions of the DownFaceConv layer. The DownFaceConv and FaceConv blocks will be described in more detail below.
[0218] Figure 6B is a schematic diagram showing an example heterogeneous mesh encoder process according to some embodiments. For some embodiments, the heterogeneous mesh encoder process can be Figure 6B the example HetMeshEncoder process shown in [[ ]]. Figure 6B The schematic diagram of [[ ]] shows how the example mesh object (table) 652 is transformed at each stage in [[ ]]. In some embodiments, Figure 6B each stage of [[ ]] shows how the example mesh object (table) 652 is transformed. In some embodiments, Figure 6B shows the same example process as [[ ]]. Figure 6A the same example process as [[ ]]. Figure 6B the encoding process shown in [[ ]] includes K repetitions of the DownFaceConv layers 654, 658, 660 to encode the input mesh 652 into the base mesh 662 via a series of intermediate meshes 656.
[0219] Figure 7A is a functional block diagram showing an example heterogeneous mesh decoder process according to some embodiments. For some embodiments, the heterogeneous mesh decoder process may be Figure 7A the example HetMeshDecoder process shown in to transform the received base mesh 702. Figure 7A The example decoder named HetMeshDecoder shown in includes the following pair of blocks (which may be attached at the position of the dashed arrow in Figure 7A ) repeated K times: UpFaceConv layers 704, 708, 712 and Face2Node layers 706, 710, 714. Figure 7A Each UpFaceConv layer shown in is a pair of FaceConv and SubDivUnpool blocks. The FaceConv block may be the same as in the encoder, while using the subdivision-based upsampling / upsampling block SubDivUnpool. See Loop. Each Face2Node layer 706, 710, 714 may output intermediate lists 716, 718, 720 of the reconstructed positions.
[0220] Figure 7A The HetMeshDecoder block shown in is almost a mirror image of the HetMeshEncoder, except that the HetMeshDecoder block inserts a Face2Node block between each UpFaceConv block. This insertion does not necessarily reconstruct the feature map supported on the original mesh, but rather reconstructs the mesh shape itself defined by the geometric positions. For comparison, in an image, the goal is to reconstruct the feature map corresponding to the pixel values supported by the image. Regarding Figure 11A the Face2Node block discussed in further detail receives face features as input and outputs new face features as well as differential position updates for each node in the mesh. The SubdivUnpool block (or for some embodiments, a layer) inserts new nodes at the midpoints of each edge of the mesh of the previous layer, which subdivides each triangle into four. The face features of these new faces can be copied from their parent faces. The FaceConv layer updates the face features, which can be passed to the Face2Node block to update (all) node positions.
[0221] For a specific iteration a, the Face2Node block outputs a set of positions in the reconstructed mesh For example, the output from the first Face2Node block is For some embodiments, the heterogeneous mesh decoder (e.g., HetMeshDecoder) may be a series of UpFaceConv and Face2Node blocks. For example, the series may be 5 sets of such blocks.
[0222] Figure 7B A schematic diagram showing an example heterogeneous mesh decoder process according to some embodiments. For some embodiments, the heterogeneous mesh decoder process may be Figure 7B the example HetMeshDecoder process shown in Figure 7B The schematic diagram of Figure 7B shows how, for some embodiments, at each stage 754, 758 of Figure 7B the example mesh object (table) 752 is transformed to generate a series of intermediate reconstructed mesh objects 756, 760 and a final reconstructed mesh object 762. In some embodiments, Figure 7A shows the same example process as
[0223] Figure 8A A functional block diagram showing an example face convolution downsampling process according to some embodiments. For some embodiments, the face convolution downsampling process may be Figure 8A the example DownFaceConv process shown in Figure 6A The DownFaceConv layer is a FaceConv layer 804 followed by a SubdivPool layer 806. The FaceConv block determines face neighborhoods and performs a convolutional aggregation operation on the features on the faces. For some embodiments, the DownFaceConv process may transform an input mesh 802 into an output mesh 808. For some embodiments, the Figure 8A shown DownFaceConv process may be performed on the DownFaceConv block shown in
[0224] Figure 8B A schematic diagram showing an example face convolution downsampling process according to some embodiments. For some embodiments, the face convolution downsampling process may be Figure 8B the example DownFaceConv process shown in Figure 8B The schematic diagram of Figure 8B shows how, for some embodiments, at each stage 854, 858 of Figure 8B the example mesh object 852 is transformed to generate an intermediate mesh object 856 and an output mesh object 860. In some embodiments, Figure 8A shows the same example process as
[0225] Figure 9A A functional block diagram showing an example upsampling face convolution process according to some embodiments. For some embodiments, the upsampling face convolution process may be Figure 9AThe example UpFaceConv process shown in. The UpFaceConv layer is a SubdivUnpool layer 904 followed by a FaceConv layer 906. As mentioned above, the FaceConv block determines the face neighborhood and performs a convolutional aggregation operation on the features on the face. According to some embodiments, SubdivUnpool proceeds in the opposite direction of SubdivPool and follows a loop subdivision pattern to transform one mesh face into four smaller faces. The features of the final four smaller faces (if any) are copies of the original larger face. For some embodiments, the UpFaceConv process can transform the input mesh 902 into the output mesh 908. For some embodiments, Figure 7A the UpFaceConv block shown in Figure 9A is executed for the UpFaceConv process shown in.
[0226] Figure 9B is a schematic diagram showing an example upsampling face convolution process according to some embodiments. For some embodiments, the upsampling face convolution process can be Figure 9B the example UpFaceConv process shown in. Figure 9B The schematic diagram of shows how the example mesh object 952 is transformed at each stage 954, 958 for some embodiments to generate the intermediate mesh object 956 and the output mesh object 960. In some embodiments, Figure 9B shows how the example mesh object 952 is transformed at each stage 954, 958 for some embodiments to generate the intermediate mesh object 956 and the output mesh object 960. In some embodiments, Figure 9B shows the same example process as Figure 9A does.
[0227] Figure 10 is a schematic diagram showing an example aggregation of adjacent faces around a node according to some embodiments. Edge vectors are concatenated with the input face features of face i1002, where the starting point is given by the index of node j1004 in face i1002. In this example, for face i, the edge vectors are concatenated in the order of the dashed edge vector first, the solid edge vector second, and the dashed edge vector third. The orientation of the nodes in each face is predefined such that the normal vector points outwards.
[0228] Figure 11A is a functional block diagram showing an example process for converting face features into differential position updates according to some embodiments. For some embodiments, the process for converting face features into differential position updates can be Figure 11A the example Face2Node process shown in. Figure 11A The example Face2Node block shown in directly converts a set of face features into the associated node position updates of the nodes and the updated set of face features. The layer architecture is described below. For some embodiments, Figure 7A the Face2Node block shown in Figure 11AThe Face2Node process shown in
[0229] Loop subdivision-based upsampling / upsampling performs upsampling on the input mesh in a deterministic manner and is similar to naive upsampling in the 2D image domain. Thus, given the input mesh node positions, the output node positions in the upsampled mesh are fixed. According to some embodiments, in order to output the best reconstruction of the input mesh given a codeword, intermediate lower-resolution reconstructions can also be monitored. This monitoring can enable scalable decoding according to the desired decoding resolution and decoder resources, rather than being limited to (always) outputting a reconstruction that matches the resolution of the input mesh.
[0230] The Face2Node block transforms face features into differential position updates in a permutation-invariant manner (with respect to both face and node orderings). Superficially, each face feature carries some information about its region on the surface where the feature is located. All face features on the faces containing node v can be aggregated to update the position of the node.
[0231] The Face2Node layer receives a face list, face features, node positions, and face connectivity as input 1102, and outputs updated node positions corresponding to an intermediate approximation of the input mesh and associated updated face features. The face features can be represented as and the set of node positions can be represented as The Face2Node block reconstructs an enhanced set of node-specific face features G, which, from the perspective of a particular node that is part of these faces, can be regarded as face features.
[0232] For example, let denote the neighborhood of all faces that contain node j as a vertex, and let f i denote the feature of the i-th face. Assume that node j is the k-th node of the i-th mesh face, where k can be 0, 1, or 2. Face2Node concatenates the edge vectors to f i . If x0, x1, and x2 are the geometric positions of the three nodes of face i, the predefined edge vectors are given by equations 2 to 4:
[0233] e0 = x1 - x0 Equation 2
[0234] e1 = x2 - x1 Equation 3
[0235] e2 = x0 - x2 Equation 4
[0236] According to face k ijThe index (and thus the modulus) of the reference node index in [ ], the edge vectors are connected in a cyclic manner. The order of connection is used to maintain permutation invariance with respect to each face. The node indices of the face are sorted in one direction so that the normal vector points outwards. The starting point in the face is set to node j. For example, if node j exactly corresponds to position x1, the edge vectors are connected in the order e1, e2, e0, and the combined feature is given by Equation 5:
[0237] g i1 = [f i |e1|e2|e0] Equation 5
[0238] The other two possibilities for this example are given by Equations 6 and 7:
[0239] g i2 = [f i |e2|e0|e1] Equation 6
[0240] g i0 = [f i |e0|e1|e2] Equation 7
[0241] For notational convenience, for node j in face i, the connected feature will be denoted as g ij . Figure 11A Shows how this process centered on the face is implemented.
[0242] According to the face feature f of node j i Shown in Equation 8, that is:
[0243]
[0244] where k ij represents the index of node j in the i-th mesh face, %3 represents the modulus with respect to 3, and e l represents the j-th edge vector.
[0245] Use the shared MLP block 1106 that operates in parallel on each g ij to update the enhanced node-specific face feature set G1104:
[0246]
[0247] Face2Node updates the connected face features of all three node orderings initiated at node j and outputs the updated node-specific feature set G′1108.
[0248] The differential position update of node j is the average of the first 3 components of g′ ij on the adjacent faces of node j, as shown in Equation 10:
[0249]
[0250] The updated node position 1112 is obtained from the neighborhood average, as shown in Equations 11 and 12:
[0251] x′ j = x j + Δ j Equation 11
[0252]
[0253] where the neighborhood N j is defined by all the faces that include node j as a vertex. The updated face feature f′ i is the average of the updated node-specific features 1110, as shown in Equation 13:
[0254]
[0255] The face feature is updated by averaging all three versions of g′ ij [3:] for the i-th face. The symbol "[3:]" refers to matrix indices 3, 4, 5 …… up to the end of the matrix. F′ refers to the set of updated node-specific features f′ for the set of values of i i , and X′ refers to the set of updated node positions x′ for the set of values of j j .
[0256] Figure 11B is a schematic diagram showing an example process for converting a face feature into a differential position update according to some embodiments. For some embodiments, the process for converting a face feature into a differential position update may be Figure 11B the example Face2Net process shown in Figure 11B The schematic diagram of Figure 11B shows how the example mesh 1150 (a set of triangles) is transformed at each stage in Figure 11B In some embodiments, Figure 11A shows the same example process as Figure 11B For some embodiments, the example mesh 1150 is processed by the MLP process 1152 into a set of node-specific features 1154. As shown in
[0257] Figure 12A is a functional block diagram showing an example process for deforming a base mesh into a canonical sphere shape according to some embodiments. For some embodiments, the process for deforming the base mesh 1202 into the canonical sphere 1216 may be Figure 12AThe example SphereNet process shown in []. In some embodiments of a fixed-length codeword autoencoder, geometric information (mesh vertex positions) is injected into the fixed-length codeword at different scales, particularly the base mesh geometry. Without forcing this information into the codeword, the quality of the codeword may be (severely) degraded, and the codeword may only contain a summary of local face-specific information, which reduces the performance of the codeword when paired with downstream tasks such as classification or segmentation.
[0258] To achieve a (soft) decoupling of this geometric information, one can use Figure 12A the example SphereNet process shown in []. The SphereNet process attempts to match the base mesh geometry with a predefined sphere geometry that includes a set of points sampled on the unit sphere. For some embodiments, this matching can be done by deforming the base mesh geometry into an approximate sphere geometry and using the EMD (Earth Mover's Distance) or Sinkhorn algorithm to match the approximate and actual sphere geometries. For some embodiments, only the output of this matching is transmitted to the decoder. Thus, the codeword learned from the autoencoder (including the encoder and decoder) is forced to learn a better representation of the geometric information. For some embodiments, the SphereNet architecture can be trained separately or co-trained with the overall autoencoder in an end-to-end manner supervised by the chamfer distance. For some embodiments, the SphereNet architecture includes three pairs of FaceConv 1204, 1208, 1212 and Face2Node 1206, 1210, 1214 layers to output the base mesh vertex positions 1216 of the sphere mapping using the face features of the base mesh 1202.
[0259] According to some embodiments, on the decoder side, an example process such as the DeSphereNet process can be used. In some embodiments, the DeSphereNet process can have the same architecture as SphereNet (but different parameters). The DeSphereNet process can be used to reconstruct the base mesh geometry from the matching points on the actual sphere.
[0260] In another implementation, instead of using the learning-based module SphereNet, the deformation / wrapping can be performed via a traditional non-learning-based process. This process can utilize the Laplace operator obtained from the connectivity of the base mesh, i.e., the base graph (also referred to as the sphere shape or base connectivity in this application). In particular, by repeatedly applying the cotangent Laplace operator to the mesh vertex positions, the mesh surface area is minimized by moving the surface along the mean curvature normal direction. The result of this iterative application of the cotangent Laplace operator is a smooth mesh that is very similar to a sphere mesh with the same connectivity as the original base mesh.
[0261] For some embodiments, feature maps on the base mesh can be extracted on the encoder side, and super-resolution capabilities from the feature maps can be extracted on the decoder side. Such a system can be used to extract latent feature maps on the base mesh. For subdivision meshes that contain the same connectivity and face ordering but different geometries, a heterogeneous mesh encoder (e.g., HetMeshEncoder) and a heterogeneous mesh decoder (e.g., HetMeshDecoder) extract (meaningful) latent representations because the feature map sizes on different meshes are the same and aligned with each other.
[0262] However, according to some embodiments, to extend this result to meshes of different connectivity and size, a fixed-length latent code is extracted regardless of size or connectivity and the latent code is decoupled from the base mesh shape. The latter goal stems from the desire to know the connectivity of the base mesh at the decoder. This knowledge of the connectivity of the base mesh is used to perform loop subdivision. If the base mesh geometry is also sent as-is, the geometry also contains relevant information about the mesh shape and limits the information that the latent code can contain.
[0263] At the encoder, a fixed-length latent code is extracted by pooling the feature maps over all faces. For some embodiments, max pooling followed by an MLP layer process can be performed. To decouple the latent code from the base mesh, a SphereNet process is used. The goal of the SphereNet block is to deform the base mesh into a canonical 3D shape. Due to some of the equivalent properties, a sphere is chosen. Ideally, the sphere shape then sent to the decoder should have little information about the shape of the original mesh. For some embodiments, the SphereNet process can be an alternation between FaceConv and Face2Node layers, without upsampling or downsampling. The SphereNet process can be pre-trained with base mesh examples and supervised with a chamfer loss using a random point cloud sampled from the unit sphere. According to some embodiments, the input features are the same as those Figure 5A previously described.
[0264] During the training of the full architecture, the weights of the SphereNet process are fixed, and the predicted sphere geometry is index-matched with a canonical sphere mesh defined by a Fibonacci lattice of the same size as the base mesh geometry. The index matching is performed using the Sinkhorn algorithm with the Euclidean cost between each pair of 3D points. The indices of the sphere mesh corresponding to each of the base mesh geometries are sent to the decoder. This operation ensures that the decoder reconstructs points that lie entirely on the sphere.
[0265] At the decoder, the spherical lattice points are output in the order provided by the index sent from the encoder. Initially, for a heterogeneous mesh decoder (e.g., HetMeshDecoder), these spherical lattice points, along with the latent code and the base mesh connectivity, are reconstructed back to the base mesh and the feature map on the base mesh. As previously described, the face features on the mesh defined by the spherical lattice points and the base mesh connectivity are initialized. The latent code is concatenated to each of these features. These latent code-enhanced face features and the mesh are processed by DeSphereNet blocks that are architecturally equivalent to SphereNet. The output feature map and the mesh are sent to the heterogeneous mesh decoder (e.g., HetMeshDecoder).
[0266] Figure 12B is a schematic diagram showing an example index matching process according to some embodiments. For some embodiments, the index matching process can be Figure 12B the example Sinkhorn process shown in. The predicted sphere geometries 1254 in the order given by the base mesh geometry 1252 are matched one-to-one with the points on the canonical (perfect) sphere lattice. For some embodiments, the Sinkhorn algorithm can be used to compute an approximate minimum-cost bijective between two point sets of equal size.
[0267] Figure 13 is a functional block diagram showing an example fixed-length codeword autoencoder with hard decoupling according to some embodiments. For some embodiments, the heterogeneous mesh encoder 1304 encodes the input subdivision mesh 1302. For some embodiments, the encoded mesh is passed through an AdaptMaxPool process 1306 to generate a codeword and a triangle list in the base mesh 1308.
[0268] For some embodiments, using the matching index to align the input mesh and the reconstructed mesh can only be used to enforce the loss during training. During inference, the base graph can be re-ordered using the matching index before sending it to the decoder. For some embodiments, the matching index may not be sent to the decoder, and the decoder can use the SphereNet process to perform (hard) decoupling.
[0269] In addition to the codeword 1308, an example fixed-length codeword autoencoder with a hard decoupling architecture also sends some information (connectivity + matching index) from the encoder to the decoder to achieve soft decoupling. Hard decoupling can be achieved by sending the codeword and weighted connectivity information without sending the matching index This is achieved by updating the decoder-side pipeline via a block based on a graph neural network (GNN), which may be referred to as the Base Convolutional Grid Reconstruction GNN (BaseConGNN). See A Comprehensive Survey on Graph Neural Networks by Wu, Zonghan, et al., 32.1 IEEE Transactions on Neural Networks and Learning Systems 4-24 (2020).
[0270] The BaseConGNN block 1312 converts the codeword c into a set of local face-specific codewords where C b = b 1c 1310. These local codewords, along with the connectivity information presented as a weighted graph G b from the base grid connectivity, are input into a (standard) GNN architecture block. The GNN block performs graph-aware updates via a shared MLP to transform the local codewords into estimated base grid face features and geometry. The remainder of the decoding pipeline is the same as that shown previously in Figure 4A . Through the estimation of the base grid 1314, the heterogeneous grid decoder 1316 can perform decoding to produce the reconstructed grid 1318. An example fixed-length codeword autoencoder with a hard decoupling architecture is shown in Figure 13 .
[0271] Figure 14 is a functional block diagram showing an example residual face convolution process according to some embodiments. For some embodiments, the residual face convolution process may be the example ResFaceConv process shown in Figure 14 . For some embodiments, the feature aggregation block draws inspiration from the ResNet architecture, as shown in Figure 14 . See Deep Residual Learning for Image Recognition by He, Kaiming, et al., Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (2016). This example shows the architecture of a ResFaceConv (RFC) block for aggregating features with D channels. Figure 14 has a residual connection from the input for adding the input to the output of a series of FaceConv D layers 1402, 1406, 1410 and rectifier linear unit (ReLU) blocks 1404, 1408, 1412.
[0272] The ReLU block refers to the rectified linear unit function. For example, for negative input values, the ReLU block can output 0, and for positive input values, it can output the input multiplied by a scalar value. In another embodiment, the ReLU function can be replaced by other functions, such as the tanh() function and / or the sigmoid() function. For some embodiments, in addition to the rectification function, the ReLU block can also include a non-linear process.
[0273] Figure 15 is a functional block diagram showing an example initial-residual face convolution according to some embodiments. For some embodiments, the initial-residual face convolution process can be Figure 15 the example ResFaceConv process shown in. For some embodiments, the feature aggregation block draws inspiration from the Inception-ResNet architecture, as Figure 15 shown. See Szegedy, Christian, et al., Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning (Inception-v4, Inception ResNet and the Impact of Residual Connections on Learning), Thirty-First AAAI Conference on Artificial Intelligence (2017). This example shows the architecture of an Inception-ResFaceConv (IRFC) block for aggregating features with D channels.
[0274] The IRFC block divides the feature aggregation process into three parallel paths. The path with more convolutional layers ( Figure 15 the left path in) aggregates (more) global information with a larger receptive field. For some embodiments, this aggregation of global information can include two sets of FaceConv D / 4 blocks 1502, 1506, followed by ReLU blocks 1504, 1508. The path with fewer convolutional layers ( Figure 15 the middle path in) aggregates local detailed information with a smaller receptive field. For some embodiments, this aggregation of local information can include a FaceConv D / 4 block 1512, followed by a ReLU block 1514. The last path ( Figure 15 the right path in) is a residual connection that brings the input directly to the output, similar to the residual connection in Figure 14 . For some embodiments, on each of the left path and the middle path in Figure 15 , a ReLU block can be inserted after the FaceConv D / 2 blocks 1510, 1516 and before the concatenation 1518.
[0275] Figure 14 and Figure 15is an example design of HetMeshEnc / HetMeshDec shown in, for example, Figure 4A , Figure 5A and Figure 19 .
[0276] Figure 16 is a functional block diagram showing an example partition - based encoding process according to some embodiments. The previously described architecture is used to encode and decode an entire mesh. However, as the geometric data precision and the point density in the mesh increase, this process can become increasingly time - consuming and computationally expensive. Additionally, the process of converting the original mesh data to a remeshed mesh also takes longer. To address this issue, the original mesh is converted into partitions.
[0277] Figure 16 The input mesh 1602 is shown in the upper left corner. Such an input mesh can be constructed similar to the previously shown input meshes (such as the input mesh of Figure 4A ). For some embodiments, the original input mesh is converted into partitions via a shallow octree process 1604. For each partition, the origin is shifted such that the data points are shifted from the original coordinates to the local coordinates of the partition. For some embodiments, such a shift can be done as part of a local partition remeshing process 1606. Each partition mesh is individually encoded by a heterogeneous mesh encoder 1610 (e.g., HetMeshEnc) to generate a partition bitstream 1614. The auxiliary information about the partitions passing through the shallow octree process 1604 is encoded (compressed) using uniform entropy encoding 1608. The encoded partition bitstream auxiliary information 1612 is added to the partition bitstream 1614 to create a combined bitstream 1616.
[0278] For some embodiments, other partitioning schemes can be used, such as object - based or part - based partitioning. For such embodiments, a shallow octree can be constructed using only the origin of each partition in the original coordinates. Through this process, each partition contains a smaller portion of the mesh, which can be remeshed faster and in parallel for each partition. After compression (encoding) and decompression (decoding), the meshes recovered from all partitions are combined and brought back to the original coordinates.
[0279] Figure 17 is a functional block diagram showing an example partition - based decoding process according to some embodiments. In Figure 17The left side of the decoding process 1700 shows the combined bitstream input 1702. The bitstream is split into auxiliary information bits 1704 and mesh partition bits 1706. The mesh partition bits 1706 are decoded using a heterogeneous mesh decoder 1710 (e.g., HetMeshDec) to generate a reconstructed partitioned mesh 1714. The auxiliary information bits 1704 are decoded (decompressed) using a uniform entropy decoder 1708 and sent to the shallow block partition octree process 1712. The shallow block partition octree process 1712 combines the reconstructed partitioned mesh 1714 with the decoded auxiliary information to generate a reconstructed mesh 1716. The decoded auxiliary information includes information about the partitions to enable the shallow block partition octree blocks to generate the reconstructed mesh. For some embodiments, this information may include information indicating the amount by which the partitions are shifted back from local coordinates to the original coordinates.
[0280] Figure 18 is a functional block diagram showing an example mesh classification architecture based on a fixed-length codeword autoencoder according to some embodiments. Figure 18 shows an end-to-end learning-based mesh encoder framework (e.g., HetMeshEnc) that produces frames mimicking those in the image domain, which is a process 1800 capable of operating on meshes of different sizes and connectivities while having a fixed-length latent representation. Additionally, when optimal reconstruction performance is desired rather than a fixed-length codeword summarizing the global topology of the mesh, the encoder and decoder blocks can be adapted to (respectively) produce and digest latent feature maps 1802 residing on a low-resolution base mesh. The codewords produced by HetMeshEnc 1804, followed by the AdaptMaxPool process 1806, are passed through an additional MLP block 1808, the output dimension of which matches the number of different mesh classes to be classified. The Softmax process 1808 converts the output values to class scores 1810 in the range [0, 1]. The class with the highest score is the predicted class for classification.
[0281] Figure 19 is a functional block diagram showing an example fixed-length codeword autoencoder with soft decoupling and SphereNet according to some embodiments. For some embodiments, an end-to-end learning-based mesh autoencoder framework HetMeshNet can mimic those in the image domain, which is a process 1900 capable of operating on meshes of different sizes and connectivities while producing a useful fixed-length latent representation. Additionally, when optimal reconstruction performance is desired rather than a fixed-length codeword summarizing the global topology of the mesh, the proposed encoder and decoder modules can be adapted to (respectively) produce and digest latent feature maps on a low-resolution base mesh.
[0282] A heterogeneous mesh encoder (e.g., HetMeshEnc 1904) encodes the input subdivision mesh 1902 and outputs an initial feature map on the faces of the base mesh. The AdaptMaxPool process 1908 is applied on the faces to generate a latent vector 1912, while also deforming the base mesh into a canonical sphere shape using a learnable process (SphereNet 1910) and a list of sampling positions 1906 on the unit sphere, which results in a base map ( 1914) (also referred to as base connectivity in this application). For some embodiments, additional modification requests may be found in the heterogeneous mesh encoder.
[0283] For decoding, according to some embodiments, another learnable process (e.g., DeSphereNet 1916, which may have the same architecture as SphereNet 1910 in some embodiments) is first used to deform the sphere shape and the latent vector back to a list of positions in the base mesh A list of features in the base mesh and the base map 1918. For some embodiments, DeSphereNet 1916 may use the list of positions on the unit sphere 1906 as input. DeSphereNet 1916 may include a series of face convolutions and the Face2Node mesh processing layer. By estimating the base mesh and the codeword, the heterogeneous mesh decoder (e.g., HetMeshDec 1920) produces the final reconstructed mesh 1922 at the same resolution as the input subdivision mesh 1902.
[0284] Figure 20 is a flowchart showing an example encoding method according to some embodiments. For some embodiments, the process 2000 for encoding mesh data is shown in Figure 20 There is a start box 2002 shown, and the process proceeds to box 2004 to determine the initial mesh face features from the input mesh. Control proceeds to box 2006 to determine a base mesh including a set of face features based on a first learning-based process, which may include a series of mesh feature extraction layers. Control proceeds to box 2008 to generate a fixed-length codeword from the base mesh, which can be done on the mesh faces using a second learning-based pooling process. Control proceeds to box 2010 to generate a base map by matching the vertices of a predefined template mesh with the vertices of the base mesh using a third process, which may be a learning-based pooling process.
[0285] Figure 21 is a flowchart showing an example decoding method according to some embodiments. For some embodiments, in Figure 21Shown therein is a process 2100 for decoding mesh data. A start box 2102 is shown, and for some embodiments, the process proceeds to box 2104 to determine a reconstructed base mesh and a base surface feature map using a fixed codeword and a base graph via a first learning-based module in the presence of a predefined template mesh. For some embodiments, control proceeds to box 2106 to generate at least one reconstructed mesh at multiple hierarchical resolutions by a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
[0286] Figure 22 FIG. [0000834] is a flowchart showing an example encoding process according to some embodiments. For some embodiments, the example process may include accessing an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions. For some embodiments, the example process may further include generating at least two initial mesh surface features for at least one face listed in the list of faces of the input mesh. For some embodiments, the example process may further include generating a base mesh and at least two base mesh surface features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh. For some embodiments, the example process may further include generating a fixed-length codeword from the at least two base mesh surface features. For some embodiments, the example process may further include accessing a predefined template mesh. For some embodiments, the example process may further include generating a matching index set, where the matching index set indicates matching vertices between the predefined template mesh and the base mesh. For some embodiments, the example process may further include outputting a fixed-length codeword, information indicating the connectivity of the base mesh, and the matching index set.
[0287] Figure 23 FIG. [0000837] is a flowchart showing an example decoding process according to some embodiments. For some embodiments, the example process may include receiving a fixed-length codeword, information indicating the connectivity of the base mesh, and a matching index set to generate a reconstructed base mesh and at least two base surface features. For some embodiments, the example process may further include generating a reconstructed base mesh and at least two base surface features. For some embodiments, the example process may further include generating at least one reconstructed mesh at at least two hierarchical resolutions.
[0288] While methods and systems according to some embodiments are generally discussed in the context of extended reality (XR), some embodiments may be applied to any XR context, e.g., such as virtual reality (VR) / mixed reality (MR) / augmented reality (AR) contexts. Additionally, although the term "head-mounted display (HMD)" is used herein according to some embodiments, for some embodiments, some embodiments may be applied to wearable devices (which may or may not be attached to the head) capable of supporting, e.g., XR, VR, AR, and / or MR.
[0289] A first example method according to some embodiments may include: accessing a semi-regular input mesh to generate initial mesh face features for each mesh face, wherein the semi-regular input mesh includes a list of faces and a plurality of vertex positions; generating a base mesh including vertex positions and information indicating base connectivity and a set of face features on the base mesh through a learning-based feature aggregation module; generating a fixed-length codeword based on the base face features using a feature pooling module; accessing a predefined template mesh and the base mesh to generate a set of matching indices, the set of matching indices including indices indicating matching vertices between the predefined template mesh and the base mesh; and outputting the generated fixed-length codeword, the information indicating the base connectivity, and the set of matching indices.
[0290] A second example method according to some embodiments may include: accessing an input remeshed mesh to generate initial mesh face features, wherein the input remeshed mesh includes a list of faces and vertex positions; generating a base mesh and an atlas of face features on the base mesh; generating a fixed-length codeword from the base face features; accessing a predefined number of vertices and a predefined spherical mesh of base mesh vertices to generate a set of spherical matching indices; and outputting the generated fixed-length codeword, base mesh connectivity information, and the set of spherical matching indices.
[0291] A third example method according to some embodiments may include: accessing an input mesh, wherein the input mesh includes a list of faces and a plurality of vertex positions; generating at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; generating a base mesh and at least two base mesh face features on the base mesh, wherein the base mesh includes vertex positions and information indicating base mesh connectivity; generating a fixed-length codeword from the at least two base mesh face features; accessing a predefined template mesh; generating a set of matching indices, wherein the set of matching indices indicates matching vertices between the predefined template mesh and the base mesh; and outputting the fixed-length codeword, the information indicating the base mesh connectivity, and the set of matching indices.
[0292] For some embodiments of the third example method, the input mesh is a semi-regular mesh.
[0293] For some embodiments of the third example method, generating the base mesh may include: generating the vertex positions; and generating the information indicating the base mesh connectivity.
[0294] For some embodiments of the third example method, generating the at least two base mesh face features on the base mesh is performed by performing learning-based aggregation on the at least two initial mesh face features.
[0295] For some embodiments of the third example method, generating the fixed-length codeword is performed by pooling the at least two base mesh surface features.
[0296] For some embodiments of the third example method, the predefined template mesh is a mesh corresponding to a unit sphere.
[0297] For some embodiments of the third example method, the information indicating the base connectivity includes a triangle list with information having an indication index, the index corresponding to a matching vertex indicated by a set of matching indices.
[0298] For some embodiments of the third example method, generating the base mesh and the at least two base mesh surface features on the base mesh can be performed by a learned heterogeneous mesh encoder, and the heterogeneous mesh encoder can include at least one downsampling surface convolutional layer.
[0299] For some embodiments of the third example method, generating the fixed-length codeword from the at least two base mesh surface features can include using a learned AdaptMaxPool process.
[0300] For some embodiments of the third example method, the set of matching indices can be generated by a learned SphereNet process.
[0301] A first example device according to some embodiments can include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access a semi-regular input mesh to generate an initial mesh surface feature for each mesh surface, where the semi-regular input mesh includes a face list and a plurality of vertex positions; generate a base mesh including vertex positions and information indicating base connectivity and a set of face features on the base mesh by a learned feature aggregation module; generate a fixed-length codeword based on the base face features using a feature pooling module; access a predefined template mesh and the base mesh to generate a set of matching indices, the set of matching indices including indices indicating matching vertices between the predefined template mesh and the base mesh; and output the generated fixed-length codeword, the information indicating the base connectivity, and the set of matching indices.
[0302] A second example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input remeshed grid to generate initial mesh surface features, wherein the input remeshed grid includes a face list and vertex positions; generate a base mesh and an atlas of surface features on the base mesh; generate a fixed-length codeword from the base surface features; access a predefined number of vertices and a predefined spherical mesh of base mesh vertices to generate a set of sphere matching indices; and output the generated fixed-length codeword, base mesh connectivity information, and the set of sphere matching indices.
[0303] A third example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access an input mesh, wherein the input mesh includes a face list and a plurality of vertex positions; generate at least two initial mesh surface features for at least one face listed on the face list of the input mesh; generate a base mesh and at least two base mesh surface features on the base mesh, wherein the base mesh includes vertex positions and information indicating base mesh connectivity; generate a fixed-length codeword from the at least two base mesh surface features; access a predefined template mesh; generate a set of matching indices, wherein the set of matching indices indicates matching vertices between the predefined template mesh and the base mesh; and output the fixed-length codeword, information indicating the base mesh connectivity, and the set of matching indices.
[0304] A fourth example method according to some embodiments may include: accessing an input mesh; partitioning the input mesh into a first input mesh and a second input mesh, where the first input mesh includes a first list of faces and a first plurality of vertex positions, and where the second input mesh includes a second list of faces and a second plurality of vertex positions; generating at least two first initial mesh face features for at least one first face listed on the first list of faces of the first input mesh; generating a first base mesh and at least two first base mesh face features on the first base mesh, where the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; generating a first fixed-length codeword from the at least two first base mesh face features; accessing a first predefined template mesh; generating a first matching index set, where the first matching index set indicates first matching vertices between the first predefined template mesh and the first base mesh; and outputting the first fixed-length codeword, the first information indicating the connectivity of the first base mesh, and the first matching index set; generating at least two second initial mesh face features for at least one second face listed on the second list of faces of the second input mesh; generating a second base mesh and at least two second base mesh face features on the second base mesh, where the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; generating a second fixed-length codeword from the at least two second base mesh face features; accessing a second predefined template mesh; generating a second matching index set, where the second matching index set indicates second matching vertices between the second predefined template mesh and the second base mesh; and outputting the second fixed-length codeword, the second information indicating the connectivity of the second base mesh, and the second matching index set.
[0305] A fifth example method according to some embodiments may include: accessing base connectivity information and a sphere matching index set to generate a reconstructed base mesh and a base face feature map through a learning-based module DeSphereNet; and generating K reconstructed meshes at K hierarchical resolutions through a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
[0306] A sixth example method according to some embodiments may include: accessing base mesh connectivity information, a fixed-length codeword, and a sphere matching index set to generate a reconstructed base mesh and a base face feature map; and generating K reconstructed meshes at K hierarchical resolutions.
[0307] A seventh example method according to some embodiments may include: receiving a fixed-length codeword, information indicating base grid connectivity, and a set of matching indices to generate a reconstructed base grid and at least two base surface features; generating the reconstructed base grid and the at least two base surface features; and generating at least one reconstructed grid for at least two hierarchical resolutions.
[0308] For some embodiments of the seventh example method, generating the at least one reconstructed grid generates K reconstructed grids for K hierarchical resolutions.
[0309] For some embodiments of the seventh example method, a heterogeneous grid decoder is used to generate the K reconstructed grids.
[0310] For some embodiments of the seventh example method, the heterogeneous grid decoder performs at least one upsampling face convolution process and at least one Face2Node process.
[0311] For some embodiments of the seventh example method, generating the at least one reconstructed grid generates at least two reconstructed grids for at least two corresponding hierarchical resolutions.
[0312] For some embodiments of the seventh example method, generating the reconstructed base grid may be performed by a learning-based DeSphereNet process.
[0313] For some embodiments of the seventh example method, generating the at least one reconstructed grid for at least two hierarchical resolutions includes: determining input surface features from the base surface feature map; generating updated surface features corresponding to the input surface features; determining updated differential positions of one or more nodes of the reconstructed grid; and using the corresponding updated differential positions to update the positions of one or more nodes of the reconstructed base grid.
[0314] A fifth example device according to some embodiments may include: a processor; a memory storing instructions that, when executed by the processor, are operable to cause a grid decoder to: access base connectivity information and a sphere matching index set to generate a reconstructed base grid and a base surface feature map through a learning-based module DeSphereNet; and generate K reconstructed grids at K hierarchical resolutions through a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
[0315] A sixth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access base mesh connectivity information, fixed-length codewords, and a sphere matching index set to generate a reconstructed base mesh and a base surface feature map; and generate K reconstructed meshes at K hierarchical resolutions.
[0316] A seventh example device according to some embodiments may include: a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh decoder to: receive fixed-length codewords, information indicating base mesh connectivity, and a matching index set to generate a reconstructed base mesh and at least two base surface features; generate a reconstructed base mesh and at least two base surface features; and generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0317] An eighth example device according to some embodiments may include a mesh decoder configured to obtain fixed-length codewords, base connectivity information, and a sphere matching index set and generate a reconstructed mesh: wherein the mesh decoder is configured to: access the base connectivity information and the sphere matching index set to generate a reconstructed base mesh and a base surface feature map via a learning-based module DeSphereNet; and generate K reconstructed meshes at K hierarchical resolutions via a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
[0318] A first example method according to some embodiments may include: accessing a semi-regular input mesh to generate initial mesh surface features for each mesh surface, wherein the semi-regular input mesh includes a face list and a plurality of vertex positions; generating a base mesh including vertex positions and information indicating base connectivity and a set of face features on the base mesh via a learning-based feature aggregation module; generating fixed-length codewords based on the base surface features using a feature pooling module; accessing a predefined template mesh and the base mesh to generate information indicating matching vertices between the predefined template mesh and the base mesh; and outputting the generated fixed-length codewords and the information indicating the base connectivity.
[0319] A second example method according to some embodiments may include: accessing an input remeshed mesh to generate initial mesh surface features, wherein the input remeshed mesh includes a face list and vertex positions; generating a base mesh and a set of face feature maps on the base mesh; generating fixed-length codewords from the base surface features; accessing a predefined number of vertices and a predefined sphere mesh of base mesh vertices to generate a match between the sphere mesh vertices and the base network vertices; and outputting the generated fixed-length codewords and the base mesh connectivity information.
[0320] The third example method according to some embodiments may include: accessing an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; generating at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; generating a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; generating a fixed-length codeword from the at least two base mesh face features; accessing a predefined template mesh; generating information indicating matching vertices between the predefined template mesh and the base mesh; and outputting the fixed-length codeword and the information indicating the connectivity of the base mesh.
[0321] For some embodiments of the third example method, the input mesh is a semi-regular mesh.
[0322] For some embodiments of the third example method, generating the base mesh may include: generating the vertex positions; and generating information indicating the connectivity of the base mesh.
[0323] For some embodiments of the third example method, generating the at least two base mesh face features on the base mesh is performed by performing a learning-based aggregation on the at least two initial mesh face features.
[0324] For some embodiments of the third example method, generating the fixed-length codeword is performed by pooling the at least two base mesh face features.
[0325] For some embodiments of the third example method, the predefined template mesh is a mesh corresponding to a unit sphere.
[0326] For some embodiments of the third example method, the information indicating the base connectivity includes a list of triangles with information having an index corresponding to the matching vertices indicated by a set of matching indices.
[0327] For some embodiments of the third example method, generating the base mesh and the at least two base mesh face features on the base mesh is performed by a learning-based heterogeneous mesh encoder, and the heterogeneous mesh encoder includes at least one downsampling face convolutional layer.
[0328] For some embodiments of the third example method, generating the fixed-length codeword from the at least two base mesh face features includes using a learning-based AdaptMaxPool process.
[0329] For some embodiments of the third example method, the matching indices are generated by a learning-based SphereNet process.
[0330] Some embodiments of the third example method may further include: outputting information indicating matching vertices, where the information indicating matching vertices includes a set of matching indices that indicate the matching vertices between the predefined template mesh and the base mesh.
[0331] A first example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access a semi-regular input mesh to generate initial mesh face features for each mesh face, where the semi-regular input mesh includes a list of faces and a plurality of vertex positions; generate a base mesh including vertex positions and information indicating base connectivity and a set of face features on the base mesh via a learning-based feature aggregation module; generate a fixed-length codeword based on the base face features using a feature pooling module; access a predefined template mesh and the base mesh to generate information indicating matching vertices between the predefined template mesh and the base mesh; and output the generated fixed-length codeword and the information indicating the base connectivity.
[0332] A second example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input re-meshed mesh to generate initial mesh face features, where the input re-meshed mesh includes a list of faces and vertex positions; generate a base mesh and an atlas of face features on the base mesh; generate a fixed-length codeword from the base face features; access a predefined number of vertices and a predefined spherical mesh of base mesh vertices to generate a match between the spherical mesh vertices and the base network vertices; and output the generated fixed-length codeword and the base mesh connectivity information.
[0333] A third example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: access an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; generate at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating base mesh connectivity; generate a fixed-length codeword from the at least two base mesh face features; access a predefined template mesh; generate information indicating a match of vertices between the predefined template mesh and the base mesh; and output the fixed-length codeword and the information indicating the base mesh connectivity.
[0334] A fourth example method according to some embodiments may include: accessing base connectivity information and a predefined sphere mesh to generate a reconstructed base mesh and a base face feature map via a learning-based module, DeSphereNet; and generating K reconstructed meshes at K hierarchical resolutions via a learning-based module, HetMeshDec, consisting of a series of K pairs of UpFaceConv and Face2Node modules.
[0335] A fifth example method according to some embodiments may include: accessing base mesh connectivity information, a fixed-length codeword, and a predefined sphere mesh to generate a reconstructed base mesh and a base face feature map; and generating K reconstructed meshes at K hierarchical resolutions.
[0336] A sixth example method according to some embodiments may include: receiving a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; generating the reconstructed base mesh and the at least two base face features; and generating at least one reconstructed mesh for at least two hierarchical resolutions.
[0337] For some embodiments of the sixth example method, generating the at least one reconstructed mesh generates K reconstructed meshes for K hierarchical resolutions.
[0338] For some embodiments of the sixth example method, a heterogeneous mesh decoder is used to generate K reconstructed meshes.
[0339] For some embodiments of the sixth example method, the heterogeneous mesh decoder performs at least one upsampling face convolution process and at least one Face2Node process.
[0340] For some embodiments of the sixth example method, generating the at least one reconstructed mesh generates at least two reconstructed meshes for at least two corresponding hierarchical resolutions.
[0341] For some embodiments of the sixth example method, generating the reconstructed base mesh is performed via a learning-based DeSphereNet process.
[0342] For some embodiments of the sixth example method, generating the at least one reconstructed mesh for at least two hierarchical resolutions includes: determining input face features from the base face feature map; generating updated face features corresponding to the input face features; determining updated differential positions of one or more nodes of the reconstructed mesh; and using the corresponding updated differential positions to update the positions of one or more nodes of the reconstructed base mesh.
[0343] A fourth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause a mesh decoder to: access base connectivity information and a predefined sphere mesh to generate a reconstructed base mesh and a base face feature map via a learning-based module, DeSphereNet; and generate K reconstructed meshes at K hierarchical resolutions via a learning-based module, HetMeshDec, consisting of a series of K pairs of UpFaceConv and Face2Node modules.
[0344] A fifth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access base mesh connectivity information, fixed-length codewords, and a predefined sphere mesh to generate a reconstructed base mesh and a base face feature map; and generate K reconstructed meshes at K hierarchical resolutions.
[0345] A sixth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause a mesh decoder to: receive fixed-length codewords, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; generate a reconstructed base mesh and at least two base face features; and generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0346] A mesh decoder configured to obtain fixed-length codewords, base connectivity information, and a set of sphere matching indices and generate a reconstructed mesh according to some embodiments may be configured to: access the base connectivity information and a predefined sphere mesh to generate a reconstructed base mesh and a base face feature map via a learning-based module, DeSphereNet; and generate K reconstructed meshes at K hierarchical resolutions via a learning-based module, HetMeshDec, consisting of a series of K pairs of UpFaceConv and Face2Node modules.
[0347] A seventh example method according to some embodiments may include: determining an initial mesh face feature from an input mesh; determining a base mesh including a set of face features based on a first learning-based module including a series of mesh feature extraction layers; generating fixed-length codewords from the base mesh on the mesh face using a second learning-based pooling module; and generating a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
[0348] A seventh example device according to some embodiments includes a memory and a processor, the processor being configured to perform: determining initial mesh surface features from an input mesh; determining a base mesh including a set of surface features based on a first learning-based module including a series of mesh feature extraction layers; generating a fixed-length codeword from the base mesh on the mesh surface using a second learning-based pooling module; and generating a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
[0349] An eighth example method according to some embodiments may include: in the presence of a predefined template mesh, determining a reconstructed base mesh and a base surface feature map via a first learning-based module using a fixed codeword and a base graph; generating at least one reconstructed mesh at multiple hierarchical resolutions via a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
[0350] An eighth example device according to some embodiments may include a memory and a processor, the processor being configured to perform: in the presence of a predefined template mesh, determining a reconstructed base mesh and a base surface feature map via a first learning-based module using a fixed codeword and a base graph; generating at least one reconstructed mesh at multiple hierarchical resolutions via a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
[0351] A ninth example method according to some embodiments may include: a heterogeneous mesh encoder including a series of layers including pairs of mesh feature extraction modules and mesh downsampling modules; and a heterogeneous mesh decoder including a learning-based module including a series of layers including pairs of mesh node generation modules and mesh upsampling modules.
[0352] For some embodiments of the ninth example device, a base mesh is transmitted from the heterogeneous mesh encoder to the heterogeneous mesh decoder.
[0353] For some embodiments of the ninth example device, multiple input features are used in addition to the directly consumed mesh.
[0354] For some embodiments of the ninth example method, the loop subdivision-based upsampling module includes: constructing an enhanced node-specific surface feature set; updating the enhanced node-specific surface feature set using a shared module; averaging the updated node-specific surface features; and performing neighborhood averaging on node positions.
[0355] Some embodiments of the ninth example method may further include: converting a codeword into a face-specific codeword set; and transforming the face-specific codeword into a base mesh feature and a geometry.
[0356] Some embodiments of the ninth example method may further include: converting an original mesh into partitions; shifting the origin of the partitions; and encoding or decoding each partition mesh separately.
[0357] For some embodiments of the ninth example method, the meshes have different sizes and connectivities.
[0358] A tenth example device according to some embodiments may include a non-transitory computer-readable medium that contains data content generated according to any of the methods listed above for playback using a processor.
[0359] A first example signal according to some embodiments may include: video data generated according to any of the methods listed above for playback using a processor.
[0360] An example computer program product according to some embodiments may include instructions that, when the program is executed by a computer, cause the computer to perform any of the methods listed above.
[0361] A first non-transitory computer-readable medium according to some embodiments may include data content that includes instructions for performing any of the methods listed above.
[0362] For some embodiments of the seventh example device, the third module is a learning-based module.
[0363] For some embodiments of the seventh example device, the third module is a traditional non-learning-based module.
[0364] An eleventh example method according to some embodiments may include: accessing an input mesh, where the input mesh includes a face list and a plurality of vertex positions; generating at least two initial mesh face features for at least one face listed in the face list of the input mesh; generating a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; accessing a predefined template mesh; generating information indicating matching vertices between the predefined template mesh and the base mesh; and outputting information indicating the connectivity of the base mesh.
[0365] The eleventh example device according to some embodiments may include a processor; a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; generate at least two initial mesh face features for at least one face listed on the list of faces of the input mesh; generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; access a predefined template mesh; generate information indicating matching vertices between the predefined template mesh and the base mesh; and output information indicating the connectivity of the base mesh.
[0366] The twelfth example method according to some embodiments may include: receiving information indicating the connectivity of a base mesh and a predefined mesh to generate a reconstructed base mesh and at least two base face features; generating a reconstructed base mesh and at least two base face features; and generating at least one reconstructed mesh for at least two hierarchical resolutions.
[0367] The twelfth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: receive information indicating the connectivity of a base mesh and a predefined mesh to generate a reconstructed base mesh and at least two base face features; generate a reconstructed base mesh and at least two base face features; and generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0368] The thirteenth example method according to some embodiments may include: accessing an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; performing a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed on the list of faces of the input mesh; performing an AdaptMaxPool process to thereby: generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; and generate a fixed-length codeword from the at least two base mesh face features; output the fixed-length codeword and information indicating the connectivity of the base mesh.
[0369] A thirteenth example device according to some embodiments may include: a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input mesh, where the input mesh includes a list of faces and a plurality of vertex positions; perform a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; perform an AdaptMaxPool process to: generate a base mesh and at least two base mesh face features on the base mesh, where the base mesh includes vertex positions and information indicating the connectivity of the base mesh; and generate a fixed-length codeword from the at least two base mesh face features; output the fixed-length codeword and the information indicating the connectivity of the base mesh.
[0370] A fourteenth example method according to some embodiments may include: receiving a fixed-length codeword, information indicating the connectivity of a base mesh, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; performing a Base Convolutional Graph Neural Network (BaseConGNN) process to generate a reconstructed base mesh and at least two base face features; and performing a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0371] A fourteenth example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: receive a fixed-length codeword, information indicating the connectivity of a base mesh, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; perform a Base Convolutional Graph Neural Network (BaseConGNN) process to generate a reconstructed base mesh and at least two base face features; and perform a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
[0372] A fifteenth example method according to some embodiments may include: accessing an input mesh; partitioning the input mesh into a first input mesh and a second input mesh, wherein the first input mesh includes a first list of faces and a first plurality of vertex positions, and wherein the second input mesh includes a second list of faces and a second plurality of vertex positions; generating, for at least one first face listed on the first list of faces of the first input mesh, at least two first initial mesh face features; generating a first base mesh and at least two first base mesh face features on the first base mesh, wherein the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; generating a first fixed-length codeword from the at least two first base mesh face features; accessing a first predefined template mesh; outputting the first fixed-length codeword and the first information indicating the connectivity of the first base mesh; generating, for at least one second face listed on the second list of faces of the second input mesh, at least two second initial mesh face features; generating a second base mesh and at least two second base mesh face features on the second base mesh, wherein the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; generating a second fixed-length codeword from the at least two second base mesh face features; accessing a second predefined template mesh; and outputting the second fixed-length codeword and the second information indicating the connectivity of the second base mesh.
[0373] Some embodiments of the fifteenth example method may further include: generating a first set of matching indices, wherein the first set of matching indices indicates first matching vertices between the first predefined template mesh and the first base mesh; outputting the first set of matching indices; generating a second set of matching indices, wherein the second set of matching indices indicates second matching vertices between the second predefined template mesh and the second base mesh; and outputting the second set of matching indices.
[0374] The fifteenth example device according to some embodiments may include a processor; and a memory storing instructions that, when executed by the processor, are operable to cause the processor to: access an input mesh; partition the input mesh into a first input mesh and a second input mesh, where the first input mesh includes a first list of faces and a first plurality of vertex positions, and where the second input mesh includes a second list of faces and a second plurality of vertex positions; for at least one first face listed on the first list of faces of the first input mesh, generate at least two first initial mesh face features; generate a first base mesh and at least two first base mesh face features on the first base mesh, where the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; generate a first fixed-length codeword from the at least two first base mesh face features; access a first predefined template mesh; output the first fixed-length codeword and the first information indicating the connectivity of the first base mesh; for at least one second face listed on the second list of faces of the second input mesh, generate at least two second initial mesh face features; generate a second base mesh and at least two second base mesh face features on the second base mesh, where the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; generate a second fixed-length codeword from the at least two second base mesh face features; access a second predefined template mesh; and output the second fixed-length codeword and the second information indicating the connectivity of the second base mesh.
[0375] The sixteenth example device according to some embodiments may include: at least one processor configured to perform any of the methods listed above.
[0376] The seventeenth example device according to some embodiments may include a computer-readable storage medium storing instructions for causing one or more processors to perform any of the methods listed above.
[0377] The eighteenth example device according to some embodiments may include: at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform any of the methods listed above.
[0378] The second example signal according to some embodiments may include: a bitstream generated according to any of the methods listed above.
[0379] The various implementations relate to decoding. As used in this application, "decoding" can cover all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include or alternatively include processes performed by the encoders of the various implementations described in this application.
[0380] As another example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a more general decoding process, and it is considered that those skilled in the art will understand well.
[0381] The various implementations relate to encoding. In a manner similar to the discussion above regarding "decoding", "encoding" as used in this application can include, for example, all or part of the processing performed on an input video sequence to produce a coded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as partitioning, differential encoding, transform, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by the decoders of the various implementations described in this application.
[0382] As another example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of entropy encoding and differential encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to a more general encoding process, and it is considered that those skilled in the art will understand well.
[0383] It should be noted that the syntactic elements used herein are descriptive terms. Therefore, they do not exclude the use of other syntactic element names.
[0384] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / device.
[0385] Various embodiments may relate to parametric models and rate - distortion optimization. In particular, during the encoding process, a balance or trade - off between rate and distortion is typically considered, often in view of constraints on computational complexity. It can be measured by a rate - distortion optimization (RDO) metric or by least mean square (LMS), mean absolute error (MAE), or other such measures. Rate - distortion optimization is typically formulated as minimizing a rate - distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate - distortion optimization problem. For example, these methods can be based on an extensive test of all encoding options, including all considered modes or encoding parameter values, where the encoding cost and the associated distortion of the reconstructed signal are comprehensively evaluated after encoding and decoding. Faster methods can also be used to save encoding complexity, particularly by calculating an approximate distortion based on the predicted or prediction - residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of various techniques to perform the optimization, but the optimization is not necessarily a comprehensive evaluation of encoding cost and associated distortion.
[0386] The implementations and aspects described herein can be implemented, for example, in a method or process, a device, a software program, a data stream, or a signal. Even when discussed in the context of only a single implementation form (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., a device or a program). A device can be implemented, for example, with appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, for example, such as a computer, a cellular phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication between end - users.
[0387] Reference to “an embodiment” or “embodiments” or “an implementation” or “implementations” and other variations thereof means that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the appearances of the phrases “in an embodiment” or “in embodiments” or “in an implementation” or “in implementations” and any other variations thereof throughout this application do not necessarily all refer to the same embodiment.
[0388] Additionally, this application may relate to “determining” various information. Determining information can include one or more of the following: for example, estimating information, calculating information, predicting information, or retrieving information from a memory.
[0389] In addition, the present application may relate to "accessing" various information. Accessing information may include one or more of the following: for example, receiving information, retrieving information (e.g., retrieving from a memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0390] Additionally, the present application may relate to "receiving" various information. Like "accessing", receiving is a broad term. Receiving information may include one or more of the following: for example, accessing information or retrieving information (e.g., retrieving from a memory). Further, "receiving" is typically involved in one way or another during operations, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information.
[0391] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of " / " and "and / or" and "at least one" is intended to cover only selecting the first-listed option (A), or only selecting the second-listed option (B), or selecting both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include only selecting the first-listed option (A), or only selecting the second-listed option (B), or only selecting the third-listed option (C), or only selecting the first-listed option and the second-listed option (A and B), or only selecting the first-listed option and the third-listed option (A and C), or only selecting the second-listed option and the third-listed option (B and C), or selecting all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to any number of items listed.
[0392] In addition, as used herein, the word "signal" particularly refers to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a particular one of a plurality of transforms, coding modes, or flags. In this way, in an embodiment, the same transform, parameter, or mode is used on both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to a decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntactic elements, flags, etc. are used to send information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.
[0393] It will be apparent to those of ordinary skill in the art that implementations can generate a variety of signals that are formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. A signal can be stored on a processor-readable medium.
[0394] The foregoing sections have described many embodiments across various claim categories and types. The features of these embodiments can be provided individually or in any combination. In addition, embodiments can individually or in any combination across various claim categories and types include one or more of the following features, devices, or aspects:
[0395] One embodiment includes an apparatus that includes a learning-based heterogeneous grid autoencoder.
[0396] Other embodiments include methods for performing learning-based heterogeneous grid autoencoding.
[0397] Other embodiments include the above-described methods and apparatuses performing face feature initialization.
[0398] Other embodiments include the above-described methods and apparatuses performing heterogeneous grid encoding and / or decoding.
[0399] Other embodiments include the method and apparatus above performing soft decoupling or hard decoupling.
[0400] Other embodiments include the method and apparatus above performing partition-based coding.
[0401] One embodiment includes a bitstream or signal that includes one or more syntax elements for performing the above functions or variations thereof.
[0402] One embodiment includes a bitstream or signal that includes syntax for conveying information generated according to any of the described embodiments.
[0403] One embodiment includes creating and / or transmitting and / or receiving and / or decoding according to any of the described embodiments.
[0404] One embodiment includes a method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the described embodiments.
[0405] One embodiment includes inserting a syntax element in a signaling that enables a decoder to determine decoded information in a manner corresponding to the way used by an encoder.
[0406] One embodiment includes creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof.
[0407] One embodiment includes a TV, set-top box, mobile phone, tablet, or other electronic device that performs a transform method according to any of the described embodiments.
[0408] One embodiment includes a TV, set-top box, mobile phone, tablet, or other electronic device that determines and displays (e.g., using a monitor, screen, or other type of display) a resultant image by performing a transform method according to any of the described embodiments.
[0409] One embodiment includes a TV, set-top box, mobile phone, tablet, or other electronic device that selects, band-limits, or tunes a channel (e.g., using a tuner) to receive a signal including an encoded image and performs a transform method according to any of the described embodiments.
[0410] One embodiment includes a TV, set-top box, mobile phone, tablet, or other electronic device that receives (e.g., using an antenna) a signal including an encoded image over the air and performs a transform method.
[0411] Note that various hardware elements of one or more of the embodiments are referred to as "modules" which implement (i.e., perform, execute, etc.) the various functions described herein in connection with the corresponding modules. As used herein, a module includes hardware that a person of ordinary skill in the relevant art deems suitable for a given implementation (e.g., one or more processors, one or more microprocessors, one or more microcontrollers, one or more microchips, one or more application specific integrated circuits (ASICs), one or more field programmable gate arrays (FPGAs), one or more memory devices). Each described module may also include executable instructions for performing one or more functions described as being performed by the corresponding module, and note that these instructions may take the form of hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc. or include hardware (i.e., hardwired) instructions, firmware instructions, software instructions, etc. and may be stored in any suitable non-transitory computer-readable medium such as those commonly referred to as RAM, ROM, etc.
[0412] Although the features and elements have been described above in particular combinations, one of ordinary skill in the art will understand that each feature or element can be used alone or in any combination with other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware that is incorporated into a computer-readable medium for execution by a computer or processor. Examples of computer-readable storage media include, but are not limited to, read only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for a WTRU, UE, terminal, base station, RNC, or any host computer.
Claims
1. A method, comprising: Accessing an input mesh, wherein the input mesh includes a list of faces and a plurality of vertex positions; Generating at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; Generating a base mesh and at least two base mesh face features on the base mesh, wherein the base mesh includes vertex positions and information indicating the connectivity of the base mesh; Generating a fixed-length codeword from the at least two base mesh face features; Accessing a predefined template mesh; Generating information indicating matching vertices between the predefined template mesh and the base mesh; and Outputting the fixed-length codeword and the information indicating the connectivity of the base mesh.
2. The method according to claim 1, wherein, The input mesh is a semi-regular mesh.
3. The method according to claim 1, wherein, Generating the base mesh includes: Generating the vertex positions; and Generating the information indicating the connectivity of the base mesh.
4. The method according to claim 1, wherein, Generating the at least two base mesh face features on the base mesh is performed by performing a learning-based aggregation on the at least two initial mesh face features.
5. The method according to claim 1, wherein Generating the fixed-length codeword is performed by pooling the at least two base mesh face features.
6. The method according to claim 1, wherein The predefined template mesh is a mesh corresponding to a unit sphere.
7. The method according to claim 1, wherein The information indicating the base connectivity includes a list of triangles with information having an index corresponding to the matching vertices indicated by a set of matching indices.
8. The method according to claim 1, Among them, wherein generating the base mesh and the at least two base mesh face features on the base mesh is performed by a learning-based heterogeneous mesh encoder, and wherein the heterogeneous mesh encoder includes at least one downsampling face convolutional layer.
9. The method according to claim 1, wherein, Generating the fixed-length codeword from the at least two base mesh face features includes using a learning-based AdaptMaxPool process.
10. The method according to claim 1, wherein, Generating the set of matching indices is performed by a learning-based SphereNet process.
11. The method according to claim 1, the method further comprising: Outputting information indicating matching vertices, wherein the information indicating matching vertices includes a set of matching indices indicating the matching vertices between the predefined template mesh and the base mesh.
12. An apparatus, comprising: A processor; A memory storing instructions that, when executed by the processor, are operable to cause a mesh encoder to: Access an input mesh, wherein the input mesh includes a list of faces and a plurality of vertex positions; Generate at least two initial mesh face features for at least one face listed in the list of faces of the input mesh; Generate a base mesh and at least two base mesh face features on the base mesh, wherein the base mesh includes vertex positions and information indicating the connectivity of the base mesh; Generate a fixed-length codeword from the at least two base mesh face features; Access a predefined template mesh; Generate information indicating the matching of vertices between the predefined template mesh and the base mesh; and Output the fixed-length codeword and the information indicating the connectivity of the base mesh.
13. A method, comprising: Receive a fixed-length codeword, information indicating the connectivity of the base mesh, and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; Generate a reconstructed base mesh and at least two base surface features; And Generate at least one reconstructed mesh for at least two hierarchical resolutions.
14. The method according to claim 13, wherein Generate the at least one reconstructed mesh Generate K reconstructed meshes for K hierarchical resolutions.
15. The method according to claim 14, wherein Use a heterogeneous mesh decoder to generate K reconstructed meshes.
16. The method according to claim 15, wherein, The heterogeneous mesh decoder performs at least one upsampling surface convolution process and at least one Face2Node process.
17. The method according to claim 13, wherein, Generate the at least one reconstructed mesh Generate at least two reconstructed meshes for at least two corresponding hierarchical resolutions.
18. The method according to claim 13, wherein, Perform generating the reconstructed base mesh through a learning-based DeSphereNet process.
19. The method according to claim 13, wherein, Generating the at least one reconstructed mesh for at least two hierarchical resolutions includes: Determine input surface features from the base surface feature map; Generate updated surface features corresponding to the input surface features; Determine the updated differential positions of one or more nodes of the reconstructed mesh; and Use the corresponding updated differential positions to update the positions of one or more nodes of the reconstructed base mesh.
20. An apparatus, comprising: A processor; A memory storing instructions that, when executed by the processor, can operate to cause a mesh decoder to: Receive a fixed-length codeword, information indicating the connectivity of the base mesh, and a predefined mesh to generate a reconstructed base mesh and at least two base surface features; Generate a reconstructed base mesh and at least two base surface features; And Generate at least one reconstructed mesh for at least two hierarchical resolutions.
21. A mesh encoding method, comprising: Access a semi-regular input mesh to generate initial mesh surface features for each mesh face, wherein the semi-regular input mesh includes a face list and a plurality of vertex positions; Generate a base mesh including vertex positions and information indicating base connectivity and a set of surface features on the base mesh through a learning-based feature aggregation module; Generate a fixed-length codeword based on the base surface features using a feature pooling module; Access a predefined template mesh and the base mesh to generate information indicating matching vertices between the predefined template mesh and the base mesh; and Output the generated fixed-length codeword and the information indicating the base connectivity.
22. An encoding method, comprising: Access an input remeshed mesh to generate initial mesh surface features, wherein the input remeshed mesh includes a face list and vertex positions; Generate a base mesh and an atlas of surface features on the base mesh; Generate a fixed-length codeword from the base surface features; Access a predefined number of vertices and a predefined sphere mesh of base mesh vertices to generate a match between the sphere mesh vertices and the base network vertices; and Output the generated fixed-length codeword and the base mesh connectivity information.
23. A mesh encoder, comprising: A processor; A memory storing instructions that, when executed by the processor, can operate to cause the mesh encoder to: Access a semi-regular input mesh to generate initial mesh surface features for each mesh face, Among them, the semi-regular input grid includes a face list and multiple vertex positions; Generate a base grid including vertex positions and information indicating basic connectivity, and a set of face features on the base grid through a learning-based feature aggregation module; Use a feature pooling module to generate a fixed-length codeword based on the base face features; Access a predefined template grid and the base grid to generate information indicating matching vertices between the predefined template grid and the base grid; and Output the generated fixed-length codeword and the information indicating the basic connectivity.
24. An encoding device, comprising: A processor; And A memory storing instructions, which when executed by the processor can operate to cause the processor to: Access an input remeshed grid to generate initial grid face features, wherein the input remeshed grid includes a face list and vertex positions; Generate a base grid and an atlas of face features on the base grid; Generate a fixed-length codeword from the base face features; Access a predefined number of vertices and a predefined spherical grid of base grid vertices to generate a match between the spherical grid vertices and the base network vertices; and Output the generated fixed-length codeword and the base grid connectivity information.
25. A grid decoding method, comprising: Access the basic connectivity information and a predefined spherical grid to generate a reconstructed base grid and a base face feature map through a learning-based module DeSphereNet; And Generate K reconstructed grids at K hierarchical resolutions through a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
26. A decoding method, comprising: Access the basic grid connectivity information, the fixed-length codeword, and a predefined spherical grid to generate a reconstructed base grid and a base face feature map; And Generate K reconstructed grids at K hierarchical resolutions.
27. A grid decoding device, comprising: A processor; A memory storing instructions, which when executed by the processor can operate to cause the grid decoder to: Access the basic connectivity information and a predefined spherical grid to generate a reconstructed base grid and a base face feature map through a learning-based module DeSphereNet; And Generate K reconstructed grids at K hierarchical resolutions through a learning-based module HetMeshDec composed of a series of K pairs of UpFaceConv and Face2Node modules.
28. A decoding device, comprising: A processor; And A memory storing instructions, which when executed by the processor can operate to cause the processor to: Access the basic grid connectivity information, the fixed-length codeword, and a predefined spherical grid to generate a reconstructed base grid and a base face feature map; and Generate K reconstructed grids at K hierarchical resolutions.
29. A grid decoder configured to obtain a fixed-length codeword, base connectivity information, and a set of sphere matching indices and generate a reconstructed grid, wherein, The grid decoder is configured to: Access the basic connectivity information and a predefined spherical grid to generate a reconstructed base grid and a base face feature map through a learning-based module DeSphereNet; and Generate K reconstructed meshes at K hierarchical resolutions through a learning-based module HetMeshDec consisting of a series of K pairs of UpFaceConv and Face2Node modules.
30. A method, comprising: Determine initial mesh face features from an input mesh; Determine a base mesh including a set of face features based on a first learning-based module including a series of mesh feature extraction layers; Generate a fixed-length codeword from the base mesh on the mesh faces using a second learning-based pooling module; And Generate a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
31. An apparatus comprising a memory and a processor, the processor configured to execute: Determine initial mesh face features from an input mesh; Determine a base mesh including a set of face features based on a first learning-based module including a series of mesh feature extraction layers; Generate a fixed-length codeword from the base mesh on the mesh faces using a second learning-based pooling module; And Generate a base graph by matching vertices of a predefined template mesh with vertices of the base mesh using a third module.
32. A method, comprising: In the presence of a predefined template mesh, use a fixed codeword and a base graph via a first learning-based module to determine a reconstructed base mesh and a base face feature map; And Generate at least one reconstructed mesh at multiple hierarchical resolutions through a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
33. An apparatus comprising a memory and a processor, the processor configured to execute: In the presence of a predefined template mesh, use a fixed codeword and a base graph via a first learning-based module to determine a reconstructed base mesh and a base face feature map; and Generate at least one reconstructed mesh at multiple hierarchical resolutions through a second learning-based module including a series of layers, the series of layers including a series of mesh feature extraction and node generation layers.
34. An apparatus, comprising: A heterogeneous mesh encoder including a series of layers including paired mesh feature extraction modules and mesh downsampling modules; And A heterogeneous mesh decoder including a learning-based module including a series of layers including paired mesh node generation modules and mesh upsampling modules.
35. The apparatus according to claim 34, wherein, Transmit a base mesh from the heterogeneous mesh encoder to the heterogeneous mesh decoder.
36. The device according to claim 35, wherein, Use multiple input features in addition to the directly consumed mesh.
37. The method according to claim 32, the method further comprising an upsampling module based on loop subdivision, wherein, The loop subdivision includes: Construct an enhanced node-specific face feature set; Use a shared module to update the enhanced node-specific face feature set; Average the updated node-specific face features; and perform neighborhood averaging on node positions.
38. The method according to any one of claims 30, 32, or 37, the method further comprising: Convert the codeword into a face-specific codeword set; and transform the face-specific codeword into base mesh features and geometry.
39. The method according to any one of claims 30, 32, or 37, the method further comprising: Convert the original grid into partitions; Shift the origin of the partitions; And encode or decode each partitioned grid respectively.
40. The method according to any one of claims 30, 32 or 37, wherein The grids have different sizes and connectivities.
41. A non-transitory computer-readable medium comprising data content according to the method of claim 30 or generated by the apparatus of claim 31 for playback using a processor.
42. A signal comprising video data according to the method of claim 30 or generated by the apparatus of claim 31 for playback using a processor.
43. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to perform the method according to claim 30 or 32.
44. A non-transitory computer-readable medium comprising data content, the data content including instructions for performing the method according to any one of claims 30, 32 and 37 to 39.
45. The method according to claim 30, wherein, The third module is a learning-based module.
46. The method according to claim 30, wherein The third module is a traditional non-learning-based module.
47. A method comprising: Access an input grid, wherein the input grid includes a list of faces and a plurality of vertex positions; Generate at least two initial grid face features for at least one face listed in the list of faces of the input grid; Generate a base grid and at least two base grid face features on the base grid, wherein the base grid includes vertex positions and information indicating the connectivity of the base grid; Access a predefined template grid; Generate information indicating matching vertices between the predefined template grid and the base grid; and Output information indicating the connectivity of the base grid.
48. An apparatus comprising: A processor; And A memory storing instructions which, when executed by the processor, are operable to cause the processor to: Access an input grid, wherein the input grid includes a list of faces and a plurality of vertex positions; Generate at least two initial grid face features for at least one face listed in the list of faces of the input grid; Generate a base grid and at least two base grid face features on the base grid, wherein the base grid includes vertex positions and information indicating the connectivity of the base grid; Access a predefined template grid; Generate information indicating matching vertices between the predefined template grid and the base grid; and Output information indicating the connectivity of the base grid.
49. A method comprising: Receive information indicating the connectivity of a base grid and a predefined grid to generate a reconstructed base grid and at least two base face features; Generate a reconstructed base grid and at least two base face features; And Generate at least one reconstructed grid for at least two hierarchical resolutions.
50. An apparatus comprising: A processor; And A memory storing instructions which, when executed by the processor, are operable to cause the processor to: Receive information indicating the connectivity of a base grid and a predefined grid to generate a reconstructed base grid and at least two base face features; Generate a reconstructed base grid and at least two base face features; And Generate at least one reconstructed mesh for at least two hierarchical resolutions.
51. A method, comprising: Access an input mesh, wherein the input mesh includes a list of faces and a plurality of vertex positions; Perform a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed on the list of faces of the input mesh; Perform an AdaptMaxPool process such that: Generate a base mesh and at least two base mesh face features on the base mesh, wherein the base mesh includes vertex positions and information indicating base mesh connectivity; and Generate a fixed-length codeword from the at least two base mesh face features; Output the fixed-length codeword and the information indicating the base mesh connectivity.
52. An apparatus, comprising: A processor; And A memory storing instructions that, when executed by the processor, are operable to cause the processor to: Access an input mesh, wherein the input mesh includes a list of faces and a plurality of vertex positions; Perform a heterogeneous mesh encoder process to generate at least two initial mesh face features for at least one face listed on the list of faces of the input mesh; Perform an AdaptMaxPool process such that: Generate a base mesh and at least two base mesh face features on the base mesh, wherein the base mesh includes vertex positions and information indicating base mesh connectivity; and Generate a fixed-length codeword from the at least two base mesh face features; Output the fixed-length codeword and the information indicating the base mesh connectivity.
53. A method, comprising: Receive a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; Perform a base mesh reconstruction graph neural network (BaseConGNN) process to generate a reconstructed base mesh and at least two base face features; And Perform a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
54. An apparatus, comprising: A processor; And A memory storing instructions that, when executed by the processor, are operable to cause the processor to: Receive a fixed-length codeword, information indicating base mesh connectivity, and a predefined mesh to generate a reconstructed base mesh and at least two base face features; Perform a base mesh reconstruction graph neural network (BaseConGNN) process to generate a reconstructed base mesh and at least two base face features; And Perform a heterogeneous mesh decoder process to generate at least one reconstructed mesh for at least two hierarchical resolutions.
55. A method, comprising: Access an input mesh; Partition the input mesh into a first input mesh and a second input mesh, wherein the first input mesh includes a first list of faces and a first plurality of vertex positions, and wherein the second input mesh includes a second list of faces and a second plurality of vertex positions; Generate at least two first initial mesh face features for at least one first face listed on the first list of faces of the first input mesh; Generate a first base mesh and at least two first base mesh surface features on the first base mesh, wherein the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; Generate a first fixed-length codeword from the at least two first base mesh surface features; Access a first predefined template mesh; Output the first fixed-length codeword and the first information indicating the connectivity of the first base mesh; Generate at least two second initial mesh surface features for at least one second surface listed in the second surface list of the second input mesh; Generate a second base mesh and at least two second base mesh surface features on the second base mesh, wherein the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; Generate a second fixed-length codeword from the at least two second base mesh surface features; Access a second predefined template mesh; and Output the second fixed-length codeword and the second information indicating the connectivity of the second base mesh.
56. The method according to claim 55, the method further comprising: Generate a first matching index set, wherein the first matching index set indicates first matching vertices between the first predefined template mesh and the first base mesh; Output the first matching index set; Generate a second matching index set, wherein the second matching index set indicates second matching vertices between the second predefined template mesh and the second base mesh; and Output the second matching index set.
57. An apparatus, comprising: A processor; And A memory storing instructions that, when executed by the processor, are operable to cause the processor to: Access an input mesh; Partition the input mesh into a first input mesh and a second input mesh, wherein the first input mesh includes a first surface list and a first plurality of vertex positions, and wherein the second input mesh includes a second surface list and a second plurality of vertex positions; Generate at least two first initial mesh surface features for at least one first surface listed in the first surface list of the first input mesh; Generate a first base mesh and at least two first base mesh surface features on the first base mesh, wherein the first base mesh includes first vertex positions and first information indicating the connectivity of the first base mesh; Generate a first fixed-length codeword from the at least two first base mesh surface features; Access a first predefined template mesh; Output the first fixed-length codeword and the first information indicating the connectivity of the first base mesh; Generate at least two second initial mesh surface features for at least one second surface listed in the second surface list of the second input mesh; Generate a second base mesh and at least two second base mesh surface features on the second base mesh, wherein the second base mesh includes second vertex positions and second information indicating the connectivity of the first base mesh; Generate a second fixed-length codeword from the at least two second base mesh surface features; Access a second predefined template mesh; and Output the second fixed-length codeword and the second information indicating the second base grid connectivity.
58. An apparatus comprising at least one processor configured to perform the method according to any one of claims 1 to 11, 13 to 19, 21, 22, 25, 26, 30, 32, 37 to 40, 45 to 47, 49, 51, 53, 55, and 56.
59. An apparatus comprising a computer-readable medium storing instructions for causing one or more processors to perform the method according to any one of claims 1 to 11, 13 to 19, 21, 22, 25, 26, 30, 32, 37 to 40, 45 to 47, 49, 51, 53, 55, and 56.
60. An apparatus comprising at least one processor and at least one non-transitory computer-readable medium storing instructions for causing the at least one processor to perform the method according to any one of claims 1 to 11, 13 to 19, 21, 22, 25, 26, 30, 32, 37 to 40, 45 to 47, 49, 51, 53, 55, and 56.
61. A signal comprising a bitstream generated according to any one of claims 1 to 11, 13 to 19, 21, 22, 25, 26, 30, 32, 37 to 40, 45 to 47, 49, 51, 53, 55, and 56.