Directional alignment of normal-based 3D mesh subdivision
The directional alignment of normal-based 3D mesh subdivision method addresses the challenge of efficiently compressing and transmitting 3D meshes by encoding geometry and attributes separately, ensuring high-quality visual data transmission for various applications.
Patent Information
- Application Number
- PCT/US2025/038143
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-30
- Filing Date
- 2025-07-17
- Publication Date
- 2026-01-22
AI Technical Summary
Existing methods for encoding and decoding volumetric visual data, particularly 3D meshes, face challenges in efficiently compressing and transmitting large data sizes while maintaining visual quality, especially in applications like Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR), where lossy compression compromises user experience and lossless compression is required for medical or geological applications.
A directional alignment of normal-based 3D mesh subdivision method that involves encoding geometry and attribute information separately, using displacement vectors and wavelet coefficients, and applying subdivision schemes to reduce data size effectively.
This approach achieves efficient compression and transmission of 3D meshes, preserving visual quality and enabling adaptive streaming, suitable for diverse applications including AR, VR, and MR, while meeting the needs of lossless compression in specific domains.
Smart Images

Figure US2025038143_22012026_PF_FP_ABST
Abstract
Description
Directional Alignment of Normal-based 3D Mesh Subdivision CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application Nos. 63 / 672,670, filed July 17, 2024, 63 / 712,942, filed October 28, 2024, and 63 / 714,096, filed October 30, 2024, all of which are hereby incorporated by reference in their entireties.BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Examples of several of the various embodiments of the present disclosure are described herein with reference to the drawings.
[0003] FIG. 1 illustrates an exemplary mesh coding / decoding system in which embodiments of the present disclosure may be implemented.
[0004] FIG. 2A illustrates a block diagram of an example encoder for intra encoding a 3D mesh, according to some embodiments.
[0005] FIG. 2B illustrates a block diagram of an example encoder for inter encoding a 3D mesh, according to some embodiments.
[0006] FIG. 3 illustrates a diagram showing an example decoder.
[0007] FIG. 4 is a diagram showing an example process for generating displacements of an input mesh (e.g., an input 3D mesh frame) to be encoded, according to some embodiments.
[0008] FIG. 5 illustrates an example process for approximating and encoding a geometry of a 3D mesh, according to some embodiments.
[0009] FIG. 6 illustrates an example of vertices of a subdivided mesh (e.g., a subdivided base mesh) corresponding to multiple levels of detail (LODs), according to some embodiments.
[0010] FIG. 7A illustrates an example of an image packed with displacements (e.g., displacement fields or vectors) using a packing method, according to some embodiments.
[0011] FIG. 7B illustrates an example of the displacement image with labeled LODs, according to some embodiments.
[0012] FIG. 8 illustrates an example of a lifting scheme for representing displacement information of a 3D mesh as wavelet coefficients, according to some embodiments.
[0013] FIG. 9A illustrates a diagram of an example normal-based subdivision scheme, according to some embodiments.
[0014] FIG. 9B illustrates a diagram of an example normal-based subdivision scheme, according to some embodiments.
[0015] FIG. 9C illustrates a diagram of an example normal-based subdivision scheme with a reference vector, according to some embodiments.
[0016] FIG. 9D illustrates a diagram of an example normal-based subdivision scheme with vertex adjustments, according to some embodiments.
[0017] FIGS. 10A-C illustrate diagrams showing different combinations of: signs of a first and second weight corresponding to the first and second vertex normals of a first and second of an edge, and a same directional alignment between the first and second vertex normals, according to some embodiments.
[0018] FIGS. 11 A-C illustrate diagrams showing different combinations of: signs of a first and second weight corresponding to the first and second vertex normals of a first and second of an edge, and an opposite directional alignment between the first and second vertex normals, according to some embodiments.
[0019] FIG. 12 illustrates a diagram showing two example adjustments to one or both of the first and second weights of vertex normals for the scenario shown in FIG. 11 B.
[0020] FIG. 13 illustrates a flowchart of a method for performing normal-based subdivision for decoding (e.g., reconstructing) a 3D mesh, according to some embodiments.
[0021] FIG. 14 illustrates a flowchart of a method for performing normal-based subdivision for encoding a 3D mesh, according to some embodiments.
[0022] FIG. 15 illustrates a flowchart of a method for performing normal-based subdivision for 3D mesh coding, according to some embodiments.
[0023] FIG. 16 illustrates a block diagram of an exemplary computer system in which embodiments of the present disclosure may be implemented.DETAILED DESCRIPTION
[0024] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the disclosure. However, it will be apparent to those skilled in the art that the disclosure, including structures, systems, and methods, may be practiced without these specific details. The description and representation herein are the common means used by those experienced or skilled in the art to most effectively convey the substance of their work to others skilled in the art. In other instances, well-known methods, procedures, components, and circuitry have not been described in detail to avoid unnecessarily obscuring aspects of the disclosure.
[0025] References in the specification to “one embodiment,” “an embodiment,” “an example embodiment,” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is submitted that it is within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0026] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0027] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0028] Furthermore, embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented in software, firmware, middleware or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine-readable medium. A processor(s) may perform the necessary tasks.
[0029] Traditional visual data describes an object or scene using a series of pixels that each comprise a position in two dimensions (x and y) and one or more optional attributes like color. Volumetric visual data adds another positional dimension to this traditional visual data. Volumetric visual data describes an object or scene using a series of points that each comprise a position in three dimensions (x, y, and z) and one or more optional attributes like color. Compared to traditional visual data, volumetric visual data may provide a more immersive way to experience visual data. For example, an object or scene described by volumetric visual data may be viewed from any (or multiple) angles, whereas traditional visual data may generally only be viewed from the angle in which it was captured or rendered. Volumetric visual datamay be used in many applications, including Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR). Volumetric visual data may be in the form of a volumetric frame that describes an object or scene captured at a particular time instance or in the form of a sequence of volumetric frames (referred to as a volumetric sequence or volumetric video) that describes an object or scene captured at multiple different time instances.
[0030] One format for storing volumetric visual data is three dimensional (3D) meshes (hereinafter referred to as a mesh or a mesh frame). A mesh frame (or mesh) comprises a collection of points in three-dimensional (3D) space, also referred to as vertices. Each vertex in a mesh comprises geometry information that indicates the vertex’s position in 3D space. For example, the geometry information may indicate the vertex’s position in 3D space using three Cartesian coordinates (x, y, and z). Further the mesh may comprise geometry information indicating a plurality of triangles. Each triangle comprises three vertices connected by three edges and a face. One or more types of attribute information may be stored for each face (of a triangle). Attribute information may indicate a property of a face's visual appearance. For example, attribute information may indicate a texture (e.g., color) of the face, a material type of the face, transparency information of the face, reflectance information of the face, a normal vector to a surface of the face, a velocity at the face, an acceleration at the face, a time stamp indicating when the face (and / or vertex) was captured, or a modality indicating how the face (and / or vertex) was captured (e.g., running, walking, or flying). In another example, a face (or vertex) may comprise light field data in the form of multiple view-dependent texture information. Light field data may be another type of optional attribute information.
[0031] The triangles (e.g., represented by vertexes and edges) in a mesh may describe an object or a scene. For example, the triangles in a mesh may describe the external surface and / or the internal structure of an object or scene. The object or scene may be synthetically generated by a computer or may be generated from the capture of a real-world object or scene. The geometry information of a real world object or scene may be obtained by 3D scanning and / or photogrammetry. 3D scanning may include laser scanning, structured light scanning, and / or modulated light scanning. 3D scanning may obtain geometry information by moving one or more laser heads, structured light cameras, and / or modulated light cameras relative to an object or scene being scanned. Photogrammetry may obtain geometry information by triangulating the same feature or point in different spatially shifted 2D photographs. Mesh data may be in the form of a mesh frame that describes an object or scene captured at a particular time instance or in the form of a sequence of mesh frames (referred to as a mesh sequence or mesh video) that describes an object or scene captured at multiple different time instances.
[0032] The data size of a mesh frame or sequence in addition with one or more types of attribute information may be too large for storage and / or transmission in many applications. For example, a single mesh frame may comprise thousands or tens or hundreds of thousands of triangles, where each triangle(e.g vertexes and / or edges) comprises geometry information and one or more optional types of attribute information. The geometry information of each vertex may comprise three Cartesian coordinates (x, y, and z) that are each represented, for example, using 8 bits or 24 bits in total. The attribute information of each point may comprise a texture corresponding to three color components (e.g., R, G, and B color components) that are each represented, for example, using 8 bits or 24 bits in total. A single vertex therefore comprises 48 bits of information in this example, with 24 bits of geometry information and 24 bits of texture. Encoding may be used to compress the size of a mesh frame or sequence to provide for more efficient storage and / or transmission. Decoding may be used to decompress a compressed mesh frame or sequence for display and / or other forms of consumption (e.g., by a machine learning based device, neural network based device, artificial intelligence based device, or other forms of consumption by other types of machine based processing algorithms and / or devices).
[0033] Compression of meshes may be lossy (e.g., introducing differences relative to the original data) for the distribution to and visualization by an end-user, for example on AR / VR glasses or any other 3D- capable device. Lossy compression allows for a very high ratio of compression but incurs a trade-off between compression and visual quality perceived by the end-user. Other frameworks, like medical or geological applications, may require lossless compression to avoid altering the decompressed meshes.
[0034] Volumetric visual data may be stored after being encoded into a bitstream in a container, for example, a file server in the network. The end-user may request for a specific bitstream depending on the user’s requirement. The user may also request for adaptive streaming of the bitstream where the trade-off between network resource consumption and visual quality perceived by the end-user is taken into consideration by an algorithm.
[0035] FIG. 1 illustrates an exemplary mesh coding / decoding system 100 in which embodiments of the present disclosure may be implemented. Mesh coding / decoding system 100 comprises a source device 102, a transmission medium 104, and a destination device 106. Source device 102 encodes a mesh sequence 108 into a bitstream 110 for more efficient storage and / or transmission. Source device 102 may store and / or transmit bitstream 110 to destination device 106 via transmission medium 104. Destination device 106 decodes bitstream 110 to display mesh sequence 108 or for other forms of consumption. Destination device 106 may receive bitstream 110 from source device 102 via a storage medium or transmission medium 104 Source device 102 and destination device 106 may be any one of a number of different devices, including a cluster of interconnected computer systems acting as a pool of seamless resources (also referred to as a cloud of computers or cloud computer), a server, a desktop computer, a laptop computer, a tablet computer, a smart phone, a wearable device, a television, a camera, a video gaming console, a set-top box, a video streaming device, an autonomous vehicle, or a head mounted display. A head mounted display may allow a user to view a VR, AR, or MR scene and adjust the view of the scene based on movement of the user's head. A head mounted display may be tethered to aprocessing device (e.g., a server, desktop computer, set-top box, or video gaming counsel) or may be fully self-contained.
[0036] To encode mesh sequence 108 into bitstream 110, source device 102 may comprise a mesh source 112, an encoder 114, and an output interface 116. Mesh source 112 may provide or generate mesh sequence 108 from a capture of a natural scene and / or a synthetically generated scene. A synthetically generated scene may be a scene comprising computer generated graphics. Mesh source 112 may comprise one or more mesh capture devices (e.g., one or more laser scanning devices, structured light scanning devices, modulated light scanning devices, and / or passive scanning devices), a mesh archive comprising previously captured natural scenes and / or synthetically generated scenes, a mesh feed interface to receive captured natural scenes and / or synthetically generated scenes from a mesh content provider, and / or a processor to generate synthetic mesh scenes.
[0037] As shown in FIG. 1 , a mesh sequence 108 may comprise a series of mesh frames 124. A mesh frame describes an object or scene captured at a particular time instance. Mesh sequence 108 may achieve the impression of motion when a constant or variable time is used to successively present mesh frames 124 of mesh sequence 108. A (3D) mesh frame comprises a collection of vertices 126 in 3D space and geometry information of vertices 126. A 3D mesh may comprise a collection of vertices, edges, and faces that define the shape of a polyhedral object. Further, the mesh frame comprises a plurality of triangles (e.g., polygon triangles). For example, a triangle may include vertices 134A-C and edges 136A- C and a face 132. The faces usually consist of triangles (triangle mesh), Quadrilaterals (Quads), or other simple convex polygons (n-gons), since this simplifies rendering, but may also be more generally composed of concave polygons, or even polygons with holes. Each of vertices 126 may comprise geometry information that indicates the point's position in 3D space. For example, the geometry information may indicate the point’s position in 3D space using three Cartesian coordinates (x, y, and z). For example, the geometry information may indicated the plurality of triangles with each comprising three vertices of vertices 126. One or more of the triangles may further comprise one or more types of attribute information. Attribute information may indicate a property of a point's visual appearance. For example, attribute information may indicate a texture (e.g., color) of a face, a material type of a face, transparency information of a face, reflectance information of a face, a normal vector to a surface of a face, a velocity at a face, an acceleration at a face, a time stamp indicating when a face was captured, a modality indicating when a face was captured (e.g., running, walking, or flying). In another example, one or more of the faces (or triangles) may comprise light field data in the form of multiple view-dependent texture information. Light field data may be another type of optional attribute information. Color attribute information of one or more of the faces may comprise a luminance value and two chrominance values. The luminance value may represent the brightness (or luma component, Y) of the point. The chrominance values may respectively represent the blue and red components of the point (or chroma components, Cb and Cr)separate from the brightness. Other color attribute values are possible based on different color schemes (e.g., an RGB or monochrome color scheme).
[0038] In some embodiments, a 3D mesh (e.g., one of mesh frames 124) may be a static or a dynamic mesh. In some examples, the 3D mesh may be represented (e.g., defined) by connectivity information, geometry information, and texture information (e.g., texture coordinates and texture connectivity). In some embodiments, the geometry information may represent locations of vertices of the 3D mesh in 3D space and the connectivity information may indicate how the vertices are to be connected together to form polygons (e.g., triangles) that make up the 3D mesh. Also, the texture coordinates indicate locations of pixels in a 2D image that correspond to vertices of a corresponding 3D mesh (or a sub-mesh of the 3D mesh). In some examples, patch information may indicate how the texture coordinates defined with respect to a 2D bounding box map into a 3D space of a 3D bounding box associated with the patch based on how the points were projected onto a projection plane for the patch. Also, the texture connectivity information may indicate how the vertices represented by the texture coordinates are to be connected together to form polygons of the 3D mesh (or sub-meshes). For example, each texture or attribute patch of the texture image may corresponds to a corresponding sub-mesh defined using texture coordinates and texture connectivity.
[0039] In some embodiments, for each 3D mesh, one or multiple 2D images may represent the textures or attributes associated with the mesh. For example, the texture information may include geometry information listed as X, Y, and Z coordinates of vertices and texture coordinates listed as 2D dimensional coordinates corresponding to the vertices. The example texture mesh may include texture connectivity information that indicates mappings between the geometry coordinates and texture coordinates to form polygons, such as triangles. For example, a first triangle may be formed by three vertices, where a first vertex is defined as the first geometry coordinate (e.g. 64.062500, 1237.739990, 51 .757801), which corresponds with the first texture coordinate (e.g. 0.0897381 , 0.740830). A second vertex of the triangle may be defined as the second geometry coordinate (e.g. 59.570301 , 1236.819946, 54.899700), which corresponds with the second texture coordinate (e.g. 0.899059, 0.741542). Finally, a third vertex of the triangle may correspond to the third listed geometry coordinate which matches with the third listed texture coordinate. However, note that in some instances a vertex of a polygon, such as a triangle may map to a set of geometry coordinates and texture coordinates that may have different index positions in the respective lists of geometry coordinates and texture coordinates. For example, the second triangle has a first vertex corresponding to the fourth listed set of geometry coordinates and the seventh listed set of texture coordinates. A second vertex corresponding to the first listed set of geometry coordinates and the first set of listed texture coordinates and a third vertex corresponding to the third listed set of geometry coordinates and the ninth listed set of texture coordinates.
[0040] Encoder 114 may encode mesh sequence 108 into bitstream 110. To encode mesh sequence 108, encoder 114 may apply one or more prediction techniques to reduce redundant information in mesh sequence 108. Redundant information is information that may be predicted at a decoder and therefore may not be needed to be transmitted to the decoder for accurate decoding of mesh sequence 108. For example, encoder 114 may convert attribute information (e.g., texture information) of one or more of mesh frames 124 from 3D to 2D and then apply one or more 2D video encoders or encoding methods to the 2D images. For example, any one of multiple different proprietary or standardized 2D video encoders / decoders may be used, including International Telecommunications Union Telecommunication Standardization Sector (ITU-T) H.1263, ITU-T H.1264 and Moving Picture Expert Group (MPEG)-4 Visual (also known as Advanced Video Coding (AVC)), ITU-T H.1265 and MPEG-H Part 2 (also known as High Efficiency Video Coding (HEVC), ITU-T H.1265 and MPEG-I Part 3 (also known as Versatile Video Coding (WC)), the WebM VP8 and VP9 codecs, and AOMedia Video 1 (AV1). Encoder 114 may encode geometry of mesh sequence 108 based on video dynamic mesh coding (V-DMC). V-DMC specifies the encoded bitstream syntax and semantics for transmission or storage of a mesh sequence and the decoder operation for reconstructing the mesh sequence from the bitstream.
[0041] Output interface 116 may be configured to write and / or store bitstream 110 onto transmission medium 104 for transmission to destination device 106. In addition or alternatively, output interface 116 may be configured to transmit, upload, and / or stream bitstream 110 to destination device 106 via transmission medium 104. Output interface 116 may comprise a wired and / or wireless transmitter configured to transmit, upload, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as Digital Video Broadcasting (DVB) standards, Advanced Television Systems Committee (ATSC) standards, Integrated Services Digital Broadcasting (ISDB) standards, Data Over Cable Service Interface Specification (DOCSIS) standards, 3rd Generation Partnership Project (3GPP) standards, Institute of Electrical and Electronics Engineers (IEEE) standards, Internet Protocol (IP) standards, and Wireless Application Protocol (WAP) standards.
[0042] Transmission medium 104 may comprise a wireless, wired, and / or computer readable medium. For example, transmission medium 104 may comprise one or more wires, cables, air interfaces, optical discs, flash memory, and / or magnetic memory. In addition or alternatively, transmission medium 104 may comprise one or more networks (e g., the Internet) or file servers configured to store and / or transmit encoded video data.
[0043] To decode bitstream 110 into mesh sequence 108 for display or other forms of consumption, destination device 106 may comprise an input interface 118, a decoder 120, and a mesh display 122. Input interface 118 may be configured to read bitstream 110 stored on transmission medium 104 by source device 102. In addition or alternatively, input interface 118 may be configured to receive, download, and / or stream bitstream 110 from source device 102 via transmission medium 104. Inputinterface 118 may comprise a wired and / or wireless receiver configured to receive, download, and / or stream bitstream 110 according to one or more proprietary and / or standardized communication protocols, such as those mentioned above.
[0044] Decoder 120 may decode mesh sequence 108 from encoded bitstream 110. To decode attribute information (e.g., textures) of mesh sequence 108, decoder 120 may reconstruct the 2D images compressed using one or more 2D video encoders. Decoder 120 may then reconstruct the attribute information of 3D mesh frames 124 from the reconstructed 2D images. In some examples, decoder 120 may decode a mesh sequence that approximates mesh sequence 108 due to, for example, lossy compression of mesh sequence 108 by encoder 114 and / or errors introduced into encoded bitstream 110 during transmission to destination device 106. Further, decoder 120 may decode geometry of mesh sequence 108 from encoded bitstream 110, as will be further described below. Then, one or more of decoded attribute information may be applied to decoded mesh frames of mesh sequence 108.
[0045] Mesh display 122 may display mesh sequence 108 to a user. Mesh display 122 may comprise a cathode rate tube (CRT) display, a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, a 3D display, a holographic display, a head mounted display, or any other display device suitable for displaying mesh sequence 108.
[0046] It should be noted that mesh coding / decoding system 100 is presented by way of example and not limitation. In the example of FIG. 1 , mesh coding / decoding system 100 may have other components and / or arrangements. For example, mesh source 112 may be external to source device 102. Similarly, mesh display 122 may be external to destination device 106 or omitted altogether where mesh sequence is intended for consumption by a machine and / or storage device. In another example, source device 102 may further comprise a mesh decoder and destination device 106 may comprise a mesh encoder. In such an example, source device 102 may be configured to further receive an encoded bit stream from destination device 106 to support two-way mesh transmission between the devices.
[0047] FIG. 2A illustrates a block diagram of an example encoder 200A for intra encoding a 3D mesh, according to some embodiments. For example, an encoder (e.g., encoder 114) may comprise encoder 200A.
[0048] In some examples, a mesh sequence (e.g., mesh sequence 108) may include a set of mesh frames (e.g., mesh frames 124) that may be individually encoded and decoded. As will be further described below with respect to FIG. 4, a base mesh 252 may be determined (e.g., generated) from a mesh frame (e.g., an input mesh) through a decimation process. In the decimation process, the mesh topology of the mesh frame may be reduced to determine the base mesh (e.g., a decimated mesh or decimated base mesh). A mesh encoder 204 may encode base mesh 252, whose geometry information (e.g., vertices) may be quantized by quantizer 202, to generate a base mesh bitstream 254. In some examples, base mesh encoder 204 may be an existing encoder such as Draco or Edgebreaker.
[0049] Displacement generator 208 may generate displacements for vertices of the mesh frame based on base mesh 252, as will be further explained below with respect to FIGS. 4 and 5. In some examples, the displacements are determined based on a reconstructed base mesh 256. Reconstructed base mesh 256 may be determined (e.g., output or generated) by mesh decoder 206 that decodes the encoded base mesh (e.g., in base mesh bitstream 254) determined (e.g., output or generated) by mesh encoder 204. Displacement generator 208 may subdivide reconstructed base mesh 256 using a subdivision scheme (e.g., subdivision algorithm) to determine a subdivided mesh (e.g., a subdivided base mesh). Displacement 258 may be determined based on fitting the subdivided mesh to an original input mesh surface. For example, displacement 258 for a vertex in the mesh frame may include displacement information (e.g., a displacement vector) that indicates a displacement from the position of the corresponding vertex in the subdivided mesh to the position of the vertex in the mesh frame.
[0050] Displacement 258 may be transformed by wavelet transformer 210 to generate wavelet coefficients (e.g., transformation coefficients) representing the displacement information and that may be more efficiently encoded (and subsequently decoded). The wavelet coefficients may be quantized by quantizer 212 and packed (e.g., arranged) by image packer 214 into a picture (e.g., one or more images or picture frames) to be encoded by video encoder 216. Mux 218 may combine (e.g., multiplex) the displacement bitstream 260 output by video encoder 216 together with base mesh bitstream 254 to form bitstream 266.
[0051] Attribute information 262 (e.g., color, texture, etc.) of the mesh frame may be encoded separately from the geometry information of the mesh frame described above. In some examples, attribute information 262 of the mesh frame may be represented (e g., stored) by an attribute map (e.g., texture map) that associates each vertex of the mesh frame with corresponding attributes information of that vertex. Attribute transfer 232 may re-parameterize attribute information 262 in the attribute map based on reconstructed mesh determined (e.g., generated or output) from mesh reconstruction components 225. Mesh reconstruction components 225 perform inverse or decoding functions and may be the same or similar components in a decoder (e.g., decoder 300 of FIG. 3). For example, inverse quantizer 228 may inverse quantize reconstructed base mesh 256 to determine (e.g., generate or output) reconstructed base mesh 268. Video decoder 226, image unpacker 224, inverse quantizer 222, and inverse wavelet transformer 220 may perform the inverse functions as that of video encoder 216, image packer 214, quantizer 212, and wavelet transformer 210, respectively. Accordingly, reconstructed displacement 270, corresponding to displacement 258, may be generated from applying video decoder 226, image unpacker 224, inverse quantizer 222, and inverse wavelet transformer 220 in that order. Deformed mesh reconstructor 230 may determine the reconstructed mesh, corresponding to the input mesh frame, based on reconstructed base mesh 268 and reconstructed displacement 270. In some examples, the reconstructed mesh may be the same decoded mesh determined from the decoder based on decoding base mesh bitstream 254 and displacement bitstream 260.
[0052] Attribute information of the re-parameterized attribute map may be packed in images (e.g., 2D images or picture frames) by padding component 234 Padding component 234 may fill (e.g., pad) portions of the images that do not contain attribute information. In some examples, color-space converter 236 may translate (e.g., convert) the representation of color (e.g., an example of attribute information 262) from a first format to a second format (e.g., from RGB444 to YUV420) to achieve improved rate-distortion (RD) performance when encoding the attribute maps. In an example, color-space converter 236 may also perform chroma subsampling to further increase encoding performance. Finally, video encoder 240 encodes the images (e.g., pictures frames) representing attribute information 262 of the mesh frame to determine (e.g., generate or output) attribute bitstream 264 multiplexed by mux 218 into bitstream 266. In some examples, video encoder 240 may be an existing 2D video compression encoder such as an HEVC encoder or a VVC encoder.
[0053] FIG. 2B illustrates a block diagram of an example encoder 200B for inter encoding a 3D mesh, according to some embodiments. For example, an encoder (e.g., encoder 114) may comprise encoder 200B. As shown in FIG. 2B, encoder 200B comprises many of the same components as encoder 200A. In contrast to encoder 200A, encoder 200B does not include mesh encoder 204 and mesh decoder 206, which correspond to coders for static 3D meshes. Instead, encoder 200B comprises a motion encoder 242, a motion decoder 244, and a base mesh reconstructor 246. Motion encoder 242 may determine a motion field (e.g., one or more motion vectors (MVs)) that, when applied to a reconstructed quantized reference base mesh 243, best approximates base mesh 252.
[0054] The determined motion field may be encoded in bitstream 266 as motion bitstream 272. In some examples, the motion field (e.g., a motion vector in the x, y, and z directions) may be entropy coded as a codeword (e.g., for each directional component) resulting from a coding scheme such as a unary, a Golomb code (e.g., Exp-Golomb code), a Rice code, or a combination thereof. In some examples, the codeword may be arithmetically coded, e.g., using CABAC. A prefix part of the codeword may be context coded and a suffix part of the coded may be bypass coded. In some examples, a sign bit for each directional component of the motion vector may be coded separately.
[0055] In some examples, motion bitstream 272 may further include indication of the selected reconstructed quantized reference base mesh 243.
[0056] In some examples, motion bitstream 272 may be decoded by motion decoder 244 and used by base mesh reconstructor 246 to generate reconstructed quantized base mesh 256. For example, base mesh reconstructor 246 may apply the decoded motion field to reconstructed quantized reference base mesh 243 to determine (e.g., generate) reconstructed quantized base mesh 256.
[0057] In some examples, a reconstructed quantized reference base mesh m’(j) associated with a reference mesh frame with index j may be used to predict the base mesh m(i) associated with the current frame with index i. Base meshes m(i) and m(j) may comprise the same: number of vertices, connectivity,texture coordinates, and texture connectivity. The positions of vertices may differ between base meshes m(i) and m(j).
[0058] In some examples, the motion field f(i) may be computed by considering the quantized version of m(i) and the reconstructed quantized base mesh m’(j). Base mesh m’(j) may have a different number of vertices than m(j) (e.g., vertices may have been merged or removed). Therefore, the encoder may track the transformation applied to m(j) to determine (e.g., generate or obtain) m’(j) and apply it to m(i). This transformation may enable a 1 -to-1 correspondence between vertices of base mesh m’(j) and the transformed and quantized version of base mesh m(i), denoted as mA* (i). The motion field f(i) may be computed by subtracting the quantized positions Pos(i,v) of the vertex v of mA* (i) from the positions Pos(j,v) of the vertex v of m’(j) as follows: f(i,v) = Pos(j ,v) - Pos(i,v). The motion field may be further predicted by using the connectivity information of base mesh m’(j) and the prediction residuals may be entropy encoded.
[0059] In some examples, since the motion field compression process may be lossy, a reconstructed motion field denoted as f’(i) may be computed by applying the motion decoder component. A reconstructed quantized base mesh m’(i) may then be computed by adding the motion field to the positions of vertices in base mesh m'(j). To better exploit temporal correlation in the displacement and attribute map images (e.g., sequence / video of images), inter prediction may be enabled in the video encoder.
[0060] In some embodiments, an encoder (e.g., encoder 114) may comprise encoder 200A and encoder 200B.
[0061] FIG. 3 illustrates a diagram showing an example decoder 300. Bitstream 330, which may correspond to bitstream 266 in FIGS. 2A and 2B and may be received in a binary file, may be demultiplexed by de-mux 302 to separate bitstream 330 into base mesh bitstream 332, displacement bitstream 334, and attribute bitstream 336 carrying base mesh geometry information, displacement geometry information, and attribute information, respectively. Attribute bitstream 336 may include one or more attribute map sub-streams for each attribute type.
[0062] In some examples, for inter decoding, the bitstream is de-multiplexed into separate sub-streams, including: a motion sub-stream, a displacement sub-stream for positions and potentially for each vertex attribute, zero or more attribute map sub-streams, and an atlas sub-stream containing patch information in the same manner as in V3CA / -PCC.
[0063] In some examples, base mesh bitstream 332 may be decoded in an intra mode or an inter mode. In the intra mode, static mesh decoder 320 may decode base mesh bitstream 332 (e.g., to generate reconstructed base mesh m’(i)) that is then inverse quantized by inverse quantizer 318 to determine (e.g., generate or output) decoded base mesh 340 (e.g., reconstructed quantized base mesh m”(i)). In some examples, static mesh decoder 320 may correspond to mesh decoder 206 of FIG. 2A.
[0064] In some examples, in the inter mode, base mesh bitstream 332 may include motion field information that is decoded by motion decoder 324. In some examples, motion decoder 324 may correspond to motion decoder 244 of FIG. 2B. For example, motion decoder 324 may entropy decode base mesh bitstream 332 to determine motion field information. In the inter mode, base mesh bitstream 332 may indicate a previous base mesh (e.g., reference base mesh m’(j)) decoded by static mesh decoder 320 and stored (e.g., buffered) in mesh buffer 322. Base mesh reconstructor 326 may generate a quantized reconstructed base mesh m’(i) by applying the decoded motion field (output by motion decoder 324) to the previously decoded (e.g., reconstructed) base mesh m’(j) stored in mesh buffer 322. In some examples, base mesh reconstructor 326 may correspond to base mesh reconstructor 246 of FIG. 2B. The quantized reconstructed base mesh may be inverse quantized by inverse quantizer 318 to determine (e.g., generate or output) decoded base mesh 340 (e.g., reconstructed base mesh m”(i)). In some examples, decoded base mesh 340 may be the same as reconstructed base mesh 268 in FIGS. 2A and 2B.
[0065] In some examples, decoder 300 includes video decoder 308, image unpacker 310, inverse quantizer, and inverse wavelet transformer 314 that determines (e.g., generates) decoded displacement 338 from displacement bitstream 334. Video decoder 308, image unpacker 310, inverse quantizer, and inverse wavelet transformer 314 correspond to video decoder 226, image unpacker 224, inverse quantizer 222, and inverse wavelet transformer 220, respectively, and perform the same or similar operations. For example, the picture frames (e.g., images) received in displacement bitstream 334 may be decoded by video decoder 308, the displacement information may be unpacked by image unpacker 310 from the decoded image, inverse quantized by inverse quantizer 312 to determined inverse quantized wavelet coefficients representing encoded displacement information. Then, the unquantized wavelet coefficients may be inverse transformed by inverse wavelet transformer 314 to determine decoded displacement d”(i). In other words decoded displacement 338 (e.g., decoded displacement field d”(i)) may be the same as reconstructed displacement 270 in FIGS. 2A and 2B.
[0066] Deformed mesh reconstructor 316, which corresponds to deformed mesh reconstructor 230, may determine (e.g., generate or output) decoded mesh 342 (M”(i)) based on decoded displacement 338 and decoded base mesh 340. For example, deformed mesh reconstructor 316 may combine (e.g., add) decoded displacement 338 to a subdivided decoded mesh 340 to determine decoded mesh 342. Specifically, decoded displacement 338 may include a respective reconstructed displacement vector corresponding to each vertex of a subdivided mesh (e.g., identically generated as subdivided mesh 442 by encoder in FIG. 4) generated from decoded base mesh 340. For example, the decoder may apply a subdivision scheme to subdivide decoded base mesh 340 to generate the subdivided mesh, as described with respect to FIG. 6
[0067] In some examples, decoder 300 includes video decoder 304 that decodes attribute bitstream 336 comprising encoded attribute information represented (e.g., stored) in 2D images (or picture frames) to determined attribute information 344 (e.g., decoded attribute information or reconstructed attribute information). In some examples, video decoder 304 may be an existing 2D video compression decoder such as an HEVC decoder or a VVC decoder. Decoder 300 may include a color-space converter 306, which may revert the color format transformation performed by color-space converter 236 in FIGS. 2A and 2B.
[0068] FIG. 4 is a diagram 400 showing an example process (e.g., a pre-processing operations) for generating displacements 414 of an input mesh 430 (e.g., an input 3D mesh frame) to be encoded, according to some embodiments. In some examples, displacements 414 may correspond to displacement 258 shown in FIG. 2A and FIG. 2B.
[0069] In diagram 400, a mesh decimator 402 determines (e.g., generates or outputs) an initial base mesh 432 based on (e.g., using) input mesh 430. In some examples, the initial base mesh 432 may be determined (e.g., generated) from the input mesh 432 through a decimation process. In the decimation process, the mesh topology of the mesh frame may be reduced to determine the initial base mesh (which may be referred to as a decimated mesh or decimated base mesh). As will be illustrated in FIG. 5, the decimation process may involve a down sampling process to remove vertices from the input mesh 432 so that a small portion (e.g., 6% or less) of the vertices in the input mesh 430 may remain in the initial base mesh 432.
[0070] Mesh subdivider 404 applies a subdivision scheme to generate initial subdivided mesh 434. As will be discussed in more detail with regard to FIG. 5, the subdivision scheme may involve upsampling the initial base mesh 432 to add more vertices to the 3D mesh based on the topology and shape of the original mesh to generate the initial subdivided mesh 434.
[0071] Fitting component 406 may fit the initial subdivided mesh to determine a deformed mesh 436 that may more closely approximate the surface of input mesh 430. As will be discussed in more detail with respect to FIG. 5, the fitting may be performed by moving vertices of the initial subdivided mesh 434 towards the surfaces of the input mesh 430 so that the subdivided mesh 434 can be used to approximate the input mesh 430. In some implementations, the fitting is performed by moving each vertex of the initial subdivided mesh 434 along the normal direction of the vertex until the vertex intersects with a surface of the input mesh 430. The resulting mesh is the deformed mesh 436. The normal direction may be indicated by a vertex normal at the vertex, which may be obtained from face normals of triangles formed by the vertex.
[0072] Base mesh generator 408 may perform another fitting process to generate a base mesh 438 from the initial base mesh 432. For example, the base mesh generator 408 may deform the initial base mesh 432 according to the deformed mesh 436 so that the initial base mesh 432 is close to the deformed mesh436. In some implementations, the fitting process may be performed in a similar manner to the fitting component 406. For example, the base mesh generator 408 may move each of the vertices in the initial base mesh 432 along its normal direction (e.g., based on the vertex normal at each vertex) until the vertex reaches a surface of the deformed mesh 436. The output of this process is the base mesh 438.
[0073] Base mesh 438 may be output to a mesh reconstruction process 410 to generate a reconstructed base mesh 440. Reconstructed base mesh 440 may be subdivided by mesh subdivider 418 and the subdivided mesh 442 may be input to displacement generator 420 to generate (e.g., determine or output) displacement 414, as further described below with respect to FIG. 5. In some examples, mesh subdivider 418 may apply the same subdivision scheme as that applied by mesh subdivider 404. In these examples, vertices in the subdivided mesh 442 have a one-to-one correspondence with the vertices in the deformed mesh 436. As such, the displacement generator 420 may generate the displacements 414 by calculating the difference between each vertex of the subdivided mesh 442 and the corresponding vertex of the deformed mesh 436. In some implementations, the difference may be projected onto a normal direction of the associated vertex and the resulting vector is the displacement 414. In this way, only the sign and magnitude of the displacement 414 need to be encoded in the bitstream, thereby increasing the coding efficiency. In addition, because the base mesh 438 has been fitted toward the deformed mesh 436, the displacements 414 between the deformed mesh 436 and the subdivided mesh 442 (generated from the reconstructed base mesh 440) will have small magnitudes, which further reduces the payload and increases the coding efficiency.
[0074] In some examples, one advantage of applying the subdivision process is to allow for more efficient compression, while offering a faithful approximation of the original input mesh 430 (e.g., surface or curve of the original input mesh 430). The compression efficiency may be obtained because the base mesh (e.g., decimated mesh) has a lower number of vertices compared to the number of vertices of input mesh 430 and thus requires a fewer number of bits to be encoded and transmitted. Additionally, the subdivided mesh may be automatically generated by the decoder once the base mesh has been decoded without any information needed from the encoder other than a subdivision scheme (e.g., subdivision algorithm) and parameters for the subdivision (e.g., a subdivision iteration count). The reconstructed mesh may be determined by decoding displacement information (e.g., displacement vectors) associated with vertices of the subdivided mesh (e.g., subdivided curves / surfaces of the base mesh). Not only does the subdivision process allow for spat! al / qual ity scalability, but also the displacements may be efficiently coded using wavelet transforms (e.g., wavelet decomposition), which further increases compression performance.
[0075] In some embodiments, mesh reconstruction process 410 includes components for encoding and then decoding base mesh 438. FIG. 4 shows an example for the intra mode, in which mesh reconstruction process 410 may include quantizer 411, static mesh encoder 412, static mesh decoder 413, and inverse quantizer 416, which may perform the same or similar operations as quantizer 202,mesh encoder 204, mesh decoder 206, and inverse quantizer 228, respectively, from FIG. 2A. For the inter mode, mesh reconstruction process 410 may include quantizer 202, motion encoder 242, motion decoder 244, base mesh reconstructor 246, and inverse quantizer 228.
[0076] FIG. 5 illustrates an example process for approximating and encoding a geometry of a 3D mesh, according to some embodiments. For illustrative purposes, the 3D mesh is shown as 2D curves. An original surface 510 of the 3D mesh (e.g., a mesh frame) includes vertices (e.g., points) and edges that connect neighboring vertices. For example, point 512 and point 513 are connected by an edge corresponding to surface 514.
[0077] In some examples, a decimation process (e.g., a down-sampling process or a decimation / down- sampling scheme) may be applied to an original surface 510 of the original mesh to generate a down- sampled surface 520 of a decimated (or down-sampled) mesh In the context of mesh compression, decimation refers to the process of reducing the number of vertices in a mesh while preserving its overall shape and topology. For example, original mesh surface 510 is decimated into a surface 520 with fewer samples (e.g., vertices and edges) but still retains the main features and shape of the original mesh surface 510. This down-sample surface 520 may correspond to a surface of the base mesh (e.g., a decimated mesh).
[0078] In some examples, after the decimation process, a subdivision process (e.g., subdivision scheme or subdivision algorithm) may be applied to down-sampled surface 520 to generate an up-sampled surface 530 with more samples (e.g., vertices and edges). Up-sampled surface 530 may be part of the subdivided mesh (e.g , subdivided base mesh) resulting from subdividing down-sampled surface 520 corresponding to a base mesh.
[0079] Subdivision is a process that is commonly used after decimation in mesh compression to improve the visual quality of the compressed mesh. The subdivision process involves adding new vertices and faces to the mesh based on the topology and shape of the original mesh. In some examples, the subdivision process starts by taking the reduced mesh that was generated by the decimation process and iteratively adding new vertices and edges. For example, the subdivision process may comprise dividing each edge (or face) of the red uced / d eci mated mesh into shorter edges (or smaller faces) and creating new vertices at the points of division. These new vertices are then connected to form new faces (e.g., triangles, quadrilaterals, or another polygon). By applying subdivision after the decimation process, a higher level of compression can be achieved without significant loss of visual fidelity. Various subdivision schemes may be used such as, e.g., mid-point, Catmull-Clark subdivision, Butterfly subdivision, Loop subdivision, etc., or a combination thereof.
[0080] For example, FIG. 5 illustrates an example of the mid-point subdivision scheme applied to down- sampled surface 520 to generate up-sampled surface 530. In this scheme, each subdivision iteration subdivides each triangle into four sub-triangles. New vertices are introduced in the middle of each edge.The subdivision process may be applied independently to the geometry and to the texture coordinates since the connectivity for the geometry and for the texture coordinates are usually different. The subdivision scheme computes the position Pos(v12) of a newly introduced vertex v12at the center of an edge (v1;v2) formed by a first vertex (v and a second vertex (v2), as follows:Pos(v12) = i (Pos(Vi) + Pos(v2)), (Eq 1) where Pos v and Pos(v2) are the positions of the vertices VT and v2. In some examples, the same process may be used to compute the texture coordinates of the newly created vertex. For normal vectors, a normalization step may be applied as follows:and N(v2) are the normal vectors associated with the vertices v12,and v2, respectively. 11 x 11 is the norm2 of the vector x.
[0081] Using the mid-point subdivision scheme, as shown in up-sampled surface 530, point 531 may be generated as the mid-point of edge 522 which is an edge connecting point 532 and point 533. Point 531 may be added as a new vertex. Edge 534 and edge 542 are also added to connect the added new vertex corresponding to point 531 . In some examples, the original edge 522 may be replaced by two new edges 534 and 542.
[0082] In some examples, down-sampled surface 520 may be iteratively subdivided to generate up- sampled surface 530. For example, a first subdivided mesh resulting from a first iteration of subdivision applied to down-sampled surface 520 may be further subdivided according to the subdivision scheme to generate a second subdivided mesh, etc. In some examples, a number of iterations corresponding to levels of subdivision may be predetermined. In other examples, an encoder may indicate the number of iterations to a decoder, which may similarly generate a subdivided mesh, as further described above.
[0083] In some examples, the subdivision scheme is applied identically at the encoder and the decoder to generate the same up-sampled surface 530 from the same down-sampled surface 520. For example, the encoder may signal (e.g., encode) information representing vertices of down-sampled surface 520, which may be the base mesh, in a bitstream. The decoder may decode the information from the bitstream to obtain the vertices of down-sampled surface 520.
[0084] In some embodiments, the subdivided mesh may be deformed towards (e.g., approximates) the original mesh to determine (e.g., get or obtain) a prediction of the original mesh having original surface 510. The points on the subdivided mesh may be moved along a computed normal vertex / orientation until it reaches an original surface 510 of the original mesh. For example, each point (also referred to as vertex) may be associated with a vertex normal computed from a normalized average of face normals (also referred to as surface normals) ef faces containing that point. The vertex normal is a directional vertex indicating the normal direction of the point. The distance between the intersected point on theoriginal surface 510 and the subdivided point may be computed as a displacement (e.g., a displacement vector). For example, point 531 may be moved towards the original surface 510 along a computed normal orientation of surface (e.g., represented by edge 542). When point 531 intersects with surface 514 of the original surface 510 (of original / input mesh), a displacement vector 548 can be computed. Displacement vector 548 applied to point 531 may result in displaced surface 540, which may better approximate original surface 510. In some examples, displacement information (e.g., displacement vector 548) for vertices of the subdivided mesh (e.g., up-sampled surface 530 of subdivided mesh) may be encoded and transmitted in displacement bitstream 260 shown in examples encoders of FIGS. 2A and 2B. Note, as explained with respect to FIG. 4, the subdivided mesh corresponding to up-sampled surface may be subdivided mesh 442 that is compared to deformed mesh 436 representative of original surface 510 of the input mesh.
[0085] In some embodiments, displacements d(i) (e.g., a displacement field or displacement vectors) may be computed and / or stored based on local coordinates or global coordinates. For example, a global coordinate system is a system of reference that is used to define the position and orientation of objects or points in a 3D space. It provides a fixed frame of reference that is independent of the objects or points being described. The origin of the global coordinate system may be defined as the point where the three axes intersect. Any point in 3D space can be located by specifying its position relative to the origin along the three axes using Cartesian coordinates (x, y, z). For example, the displacements may be defined in the same cartesian coordinate system as the input or original mesh. Accordingly, a displacement may comprise three components (in the x, y, and z directions).
[0086] In a local coordinate system, a normal, a tangent, and / or a binormal vector (which are mutually perpendicular) may be determined that defines a local basis for the 3D space to represent the orientation and position of an object in space relative to a reference frame. In some examples, displacement field d(i) may be transformed from the canonical coordinate system to the local coordinate system, e.g., defined by a normal to the subdivided mesh at each vertex (e.g., commonly referred to as a vertex normal). The normal at each vertex may be obtained from combining the face normals of triangles formed by the vertex. In some examples, using the local coordinate system may enable further compression of tangential components of the displacements compared to the normal component. For example, the displacements may be signaled as a scalar value (e.g., including a sign and a magnitude) which may be used to derive a displacement vector based on the normal at the vertex. A normal vector at the vertex may be computed based on a normalized average of face normals (or surface normals) of faces (of the 3D mesh) containing that vertex. Accordingly, using local coordinate system, displacements need not be signaled as three components corresponding to the directions of the canonical coordinate system.
[0087] In some embodiments, a decoder (e g., decoder 300 of FIG. 3) may receive and decode a base mesh corresponding to (e.g., having) down-sampled surface 520. Similar to the encoder, the decodermay apply a subdivision scheme to determine a subdivided mesh having up-sampled surface 530 generated from down-sampled surface 520. The decoder may receive and decode displacement information including displacement vector 548 and determine a decoded mesh (e.g., reconstructed mesh) based on the subdivided mesh (corresponding to up-sampled surface 530) and the decoded displacement information. For example, the decoder may add the displacement at each vertex with a position of the corresponding vertex in the subdivided mesh. The decoder may obtain a reconstructed 3D mesh by combining the obtained / decoded displacements with positions of vertices of the subdivided mesh.
[0088] FIG. 6 illustrates an example of vertices of a subdivided mesh (e.g., a subdivided base mesh) corresponding to multiple levels of detail (LODs), according to some embodiments. As described above with respect to FIG. 5, the subdivision process (e.g., subdivision scheme) may be an iterative process, in which a mesh can be subdivided multiple times and a hierarchical data structure is generated containing multiple levels. Each level of the hierarchical data structure may include different numbers of data samples (e.g., vertices and edges in mesh) representing (e.g., forming) different density / resolution (e.g., also referred to as levels of details (LoDs)). For example, a down-sampled surface 520 (of a decimated mesh) can be subdivided into up-sampled surface 530 after a first iteration of subdivision.
[0089] Down-sampled surface 520 may represent a decimated mesh, which is also referred to as a base mesh (e.g., reconstructed base mesh 256 or reconstructed base mesh 268 of FIGS. 2A-B, decoded base mesh 340 of FIG. 3, or reconstructed base mesh 440 of FIG. 4).
[0090] Up-sampled surface 530 may be further subdivided into up-sampled surface 630 and so forth. In this case, vertices of the mesh with down-sampled surface 520 may be considered as being in or associated with LODO. Vertices, such as vertex 632, generated in up-sampled surface 530 after a first iteration of subdivision may be at LOD1 . Vertices, such as vertex 634, generated in up-sampled surface 630 after another iteration of subdivision may be at LOD2, etc. In some examples, an LODO may refer to the vertices resulting from decimation of an input (e.g., original) mesh resulting in a base mesh with (e.g., having) down-sampled surface 520. For example, vertices at LODO may be vertices of a reconstructed quantized base mesh 256 of FIGS. 2A-B, reconstructed / decoded base mesh 340 of FIG. 3, reconstructed base mesh 440 of FIG. 4.
[0091] In some examples, the computation of displacements in different LODs follows the same mechanism as described above with respect to FIG. 5. In some examples, a displacement vector 643 may be computed from a position of a vertex 641 in the original surface 510 (of original mesh) to a vertex 642, from displace surface 640 of the deformed mesh, at LODO. The displacement vectors 644 and 645 of corresponding vertices 632 and 634 from LOD1 and LOD 2, respectively, may be similarly calculated. Accordingly, in some examples, a number of iterations of subdivision may correspond to a number of LODs and one of the iterations may correspond to one LOD of the LODs.
[0092] The encoder may signal transformed wavelet coefficients in a bitstream to represent the computed displacements. A displaced surface 640 is of a deformed mesh resulting from applying the displacements to vertices of up-sampled surface 630. This deformed mesh may be the 3D mesh reconstructed at the decoder.
[0093] In some examples, at the decoder, the displacements for vertices of up-sampled surface 630 may be reconstructed from the bitstream based on inverse transforming the transformed wavelet coefficients. Then, the decoder may apply each reconstructed displacement to a respective vertex of up-sampled surface 630 to obtain displaced surface 640 of the reconstructed 3D mesh. For example, displacement vector 643 may be reconstructed and applied to a respective vertex 642 of up-sampled surface 630 to obtain vertex 641 of the reconstructed 3D mesh.
[0094] FIG. 7A illustrates an example of an image 720 (e.g., picture or a picture frame) packed with displacements 700 (e.g., displacement fields or vectors) using a packing method (e.g., a packing scheme or a packing algorithm), according to some embodiments. Specifically, displacements 700 may be generated, as described above with respect to FIG. 5 and FIG. 6, and packed into 2D images. In some examples, a displacement can be a 3D vector containing the values for the three components of the distance. For example, a delta x value represents the shift on the x-axis from a point A to a point B in a Cartesian coordinate system. In some examples, a displacement vector may be represented by less than three components, e.g., by one or two components. For example, when a local coordinate system is used to store the displacement value, one component with the highest significance may be stored as being representative of the displacement and the other components may be discarded.
[0095] In some examples, as will be further described below, a displacement value may be transformed into other signal domains for achieving better compression. For example, a displacement can be wavelet transformed and be decomposed into and represented as wavelet coefficients (e.g., coefficient values or transform coefficients). In these examples, displacements 700 that are packed in image 720 may comprise the resulting wavelet coefficients (e.g., transform coefficients), which may be more efficiently compressed than the un-transformed displacement values. At the decoder side, a decoder may decode displacements 700 as wavelet coefficients and may apply an inverse wavelet transform process to reconstruct the original displacement values obtained at the encoder.
[0096] In some examples, one or more of displacements 700 may be quantized by the encoder before being packed into displacement image 720. In some examples, one or more displacements may be quantized before being wavelet transformed, after being wavelet transformed, or quantized before and after being wavelet transformed. For example, FIG. 7A shows quantized wavelet transform values 8, 4, 1 , -1 , etc. in displacements 700. At the decoder side, the decoder may perform inverse quantization to reverse or undo the quantization process performed by the encoder.
[0097] In general, quantization in signal processing may be the process of mapping input values from a larger set to output values in a smaller set. It is often used in data compression to reduce the amount, the precision, or the resolution of the data into a more compact representation. However, this reduction can lead to a loss of information and introduce compression artifacts. The choice of quantization parameters, such as the number of quantization levels, is a trade-off between the desired level of precision and the resulting data size. There are many different quantization techniques, such as uniform quantization, non- uniform quantization, and adaptive quantization that may be selected / enabled / applied. They can be employed depending on the specific requirements of the application.
[0098] In some examples, wavelet coefficients (e.g., displacement coefficients representing displacement signals) may be adaptively quantized according to LODs. As explained above, a mesh may be iteratively subdivided to generate a hierarchical data structure comprising multiple LODs. In this example, each vertex and its associated displacement belong to the same level of hierarchy in the LOD structure, e.g., an LOD corresponding to a subdivision iteration in which that vertex was generated. In some examples, a vertex at each LOD may be quantized according to quantization parameters, corresponding to LODs, that specify different levels of intensity / precision of the signal to be quantized. For example, wavelet coefficients in LOD 3 may have a quantization parameter of, e.g., 42 and wavelet coefficients in LOD 0 may have a different, smaller quantization parameter of 28 to preserve more detail information in LOD 0.
[0099] In some examples, displacements 700 may be packed onto the pixels in a displacement image 720 with a width W and a height H. In an example, a size of displacement image 720 (e.g., W multiplied by H) may be greater or equal to the number of components in displacements 700 to ensure all displacement information may be packed. In some examples, displacement image 720 may be further partitioned into smaller regions (e.g., squares) referred to as a packing block 730. In an example, the length of packing block 730 may be an integer multiple of 2.
[0100] Displacements 700 (e.g., displacement signals represented by quantized wavelet coefficients) may be packed into a packing block 730 according to a packing order 732. Each packing block 730 may be packed (e.g., arranged or stored) in displacement image 720 according to a packing order 722. Once all the displacements 700 are packed, the empty pixels in image 720 may be padded with neighboring pixel values for improved compression. In the example shown in FIG. 7A, packing order 722 for packing blocks may be a raster order and a packing order 732 for displacements within packing block 730 may be, for example, a Z-order. However, it should be understood that other packing schemes both for blocks and displacements within blocks may be used. In some embodiments, a packing scheme for the blocks and / or within the blocks may be predetermined. In some embodiments, the packing scheme may be signaled by the encoder in the bitstream per patch, patch group, tile, image, or sequence of images. Relatedly, the signaled packing scheme may be obtained by the decoder from the bitstream.
[0101] In some examples, packing order 732 may follow a space-filling curve, which specifies a traversal in space in a continuous, non-repeating way. Some examples of space-filling curve algorithms (e.g , schemes) include Z-order curve, Hilbert Curve, Peano Curve, Moore Curve, Sierpinski Curve, Dragon Curve, etc. Space-filling curves have been used in image packing techniques to efficiently store and retrieve images in a way that maximizes storage space and minimizes retrieval time. Space-filling curves are well-suited to this task because they can provide a one-dimensional representation of a two- dimensional image. One common image packing technique that uses space-filling curves is called the Z- order or Morton order. The Z-order curve is constructed by interleaving the binary representations of the x and y coordinates of each pixel in an image. This creates a one-dimensional representation of the image that can be stored in a linear array. To use the Z-order curve for image packing, the image is first divided into small blocks, typically 8x8 or 16x16 pixels in size. Each block is then encoded using the Z-order curve and stored in a linear array. When the image needs to be retrieved, the blocks are decoded using the inverse Z-order curve and reassembled into the original image.
[0102] In some examples, once packed, displacement image 720 may be encoded and decoded using a conventional 2D video codec.
[0103] FIG. 7B illustrates an example of displacement image 720, according to some embodiments. As shown, displacements 700 packed in displacement image 720 may be ordered according to their LODs. For example, displacement coefficients (e.g., quantized wavelet coefficients) may be ordered from a lowest LOD to a highest LOD. In other words, a wavelet coefficient representing a displacement for a vertex at a first LOD may be packed (e.g., arranged and stored in displacement image 720) according to the first LOD. For example, displacements 700 may be packed from a lowest LOD to a highest LOD. Higher LODs represent a higher density of vertices and corresponds to more displacements compared to lower LODs. The portion of displacement image 720 not in any LOD may be a padded portion.
[0104] In some examples, displacements may be packed in inverse order from highest LOD to lowest LOD. In an example, the encoder may signal whether displacements are packed from lowest to highest LOD or from highest to lowest LOD.
[0105] In some examples, a wavelet transform may be applied to displacement values to generate wavelet coefficients (e.g., displacement coefficients) that may be more easily compressed. Wavelet transforms are commonly used in signal processing to decompose a signal into a set of wavelets, which are small wave-like functions allowing them to capture localized features in the signal. The result of the wavelet transform is a set of coefficients that represent the contribution of each wavelet at different scales and positions in the signal. It is useful for detecting and localizing transient features in a signal and is generally used for signal analysis and data compression such as image, video, and audio compression.
[0106] Taking a 2D image as an example, wavelet transform is used to decompose an image (signals) into two discrete components, known as approximations / predictions and details. The decomposed signalsare further divided into a high frequency component (details) and a low frequency component (approximations / predictions) by passing through two filters, high and low pass filters. In the example of 2D image, two filtering stages, a horizontal and a vertical filtering are applied to the image signals. A downsampling step is also required after each filtering stage on the decomposed components to obtain the wavelet coefficients resulting in four sub-signals in each decomposition level. The high frequency component corresponds to rapid changes or sharp transitions in the signal, such as an edge or a line in the image. On the other hand, the low frequency component refers to global characteristics of the signal. Depending on the application, different filtering and compression can be achieved. There are various types of wavelets such as Haar, Daubechies, Symlets, etc., each with different properties such as frequency resolution, time localization, etc.
[0107] In signal processing, a lifting scheme is a technique for both designing wavelets and performing the discrete wavelet transform (DWT). It is an alternative approach to the traditional filter bank implementation of the DWT that offers several advantages in terms of computational efficiency and flexibility. It decomposes the signal using a series of lifting steps such that the input signal, e.g., representing displacements for 3D meshes, may be converted to displacement coefficients in-place. In the lifting scheme, a series of lifting operations (e.g. lifting steps) may be performed. Each lifting operation involves a prediction step (e.g., prediction operation) and an update step (e.g., update operation). These lifting operations may be applied iteratively to obtain the wavelet coefficients.
[0108] FIG. 8 illustrates an example of a lifting scheme for representing displacement information of a 3D mesh as wavelet coefficients, according to some embodiments. The lifting scheme may refer to a forward lifting scheme 802 and / or an inverse lifting scheme 804. The lifting scheme comprises a plurality of lifting operations, which may be iteratively performed. Each lifting operation may include a prediction operation (e.g., prediction step) and an update operation (e.g., an update step). An encoder may perform (e.g., apply) forward lifting scheme 802 to determine (e.g., derive, generate, or obtain) wavelet coefficients representing displacement information. A decoder may perform (e.g., apply) inverse lifting scheme 804 to reverse the operations of forward lifting scheme 802 to determine (e.g., derive, generate, or obtain) the displacement information from wavelet coefficients decoded from a bitstream. The decoded displacement information may include displacement values (e.g., displacement vectors) corresponding to vertices of a 3D mesh frame, which may be used by the decoder to generate a decoded mesh (e.g , a reconstructed mesh).
[0109] Forward lifting scheme 802 comprises a plurality of iterations corresponding to a plurality of LODs, e.g., shown as LODN 810, LODN-I 812, LODN-2814, and LODo 816. Each iteration of forward lifting scheme 802 (e.g., four iterations are shown as four dotted boxes corresponding to LODs 810-816) includes a splitting operation (e.g., a splitting step shown as a “Split” component), a prediction operation(e.g a prediction step shown as a “P” component), and an update operation (e.g., an update step shown as a “U” component).
[0110] The splitting operation separates (or splits) signal Sj (j > 1) into two signals (e.g., nonoverlapping signals): the even samples signal denoted by seveilk(k e [0, j - 1]) and the odd -samples signal denoted by soddk. Signal s, represents the displacement values (e.g., displacement signals) determined for vertices of the 3D mesh frame. For example, a displacement value comprises a displacement field (e.g., a displacement vector), which may be one, two, or three components, as explained above. In each lifting operation iteration, the odd samples soddkinclude the displacement coefficients of vertices at an LOD corresponding to the iteration. For each odd sample of the odd samples soddk. the even samples seve„kmay include the two displacement coefficients of the two vertices, of the 3D mesh frame, from which the vertex corresponding to the odd sample was generated during a mesh subdivision or down-sampling process, as explained above with respect to FIG. 6. Since vertices at the LOD are generated from vertices at lower LODs, the two vertices of the 3D mesh frame are at LODs (that may be the same or different) that are lower than the LOD of the lifting operation iteration.
[0111] The prediction operation determines (e.g., computes) a prediction for the odd samples based on the even samples. For example, the prediction may be subtracted from the odd samples (e.g., shown as circles with negative signs) to create / generate a prediction error, e.g., error signal dk(k e [0, j — 1]). Forward lifting scheme 802 also includes an update operation (e.g., an update step shown as “U” block / step) that recalibrates the low-frequency signals (e.g., corresponding to signals at lower LODs) with some of the energy removed during the subsampling. In the case of classical lifting, this is used to prepare the even signals for the next prediction operation in the next iteration of forward lifting scheme 802. For example, the update operation updates (e.g., prepares) the even signals based on the error signal dkrepresenting a difference between odd sample soddkand a corresponding predicted odd sample. In some examples, the update operation may update the even signal sevenkbased on adding the prediction error dkto each of the even signal sevenk(e.g., shown as circle with positive signs). In some examples, the prediction error dkmay be adjusted by an update weight, and the even signal may be updated based on the adjusted prediction error. For example, the prediction error dkmay be scaled by the update weight before being added to each of the even signals. In some examples, the update weight may be a scaling value determined based on an index indicating / identifying the LOD of the odd sample.
[0112] In some embodiments, a decoder performs inverse lifting scheme 804 to reverse the operations of forward lifting scheme 802. For example, whereas forward lifting scheme 802 comprises lifting operations that are iteratively performed from higher LODs (e.g., LODN 810) to lower LODs (e.g., LODo 816), inverse lifting scheme 804 comprises lifting operations that are iteratively performed from lower LODs (e.g., LODo 816) to higher LODs (e.g., LODN 810). Each iteration of inverse lifting scheme 804 (e.g., four iterationsare shown as four dotted boxes corresponding to LODs 810-816) includes an update operation (e.g., an update step shown as a “U” component), a prediction operation (e.g., a prediction step shown as a “P” component), and a merge operation (e.g., a merge step shown as a “Merge” component).
[0113] Different from forward lifting scheme 802, an update operation, in each lifting operation of inverse lifting scheme 804, may update the even signals sk(e.g., corresponding to transformed displacement coefficients) by subtracting prediction error dk(corresponding to odd signals at the LOD corresponding to the lifting operation iteration) from the even samples to determine the updated even samples sevenk. In some examples, the prediction error dkmay be adjusted by an update weight, and the even samples may be updated based on the adjusted prediction error. For example, the update weight corresponding to an LOD of an odd signal may be determined. In some examples, the prediction error dkmay be scaled by the update weight before being subtracted from each of the even signals. In some examples, the update weight may be a scaling value determined based on an index indicating / identifying the LOD of the odd sample. In some examples, the update weight may be indicated by an indication signaled in the bitstream by an encoder and obtained from the bitstream by the decoder. For example, an update weight may be signaled or indicated for each respective LOD of the LODs.
[0114] A prediction operation, in each lifting operation of inverse lifting scheme 804, may determine a reconstructed odd sample soddkbased on the updated even samples seVenk. For example, the prediction operation may determine a linear combination (e.g., an average) of two updated even samples sevenkto determine a prediction of the reconstructed odd sample. In some examples, two updated even samples sevenkmay be combined as a weighted average using respective weights. Each lifting operation of inverse lifting scheme 804 combines (e.g., shown as circles with positive signs) the prediction error dkcorresponding to the odd sample with the prediction of the reconstructed odd sample to determine (e.g., reconstruct) a displacement signal soddkcorresponding to a displacement value determined at the encoder. In other words, the plurality of iterations of inverse lifting scheme 804 converts the wavelet coefficients (representing displacements), generated by the encoder and representing displacement information, into displacement values that may be used to reconstruct the (3D) mesh frame. Further, to revert the splitting operation of forward lifting scheme 802, each lifting operation of inverse lifting scheme 804 includes a merge operation that merges (e.g., orders or combines as a sequence of signals or values) the updated even samples sevenkwith the reconstructed odd sample soddk.
[0115] Note that the value j in FIG. 8 corresponds to a number of iterations for the lifting operations which varies depending on the specific requirement of the application for 3D meshes. For example, the number of levels in LOD defined by the mesh decimation process may be used for the lifting operations. In some examples, a mid-point subdivision scheme may be used in the mesh decimation process. In these examples, since each vertex in a higher LOD level is a generated mid-point of an edge defined by twovertices in lower LOD levels, the signal (e.g displacement value or its wavelet coefficient representation) associated with that vertex may be decomposed and represented by two sub-signals (e.g., displacement values or their wavelet coefficient representations) which belong to the corresponding two vertices. For example, a vertex v in LODi (e.g., an LOD of level 1) may be the mid-point of the edge defined by two vertices v1 and v2 in LODo (e.g., an LOD of level 0). In this example, the displacement associated with v can be wavelet transformed by using the lifting scheme. For an odd signal soddkcorresponding to vertex v (e.g., the signal being the displacement signal or its wavelet coefficient representation), the even samples sevenkdetermined for odd signal soddkmay correspond to vertices v1 and v2 (e.g., the signals being displacement signals or their wavelet coefficient representations) from which vertex v was generated.
[0116] In the lifting scheme, prediction weight and update weight are the coefficient values used to modify the input data during the prediction and update steps, respectively. The prediction weight may be a scalar value or a set of coefficients that define the linear combination of the neighboring signals used for prediction while the update weight determines the contribution of the prediction error to the final updated value. For example, the prediction may be determined from two input even samples based on a prediction weight equal to one half, which effectively averages signal values of the two input even samples. The prediction and update weights are often selected to satisfy certain properties or conditions to achieve desired characteristics in the transformed data. For example, in lossless lifting schemes, the weights may be selected to ensure perfect reconstruction of the original signal. In lossy lifting schemes, the weights may be selected to achieve specific frequency response characteristics or to minimize distortion based on the compression or denoising requirements.
[0117] In various implementations of 3D mesh coding, the prediction weight and the update weight may be determined (e.g., selected) for the lifting scheme, applied to displacements for vertices of a 3D mesh (e.g., each mesh frame of a sequence of mesh frames), such as to balance accuracy and properties resulting from the wavelet transforms corresponding to the displacements. As explained above, prediction operations of each iteration of the inverse lifting scheme may be dependent on (e.g., impacted by) updated signals inputs to the prediction operation However, the update weight may be a value (e.g., 1 / 8, 1 / 4, or 1 / 16, etc.) selected to be uniformly applied to wavelet coefficients corresponding (e.g., representing) the displacements. Due to characteristics and geometry of the mesh frame, characteristics at each LOD may not be the same. Therefore, applying the same update weight may results in reduced compression for displacements (e.g., displacement signals) for vertices at certain LODs.
[0118] In some embodiments, adaptive update weights in the lifting scheme are applied to displacements for vertices of 3D meshes (e.g., mesh frames of a sequence of mesh frames of a 3D mesh). For example, an update weight for each wavelet coefficient may be determined based on an LOD associated with that wavelet coefficient. As explained above, the lifting scheme may include a plurality of lifting operationscorresponding to a plurality of LODs in the 3D mesh (e.g., mesh frame). For a forward lifting scheme, each iteration of the lifting operation may update (e.g., lift) a sequence of displacement signals (e.g., displacement values or corresponding wavelet coefficients representing the displacement values) from a higher LOD (e.g., denser vertices) to one or more lower LODs (e.g., sparser vertices) and accumulate the prediction towards vertices at the lowest LOD (e.g., vertices of the base mesh). Similarly, but reciprocally, for an inverse lifting scheme, each iteration of the lifting operation may update (e.g., lift) a sequence of displacement signals (e.g., displacement values or corresponding wavelet coefficients representing the displacement values) from lower LOD (e.g., sparser vertices) to higher LODs (e.g., denser vertices). Since the update weight determines the amount of contribution of the prediction error to the final updated value, adapting uniform weight values to consider the impact of different LOD levels may result in more accurate predicting signals across different LOD levels. In some examples, lower LODs may be associated with smaller update weights and higher LODs may be associated with larger update weights. In some examples, lower LODs may be associated with larger update weights and higher LODs may be associated with smaller update weights.
[0119] In some embodiments, a decoder obtains from a bitstream transformed coefficients representing displacements of vertices of a three-dimensional (3D) mesh. After inverse quantizing the obtained transformed coefficients, the decoder applies a lifting wavelet transform scheme, such as that described above in FIG. 8, to obtain inverse-transformed wavelet coefficients representing the reconstructed displacements (e.g., reconstructed displacement vectors). As described with respect to FIGS. 5-6, the decoder and encoder may identically generate a subdivided mesh from a base mesh (e.g., down-sampled surface 520 of FIGS. 5-6). The decoder may reconstruct the 3D mesh by adding the displacements to respective vertices (e.g., positions of vertices) of the subdivided mesh.
[0120] In various implementations of 3D mesh coding, the subdivided mesh may be generated by applying a mid-point subdivision scheme to the base mesh, as described with respect to FIGS. 4-6. The base mesh, with many less vertices, was derived from an original 3D mesh as being representative of the 3D mesh. However, the additional vertices generated by the mid-point subdivision scheme are not necessarily close to or representative of the surface of the original 3D mesh (or a deformed mesh representative of the original 3D mesh). This is because the mid-point subdivision scheme subdivides the faces and edges of the base mesh into additional vertices, edges, and faces that are part of the surfaces of the base mesh.
[0121] In some implementations, a normal-based subdivision scheme may be used instead to subdivide the base mesh. In this normal-based subdivision scheme, each vertex generated by subdividing an edge of an input mesh (e.g., the base mesh or a subdivided mesh) may be generated by refining (e.g., adjusting a position of) the mid-point point according to vertex normals of vertices forming the edge. By doing so, each iteration of subdividing the base mesh and subsequent subdivided base meshes maygenerate new vertices that are not on the same plane as the surface of the base mesh. Accordingly, the final subdivided mesh may be closer to the original 3D mesh (or a deformed mesh representative of the original 3D mesh) and result in smaller displacements being signaled in the bitstream.
[0122] In some embodiments, the vertex normals of the vertices are normalized vertex normals. For example, a vertex normal of a vertex may be normalized by dividing the vertex normal by a length of the vertex normal. The vertex normal may be a unit vertex normal. Reference to vertex normals of vertices in the present disclosure may refer to normalized vertex normals (e.g., unit vertex normals).
[0123] FIG. 9A illustrates a diagram 900A of an example normal-based subdivision scheme, according to some embodiments. The normal-based subdivision scheme may be performed identically at the encoder and the decoder such that the same subdivided mesh may be generated from a base mesh. For example, the encoder may encode, in a bitstream, information indicating a geometry of the base mesh. For example, the decoder may obtain the base mesh based on decoding, from the bitstream, the information indicating the geometry of the base mesh. The geometry of the base mesh may include vertices (e.g., positions of vertices) of the base mesh, edges formed by pairs of the vertices, and triangles formed by a triple of edges of the edges.
[0124] Diagram 900A shows vertices 902A-D of a mesh edge / surface 940. Vertices 902A-D may be at a first LOD (e.g., level 0) and correspond, e.g., to vertices of a base mesh. It should be noted that for illustration purposes, the mesh surface 940 is represented as lines in diagram 900A. Diagram 900A also shows initial vertices 906A-C that would have been generated by a mid-point subdivision scheme in which each edge of mesh surface 940 is subdivided to generate edges and surfaces of a subdivided mesh at a next LOD (e.g., level 1). Subdividing the edges also results in subdivided surfaces of mesh edge / surface 940.
[0125] In some examples, vertex normals 912A-D are computed for respective vertices 902A-D of mesh edge / surface 940. In some examples, a vertex normal of a vertex may be determined based on combining (e.g., averaging) face normals (e.g., surface normals) of faces (e.g., surfaces or triangles)) containing that vertex. For example, a face normal of a face of mesh surface 940 may be determined based on a cross product of two edges forming the face. In some examples, the vertex normal may be a normalized vector resulting from normalizing the combination of the face normals. For example, the vertex normal may be a unit vertex normal.
[0126] As shown, the normal-based subdivision scheme may subdivide edge 920 of mesh surface 940 to generate a vertex 904 and subdivided edges 922-924 of subdivided edge / surface 942 at the next LOD. In some examples, an initial vertex 906A (q‘) along edge 920 may be adjusted (e.g., displaced) by a vector 908 (e.g., refinement vector v^) to determine vertex 904 (q). For example, vertex 904 may result from vector 908 added to initial vertex 906A. In some examples, initial vertex 906A may be a midpoint point along edge 920 resulting from averaging vertices 902A and 902B.
[0127] In some examples, vector 908 may be generated based on obtaining (e.g., selecting) vertices 930 (e.g., vertices 902A and 902B) forming edge 920 that is being subdivided. For example, vector 908 may be generated based on vertex normals 912A-B of vertices 902A-B (e.g., labeled vertices a and b) forming edge 920. For example, vector 908 may be obtained based on combining vertex normals 912A-B of vertices 902A-B, respectively.
[0128] FIG. 9B illustrates a diagram 900B of an example normal-based subdivision scheme and that provides more details for how vertex 904 is generated based on subdividing edge 920 formed by vertices 902A-B, according to some embodiments. As shown between diagrams 900A-B, like components have the same labels.
[0129] In some examples, vector 908 (v^) may be a linear combination of vertex normals 912A-B (n^ and n^). For example, each vertex normal of the vertex normals 912A-B may be weighted by a respective normal weight (daand db) that is based on the edge and the vertex normal. For example, vertex 904 (q) may be determined according to equation (3) as follows:In an example, initial vertex q’ may be determined as follows: (v + p).
[0130] In some embodiments, a weight (d, where i represents a specific vertex) for a vertex normal (fQ may be determined (e.g., derived) as a dot product between the subdivided edge 920 and the vertex normal. For example, the weights daand db for vertex normals 912A-B (n^ and n ) may be determined according to the dot products shown in equations (4) and (5) as follows:In some examples, the dot product in each of Eq. 4 and 5 may be multiplied / scaled by a constant c, which may be, for example,1 / 2. In an example, the constant c is omitted or equal to 1 .
[0131] In some examples, for deriving a normal weight for a vertex, the edge may be converted to a vector defined from a further vertex (forming the edge) to a closer vertex (forming the edge) with respect to the vertex. Accordingly, the normal weights may be based on the length of edge 920 and angles between edge 920 and each of the vertex normals 912A-B.
[0132] Regarding Eq. 4, a dot product between a first vector (e.g.,from vertex 902B to vertex 902A) and vertex normal 912A results in a first value between -1 and 1 (inclusive) that corresponds to an angle 950A between the first vector and vertex normal 912A. Similarly, regarding Eq. 5, a dot product between a second vector (e.g.,from vertex 902A to vertex 902B) and vertex normal 912B results in a second value between -1 and 1 (inclusive) that corresponds to an angle 950B between the second vector and vertex normal 912B. Each of the first value and the second value represents a directional alignment between the first vector and vertex normal 912A and between the second vector and vertex normal 912B,respectively. For example, a larger value of the first value (or second value) represents angle 950A (or angle 950B) being smaller.
[0133] A value resulting from the dot product between two vectors being positive (i.e. , greater than zero) represents an angle between the two vectors is acute (e.g., between 0 and 90 degrees), meaning the two vectors are generally pointing in the same direction. For example, the first value being 1 represents angle 950A being zero or the first vector and vertex normal 912A being co-linear and pointing in the same direction. A value resulting from the dot product between two vectors being zero represents the angle between the two vectors is 90 degrees, meaning the two vectors are perpendicular (orthogonal) to each other. For example, the first value being 0 represents angle 950A being 90 degrees or the first vector and vertex normal 912A being orthogonal. A value resulting from the dot product between two vectors being negative (e.g., less than zero) represents an angle between the two vectors is obtuse (e g., between 90 and 180 degrees), meaning the two vectors are pointing in generally opposite directions. For example, the first value being -1 represents angle 950A being 180 degrees or the first vector and vertex normal 912A being co-linear and pointing in the opposite direction.
[0134] In some embodiments, instead of computing the dot products in equations (4) and (5), the angles 950A and 950B may be computed directly.
[0135] In some embodiments, the linear combination of equation (3) further includes a vector weight w, that weights vector 908. The vector weight may control a smoothness of the subdivision surface to reduce or increase the adjustment. In some examples, vector weight w, may be based on an LOD (level i) of vertex 904. For example, the vector weight Wj may be determined according to equation (6) as follows, where T may represent the highest LOD level:Wj = pow(2, i — T) (Eq 6)
[0136] Accordingly, by adjusting or refining a position of initial vertex 906A to that of vertex 904, the subdivided vertices and surface may be spatially closer to the original / deformed mesh edge / surface 944 (e.g., corresponding to deformed mesh 436 of FIG. 4). For example, vertex 904 is closer to original mesh vertex 936 than initial vertex 906A and similarly edges 922 and 924 are closer to corresponding mesh edges 932 and 934 of deformed mesh edge / surface 944. By generating a subdivided mesh with vertices that are closer to the original / deformed mesh, the displacements computed by the encoder and signaled in the bitstream to the decoder are reduced, which results in increased compression.
[0137] In some examples, as shown in Eqs. 3-5, because the normal-based subdivision scheme described above with respect to FIGS. 9A-B adjust an initial vertex 906 by a (refinement) vector 908 derived using dot products to weight vertex normals in a linear combination, signs of the weights daand db for vertices 902A (vertex A) and 902B (vertex B) may be misaligned. For example, vertex normals 912A and 912B may be generally in the same direction, but the respective weights daand db computed for vertex normals 912A and 912B, respectively, may in certain scenarios have opposite signs. In thesescenarios, the impact of the vertex normals 912A-912B in the linear combination of Eq. 3 may be reduced, which results in vertex 904 being moved further away from original mesh vertex 936 after refinement. Increasing the spatial distance between vertex 904 and original mesh vertex 936 results in larger displacements that will be signaled in a bitstream and leads to increase bitrate.
[0138] Embodiments described herein relate to determining a directional alignment between vertex normals used to subdivide an edge to generate a vertex and adjusting (e.g., refining) a position of the vertex in accordance with the determined directional alignment. In some examples, as part of subdividing a mesh to determine a subdivided mesh, each edge of the mesh may be subdivided to generate a vertex and associated edges and triangles of the subdivided mesh. For an edge formed by a first vertex and a second vertex, the vertex may be determined based on displacing an initial vertex (on the edge) by a vector (e.g., refinement vector) determined based on a first and a second vertex normal of the first and second vertices, respectively, and the directional alignment between the first and second vertex normals.
[0139] In some examples, the initial vertex may be an average of the first and second vertices, which is equivalent to a mid-point point on the edge.
[0140] In some examples the vector may be based on a linear combination of the first vertex normal and the second vertex normal with respective first and second weights. In some embodiments, the first weight may be determined based on a dot product between the first vertex normal and a first vector indicated by a displacement from the second vertex to the first vertex. Similarly, the second weight may be determined based on a dot product between the second vertex normal and a second vector indicated by a displacement from the first vertex to the second vertex.
[0141] In some embodiments, the first and / or the second weights may be adjusted depending on (e.g., based on) the directional alignment between the first and the second vertex normal and / or based on signs of the first and second weights. For example, the directional alignment may indicate whether the first and the second vertex normals are generally pointing in the same direction (e.g., an angle between the first and second vertex normals is acute), orthogonal, or generally pointing in the opposite direction (e.g., an angle between the first and second vertex normals is obtuse). In some examples, the directional alignment may be determined based on (e.g., as being equal to) a dot product between the first and second vertex normals. As explained above, a positive value of the dot product indicates the first and second vertex normals are generally pointing in the same direction, a zero value of the dot product indicates the first and second vertex normals are orthogonal, and a negative value of the dot product indicates the first and second vertex normals are generally pointing in the opposite direction.
[0142] In some examples, an adjustment to a weight (d) of the first and second weights may include setting the weight to 0 (e.g., d = 0).
[0143] In some examples, an adjustment to a weight (d) of the first and second weights may include scaling the weight by a fixed value (e.g., d = d *0.5). In an example, the fixed value may be determined based on an LOD of the vertex and / or the LOD of the first and / or second vertices.
[0144] In some examples, an adjustment to a weight (d) of the first and second weights may include adding an offset to the weight (e.g., d = d + offset). For example, offset may be a fixed value. For example, the offset may be determined based on an LOD of the vertex and / or the LOD of the first and / or second vertices.
[0145] In some examples, an adjustment to a weight (d) of the first and second weights may include inverting a sign of the weight (e.g., weight d = -1 *d).
[0146] In some examples, an adjustment to a weight (d) of the first and second weights may include a combination of one or more adjustments including, for example, scaling the weight, adding an offset to the weight, and / or inverting a sign of the weight.
[0147] In some examples, a first indication of enabling or disabling the weight adjustment in the normalbased subdivision scheme may be signaled in or obtained from the bitstream.
[0148] In some examples, a second indication of one of a plurality of adjustments may be signaled in or obtained from the bitstream. For example, the plurality of adjustments may include sign inversion, scaling, adding an offset, skipping or setting to zero, etc.
[0149] For example, the first and / or second indication may be signaled by an encoder into the bitstream and obtained from the bitstream by a decoder.
[0150] FIG. 9C illustrates a diagram 900C of an example normal-based subdivision scheme with a reference vector 962, according to some embodiments. For example, this normal-based subdivision scheme generates vertex 960 (q'”) based on projecting vector 908 (q'), described above with respect to FIG. 9B, onto a reference vector 962 (nave). In other words, vertex 960 resulting from subdividing edge 920 is along the direction of reference vector 962.
[0151] In some examples, reference vector 962 (nave’) may be determined based on a linear combination of vertex normals 912A-B (n^ and n^). For example, reference vector 962 (nave’) may be an average of the vertex normals 912A-B shown as follows:
[0152] In some examples, the reference vector 962 (nave’) may be a unit vector obtained by, for example, dividing the reference vector 962 by a length of the reference vector 962.
[0153] In some examples, vertex 960 (q"’) may be determined based on projecting vertex 904 onto reference vector 962. For example, vector 964 (v^) may be determined based on multiplying reference vector 962 (nave’) to a dot product of vector 908 and reference vector 962 as follows:
[0154] Then, vertex 960 may be determined by adding vector 964 to initial vertex 906A.
[0155] FIG. 9D illustrates a diagram 900D of an example normal-based subdivision scheme with vertex adjustments, according to some embodiments. For example, edge 920 may be subdivided to determine vertex 970. For example, vertex 970 may be initial vertex 906A of FIG. 9A representing a midpoint of edge 920, vertex 904 of FIG. 9B representing an adjusted initial vertex 906A, or vertex 960 of FIG. 9C representing an adjusted initial vertex 906A.
[0156] In some embodiments, as part of subdividing edge 920 formed by vertices 902A and 902B, positions of vertices 902A and 902B may be adjusted in the opposite direction of vertex normals 912A and 912B of respective vertices 902A and 902B. For example, vertex 902A-B may be adjusted along vertex normals 972A-B to vertex 974A-B, respectively. As shown in FIG. 9D, vertex normals 972A-B have the opposite signs as vertex normals 912A-B, respectively.
[0157] In some examples, the vertex adjustments may be performed on vertices of the base mesh. For example, the determination to adjust a vertex of vertices 902A-902B may be based the vertex being of the base mesh, for example, of the lowest LOD (e.g., with LOD index 0).
[0158] In some examples, a vertex such as vertex 902A (a) may be adjusted along vertex normal 972A (- r . For example, vertex 972A may be determined by adding vertex 902A to a scaled vertex normal 972A. For example, the scaling may be based on a product of a weight (weight) and an average (dave) of vertex weights (dn) of vertices connected to vertex 902A as shown below: dn = dave a’ = a -
[0159] For example, the average vertex weight (dn) may be considered as a valence weight value because a valence of the vertex 902A indicates the number of vertices connected to vertex 902A.
[0160] In some examples, each vertex weight (dn) of a connected vertex may be determined as a dot product of vertex normal 912A and a vector associated with the edge formed by vertex 902A and the connected vertex. For example, a vertex weight for the connected vertex being vertex 902B may be determined as one half of the vector represented as a displacement from vertex 902A to vertex 902B, as shown above.
[0161] In some examples, the weight (weight) may be a fixed value. In this case, the weight may be a predetermined value and is not signaled in a bitstream.
[0162] In some examples, the weight (weight) may be a value signaled by the encoder in the bitstream. The decoder may obtain the value of the weight from the bitstream.
[0163] FIGS. 10A-C illustrate diagrams 1000A-C showing different combinations of signs of the first and second weights (i ,e. , weight dafor vertex normal rFj of vertex a and weight db for vertex normal ofvertex b) corresponding to the first and second vertex normals (i.e., vertex normals na' and n^) of a first and second vertex (i.e., vertices a and b) of an edge that is subdivided during a mesh subdivision process, according to some embodiments. Diagrams 1000A-C show examples of when a directional alignment between the first and second vertex normals (i.e., vertex normalsand n ) indicate that they are generally in the same directional. For example, the directional alignment may be determined based on a dot product between the first and second vertex normals and the first and second vertex normals generally being in the same directional may be indicated by the dot product being positive.
[0164] FIG. 10A shows vector 908 used to adjust initial vertex 906A to determine vertex 904. Vector 908 may be determined based on a linear combination of vertex normals 912A and 912B where the signs of each of the weights are positive as represented by angles 950A-B that are each acute (e.g., less than 90 degrees).
[0165] In this example, the combination of the positive signs and the directional alignment between the first and second normals being in alignment may indicate that the surface represented by the edge may be a smooth surface (e.g., convex surface). Thus, less or no adjustment may be needed. For example, the default weight determinations of Eqs. 4 and 5 may be used.
[0166] FIG. 10B shows vector 1008A used to adjust initial vertex q' to determine vertex 1004A. Vector 1008A may be determined based on a linear combination of vertex normals 1002A and 1002B where the sign of a first weight of vertex normal 1002A is positive (e.g., da> 0 indicated by angle 1000A being acute) and the sign of a second weight of vertex normal 1002B is negative (e.g., db < 0 indicated by angle 1000B being obtuse). Due to the conflicting signs, vector 1008A adjusts initial q' to vertex 1004A that is further from a corresponding original mesh vertex q”, which results in a larger displacement between subdivided vertex 1004A and vertex q”.
[0167] In this example, the combination of opposite signs and the directional alignment between the first and second normals being in alignment may indicate that the surface represented by the edge may have a slight bend or minor inversion in curvature. For example, one or more of the above adjustments may be applied to one or both weights of the first and second vertex normals in Eqs. 4 and 5. For example, the adjustment may include inverting the sign of the second weight such that both signs of the first and second weights are positive.
[0168] It should be noted that conflicting signs may also occur when the sign of the first weight of vertex normal 1002A is negative and the sign of the second weight of vertex normal 1002B is positive.Accordingly, there are two possible cases of the conflicting signs. FIG. 10B shows one of these two possible cases.
[0169] FIG. 10C shows vector 1018A used to adjust initial vertex q’ to determine vertex 1014A. Vector 1018A may be determined based on a linear combination of vertex normals 1012A and 1012B where the sign of a first weight of vertex normal 1012A is negative (e.g., da< 0 indicated by angle 1010A beingobtuse) and the sign of a second weight of vertex normal 1012B is negative (e.g., db < 0 indicated by angle 101 OB being obtuse).
[0170] In this example, the combination of the negative signs and the directional alignment between the first and second normals being in alignment may indicate that the surface represented by the edge may a smooth concave surface. Thus, less or no adjustment may be needed because vector 1018A adjusts vertex q’ towards vertex q” which is positioned inwards according to the concave surface.
[0171] FIGS. 11A-C illustrate diagrams 1100A-C showing different combinations of signs of the first and second weights (i ,e. , weight dafor vertex normal rFj of vertex a and weight db for vertex normal nj^ of vertex b) corresponding to the first and second vertex normals (i.e., vertex normalsand n ) of a first and second vertex (i.e., vertices a and b) of an edge that is subdivided during a mesh subdivision process, according to some embodiments. Diagrams 1100A-C show examples of when a directional alignment between the first and second vertex normals (i.e., vertex normalsand n ) indicate that they are generally in the opposite directional. For example, the directional alignment may be determined based on a dot product between the first and second vertex normals and the first and second vertex normals generally being in the opposite directional may be indicated by the dot product being negative.
[0172] FIG. 11A shows vector 1108A used to adjust initial vertex q' to determine vertex 1104A. Vector 1108A may be determined based on a linear combination of vertex normals 1102A and 1102B where the signs of each of the weights are positive as represented by angles 1100A-B that are each acute (e.g., less than 90 degrees).
[0173] In this example, the combination of the positive signs and the directional alignment between the first and second normals being not in alignment may indicate that the surface represented by the edge may be a sharp edge or fold. In some examples, one or more of the above adjustments may be applied to one or both weights of the first and second vertex normals in Eqs. 4 and 5. For example, the first and second weights may each be scaled (e.g., by1 / 2) to reduce the impact of the vector normals 1112A-B, which results in vertex 1114A moving closer towards initial vector q', which is also closer to vector q”.
[0174] FIG. 11 B shows vector 1118A used to adjust initial vertex q’ to determine vertex 1114A. Vector 1118A may be determined based on a linear combination of vertex normals 1112A and 1112B where the sign of a first weight of vertex normal 1112A is positive (e.g., da> 0 indicated by angle 1110A being acute) and the sign of a second weight of vertex normal 1112B is negative (e.g., db < 0 indicated by angle 1110B being obtuse). Due to the conflicting signs, vector 1118A adjusts initial q’ to vertex 1114A that is further from a corresponding original mesh vertex q”, which results in a larger displacement between subdivided vertex 1114A and vertex q”.
[0175] In this example, the combination of opposite signs and the directional alignment between the first and second normals being not in alignment may indicate that the surface represented by the edge may have a sharp feature or edge / bend in curvature. Accordingly, one or more of the above adjustments maybe applied to one or both weights of the first and second vertex normals in Eqs. 4 and 5. For example, the adjustment may include setting the second weight with the negative sign to a fixed value. For example, the fixed value may be zero. Accordingly, vector 11 18A will be shortened and result in vertex 11 14A moving closer towards vertex q', which is also closer to vertex q”.
[0176] It should be noted that conflicting signs may also occur when the sign of the first weight of vertex normal 1112A is negative and the sign of the second weight of vertex normal 1 112B is positive. Accordingly, there are two possible cases of the conflicting signs. FIG. 11 B shows one of these two possible cases.
[0177] FIG. 1 1 C shows vector 1 128A used to adjust initial vertex q’ to determine vertex 1124A. Vector 1 128A may be determined based on a linear combination of vertex normals 1122A and 1122B where the sign of a first weight of vertex normal 1122A is negative (e.g , da< 0 indicated by angle 1120A being obtuse) and the sign of a second weight of vertex normal 1212B is negative (e.g., db < 0 indicated by angle 1 120B being obtuse).
[0178] In this example, the combination of the negative signs and the directional alignment between the first and second normals being not in alignment may indicate that the surface represented by the edge may a sharp inward fold. Thus, vector 1128A may over fit the edge and adjust initial vertex q1to be further from vertex q” of original / deformed mesh. In some examples, one or more of the above adjustments may be applied to one or both weights of the first and second vertex normals in Eqs. 4 and 5. For example, the first and second weights may each be scaled (e.g., by >2) to reduce the impact of the vector normals 1122A-B, which results in vertex 1124A moving closer towards initial vector q', which is also closer to vector q”.
[0179] The following table summarizes the combinations of the signs of the first and second weights for the first and second vertex normals, respectively, and a directional alignment of the first and second vertex normals:
[0180] The above table also shows, for each combination of signs of the first and second weights and directional alignment between the first and second vertex normals, an example adjustment to one or more both of the first and second weights for the first and second vertex normals. It should be understood that other types of adjustments may be possible or a combination of the various adjustments described above, including, but not limited, to adding an offset, scaling the weight, inverting a sign of the weight, setting the weight to a fixed value, or a combination thereof.
[0181] FIG. 12 illustrates a diagram 1200 showing two example adjustments 1202A-B to one or both of the first and second weights of vertex normals 1112A-B for the scenario shown in FIG. 11 B.
[0182] For example, a possible adjustment 1202A may include scaling vertex normal 1112B associated with a negative weight by a fixed value (e.g. , V2} to reduce the impact of vertex normal 1112B in determining vector 1118A. Accordingly, the scaled vertex normal 1112B, shown as vertex normal 1208, may be used in Eq. 5 to generate a vector 1206A that is smaller and results in vertex 1204A that is closer to vertex q”.
[0183] For example, a possible adjustment 1202B may include setting the second weight of vertex normal 1112B to a fixed value such as zero as shown by circle 1210 to reduce the impact of vertex normal 1112B in determining vector 1118A. Accordingly, in the example shown in FIG. 12, a vector 1206B is generatedbased on vertex normal 1112A and not on vector normal 1112B, which results in reducing vector 1118A such that the adjusted vertex 1204B may be closer to vertex q”.
[0184] FIG. 13 illustrates a flowchart of a method 1300 for performing normal-based subdivision for reconstructing (e.g., decoding) a 3D mesh, according to some embodiments. In some examples, method 1300 may be performed by a decoder (e.g., decoder 120 of FIG. 1 or decoder 300 of FIG. 3). The following descriptions of various steps may refer to operations performed by deformed mesh reconstructor 316 of FIG. 3 and described with respect to FIGS. 9-12.
[0185] At block 1302, a decoder obtains, from a bitstream, a base mesh for a 3D mesh.
[0186] At block 1304, the decoder subdivides a first mesh associated with the base mesh. For example, the first mesh may be the base mesh or the first mesh may be a result of one or more iterations of a subdivision scheme applied to the base mesh. The subdivision scheme may be a normal-based subdivision scheme, as described with respect to FIGS. 9-12.
[0187] In some embodiments, the decoder obtains, from the bitstream, a mode indication of a subdivision scheme from a plurality of subdivision schemes. The subdivision scheme may be associated with normal-based subdivision.
[0188] In some examples, the plurality of subdivision schemes may include a first subdivision scheme that subdivides each edge of a mesh according to vertex normals of a pair of vertices forming each edge.
[0189] In some examples, the plurality of subdivision schemes may include a second subdivision scheme that subdivides each edge of the mesh according to vertex normals of the pair of vertices and a directional alignment between the pair of vertices.
[0190] In some examples, the plurality of subdivision schemes may include a midpoint subdivision scheme.
[0191] As part of subdividing the first mesh, the decoder may subdivide edges of the first mesh, which may include blocks 1306-1312.
[0192] At block 1306, the decoder subdivides an edge, formed by a first vertex and a second vertex from the first mesh associated with the base mesh, to determine a vertex of the subdivided first mesh. The process of subdividing the edge may include blocks 1308-1312.
[0193] At block 1308, the decoder determines a directional alignment between a first vertex normal of the first vertex and a second vertex normal of the second vertex For example, as described above with respect to FIGS. 10-11 , the decoder may determine the directional alignment based on a dot product between the first and second vertex normals. For example, the dot product may represent the directional alignment such that the dot product being a positive value indicates alignment (e.g., generally pointing in the same direction) and the dot product being a negative value indicates misalignment (e.g., generally pointing in the opposite direction). For example, the dot product being zero may indicate the first and second vertex normals are perpendicular to each other.
[0194] At block 1310, the decoder determines a refinement vector based on the directional alignment, the first vertex normal, and the second vertex normal.
[0195] In some embodiments, the refinement vector is determined based on combining the first and second vertex normals in a linear combination where each of the first and second vertex normals are weighted (e.g., scaled) by a first weight and a second weight, respectively, as described above with respect to FIGS. 9A-B . For example, the first weight may be determined based on a dot product between the first vertex normal and a first vector indicated by a displacement from the second vertex to the first vertex. Similarly, the second weight may be determined based on a dot product between the second vertex normal and a second vector indicated by a displacement from the first vertex to the second vertex. In some embodiments, the refinement vector is further determined based on a combination of signs of the first and second weights.
[0196] In some embodiments, the refinement vector may be determined based on combining the first and second vertex normals, as described above with respect to FIGS. 9A-B, and further based on determining (e.g. deriving) a reference vector for the vertex, as described above with respect to FIG. 90.
[0197] In some embodiments, based on a directional misalignment between the first and second vertex normals or between signs of the first and second weights, one or both of the first and second weights may be adjusted. For example, various types of adjustments are described above with respect to FIGS. 10-12.
[0198] In some embodiments, the refinement vector (e.g., represented by any of the linear combinations described above) may be weighted by a vector weight. For example, the vector weight may be based on a level of detail (LOD) of the vertex. In some examples, the decoder may obtain, from the bitstream, a vector weight for each LOD of a plurality of LODs of the 3D mesh. For example, the bitstream may include subdivision information encoding the vector weight for each LOD.
[0199] In some embodiments, the directional alignment may be used as a condition for enabling (e.g., activating or permitting) a determination of the refinement vector as described above with respect to FIGS. 9A-B. For example, the determination of the refinement vector being enabled (or used) based on the directional alignment indicating that the first vertex normal and the second vertex normal are aligned, according to some examples.
[0200] In some embodiments, the directional alignment may be used as a condition for enabling (e.g., activating or permitting) a determination of the refinement vector using a reference vector in a normalbased subdivision scheme described above with respect to FIG. 9C. For example, the directional alignment may be used as a condition for whether to generate the reference vector to determine the refinement vector.
[0201] In some embodiments, the determination of the refinement vector being enabled may be based on the directional alignment and signs of a first weight and a second weight computed / derived for the first and second vertex normals, respectively. For example, there may be eight possible combinations of thesigns of the first and second weights and the directional alignment being aligned or misaligned. FIGS. 10A-C and FIGS. 11A-C show six possible combinations. The determination of the refinement vector being enabled may be based on one or more predetermined combinations, from the eight possible combinations, being satisfied.
[0202] In some examples, the first and second weights may be determined for the first and second vertex normals based on a vector represented using the first vertex and the second vertex, as explained above with respect to FIG. 9A and FIG. 9B.
[0203] In some examples, the refinement vector may be determined (e.g., enabled to be used) based on (e.g., in response to) the dot product between the first and second vertex normals being positive.Otherwise, the refinement vector is disabled or not determined / used based on (e.g., in response to) the dot product between the first and second vertex normals being negative or zero.
[0204] In some examples, the refinement vector may be determined (e.g., enabled to be used) based on the dot product between the first and second vertex normals being positive or zero. Otherwise, the refinement vector is disabled or not determined / used based on (e.g., in response to) the dot product between the first and second vertex normals being negative.
[0205] In some embodiments, based on the refinement vector being disabled or not used (e.g., the directional alignment indicating the first and second vertex normals are not aligned), the refinement vector is not determined and block 1310 may be omitted or skipped.
[0206] In some embodiments, based on the refinement vector being disabled or not used (e.g., the directional alignment indicating the first and second vertex normals are not aligned), the refinement vector is set to zero. Setting the refinement vector to zero at block 1310 may be logically equivalent to disabling or not using the refinement vector.
[0207] At block 1312, the decoder determines the vertex based on the refinement vector and a point along the edge. In some examples, the point may be a midpoint point of the edge. For example, the midpoint point may be determined based on averaging (the values) of the pair of vertices forming the edge.
[0208] In some examples, the vertex is determined based on adding the refinement vector to the point. For example, the refinement vector indicates a displacement from the point to a position of the vertex.
[0209] In some embodiments in which the refinement vector is disabled or not used, e.g., based on the directional alignment indicating the first and second vertex normals are not aligned, the vertex may be determined as the point along the edge.
[0210] In some embodiments, as part of subdividing the edge, block 1306 may include block 1313 (shown in dotted lines to represent a possible enhancement), in which one or more of the first and second vertices are adjusted in the opposite direction of the first and second vertex normals, respectively. Examples of adjusting the first and second vertices are described above with respect to FIG. 9D.
[0211] In some embodiments, the condition for adjusting one or more of the first and second vertices may be based on the vertices being of a specific LOD (e.g., being in LOD 0 corresponding to vertices of the base mesh). In some examples, the enhancement for adjusting the vertices forming the edge may be signaled in an indication of whether to enable / activate the enhancement. For example, the decoder may obtain this indication from a bitstream, to which the encoder signals or encodes this indication.
[0212] In some embodiments, similar to the determination of the refinement vector being based on the directional alignment, the determining of whether to adjust the one or more of the first and second vertices may be based on the directional alignment determined at block 1308. For example, in some examples, block 1313 may be enabled based on the directional alignment indicating that the first and second vertex normals are aligned.
[0213] At block 1314, the decoder decodes (e.g., reconstructs) the 3D mesh based on a second mesh (e.g., subdivided mesh) comprising vertices of the first mesh and the vertex (e.g., resulting from subdividing the first mesh). For example, the second mesh may result from iteratively subdividing the first mesh, which may result from subdividing or iteratively subdividing the base mesh.
[0214] In some examples, reconstructing the 3D mesh includes obtaining, from the bitstream, displacements of second vertices of the second mesh. Then, the 3D mesh may be reconstructed based on updating the second vertices with the displacements. For example, each vertex of the second vertices may be added to a respective displacement vector of the displacements to determine a reconstructed vertex of vertices of the reconstructed 3D mesh.
[0215] FIG. 14 illustrates a flowchart of a method 1400 for performing normal-based subdivision for encoding a 3D mesh, according to some embodiments. In some examples, method 1400 may be performed by an encoder (e.g., encoder 114 of FIG. 1 , encoder 200A of FIG. 2A, or encoder 200B of FIG. 2B). The following descriptions of various steps may refer to operations performed by displacement generator 208 and / or deformed mesh reconstructor 230 of FIGS. 2A-B, and / or mesh subdivider 404 and / or mesh subdivider 418 of FIG. 4, and described with respect to FIGS. 9-12.
[0216] At block 1402, an encoder obtains a base mesh for a 3D mesh. In some examples, the encoder encodes vertices of the base mesh in a bitstream from which a decoder may obtain (e.g., decode) the vertices of the base mesh. In some examples, the encoder may apply a downsampling (or decimation) process to obtain the base mesh from the 3D mesh, as described with respect to FIG. 4.
[0217] At block 1404, the encoder subdivides a first mesh associated with the base mesh. Block 1404 may correspond to block 1304 of FIG. 13. For example, the first mesh may be the base mesh or the result of one or more iterations of a subdivision scheme applied to the base mesh. The subdivision scheme may be a normal-based subdivision scheme, as described with respect to FIGS. 10-11.
[0218] At block 1406, the encoder subdivides an edge, formed by a first vertex and a second vertex from the first mesh associated with the base mesh, to determine a vertex. The process of subdividing the edgemay include blocks 1408-1412. Blocks 1408-1412 may correspond to blocks 1308-1312 of FIG. 13, respectively.
[0219] At block 1408, the encoder determines a directional alignment between a first vertex normal of the first vertex and a second vertex normal of the second vertex.
[0220] At block 1410, the encoder determines a refinement vector based on the directional alignment, the first vertex normal, and the second vertex normal.
[0221] At block 1412, the encoder determines the vertex based on the refinement vector and a point along the edge.
[0222] In some embodiments, as part of subdividing the edge, block 1406 may include block 1413 (shown in dotted lines to represent a possible enhancement), in which one or more of the first and second vertices are adjusted in the opposite direction of the first and second vertex normals, respectively. Block 1413 may correspond to block 1313 of FIG. 13.
[0223] As described above, the encoder and the decoder may apply the same subdivision scheme (e.g., a normal-based subdivision scheme) to the base mesh such that vertices of the subdivided mesh need not be signaled in the bitstream by the encoder. Accordingly, operations of blocks 1404-1413 performed by the encoder may be the same as those performed by the decoder and thus correspond to blocks 1304- 1313 of FIG. 13, respectively.
[0224] At block 1414, the encoder encodes the 3D mesh based on a second mesh comprising vertices of the first mesh and the vertex. For example, the second mesh may represent a final subdivided mesh resulting from iteratively subdividing the base mesh For example, the second mesh may result from iteratively subdividing the first mesh, which may result from subdividing or iteratively subdividing the base mesh.
[0225] In some examples, encoding the 3D mesh includes signaling, in the bitstream, displacements of the 3D mesh based on the second mesh including vertices of the first subdivided mesh and the vertex. For example, the encoder may determine the displacements for vertices of the second mesh based on differences between (vertices of) the second mesh and (vertices of) a representation of the 3D mesh, as described with respect to FIGS. 4-6. For example, the representation may be a fitted subdivided mesh (e.g., deformed mesh 436) generated from an original 3D mesh.
[0226] FIG. 15 illustrates a flowchart of a method 1500 for performing normal-based subdivision for coding a 3D mesh, according to some embodiments. In some examples, method 1500 may be performed by a coder such as an encoder (e.g., encoder 114 of FIG. 1 , encoder 200A of FIG. 2A, or encoder 200B of FIG. 2B) or a decoder (e.g., decoder 120 of FIG. 1 or decoder 300 of FIG. 3). The decoder and the encoder may identically perform method 1500 to identically subdivide the first mesh. For example, the following descriptions of various steps may refer to operations performed by displacement generator 208 and / or deformed mesh reconstructor 230 of FIGS. 2A-B, and / or mesh subdivider 404 and / or meshsubdivider 418 of FIG. 4, and described with respect to FIGS. 9-12. For example, following descriptions of various steps may refer to operations performed by deformed mesh reconstructor 316 of FIG. 3 and described with respect to FIGS. 9-12.
[0227] At block 1504, a coder subdivides a first mesh associated with a base mesh for a 3D mesh. Block 1504 may correspond to block 1304 and / or block 1404. For example, block 1504 may represent one example implementation of block 1304 and / or block 1404 of FIGS. 13 and 14, respectively.
[0228] As explained above with respect to block 1304 and 1404, the first mesh may be derived based on subdividing the base mesh obtained for the 3D mesh. For example, the first mesh may be the result of subdividing the base mesh or iteratively subdividing the base mesh by a number of iterations. Accordingly, the first mesh may represent an iteratively subdivided base mesh.
[0229] As part of subdividing the first mesh, each edge of the first mesh may be subdivided to determine (e.g., derive) vertices of the subdivided first mesh. Block 1506 represents one subdivision operation of one edge of the first mesh. At block 1506, the coder subdivides an edge, formed by a first vertex and a second vertex from the first mesh associated with the base mesh, to determine a vertex of the subdivided first mesh. For example, block 1506 may correspond to block 1306 and block 1406 of FIGS. 13 and 14, respectively. Block 1506 may include blocks 1508-1512.
[0230] At block 1508, the coder determines a directional alignment between a first vertex normal of the first vertex and a second vertex normal of the second vertex. Block 1508 may correspond to and may be the same as blocks 1308 and / or 1408 of FIGS. 13 and 14, respectively.
[0231] At block 1510, the coder determines, at least based on the directional alignment, whether to refine a position of a point, along the edge, to determine the vertex. For example, the point may be a midpoint of the edge, as explained above with respect to blocks 1312 and 1412 of FIGS. 13 and 14, respectively. Block 1510 may be an example of determining whether the determination of the refinement vector of blocks 1310 and 1410 of FIGS. 13 and 14, respectively, is enabled / activated / used.
[0232] As explained above with respect to blocks 1310 and 1410, the position of the point may be determined to be refined based on, for example, the directional alignment indicating that the first and second vertex normals are aligned. For example, the first and second vertex normals being aligned may be determined based on a dot product of the first and second vertex normals being positive. As explained in blocks 1310 and 1410, the determination of whether to refine the position of the point may be further based on signs of a first and second weight derived for the first and second vertices, respectively.
[0233] In some embodiments, based on determining to refine the position of the point, a refinement vector may be determined based on the first and second vertex normals, as explained in blocks 1310 and 1410. Then, the vertex may be determined based on the position of the point being adjusted based on the refinement vector, e.g., based on a sum of the position of the point and the refinement vector.
[0234] In some examples, if enabled according to the determined directional alignment, the refinement vector may be determined as a linear combination of the first and second vertex normals, as described above with respect to FIGS. 9A-B.
[0235] In some examples, if enabled according to the determined directional alignment, the refinement vector may be determined further based on a reference vector determined from the first and second vertex normals, as described above with respect to FIG. 9C.
[0236] In some embodiments, based on determining not to refine the position of the point, the vector may be determined as the position of the point.
[0237] In some embodiments, block 1506 may include block 1512 (e.g., shown in dotted lines to indicate a possible embodiment), the coder determines, at least based on the directional alignment, whether to refine (e.g., adjust) positions of one or more of the first vertex and the second vertex. For example, one or more of the first and second vertices may be adjusted along the first and second vertex normals, respectively, in the opposite direction, as explained above with respect to FIG. 9D. For example, the operation of adjusting the positions of the first and / or the second vertices are also referenced at block 1313 of FIG. 13 and block 1413 of FIG. 14.
[0238] In some embodiments, based on the directional alignment indicating that the first and second vertex normals are aligned, the coder determines that the first and second vertices may be adjusted. For example, the coder further determines whether the first and / or second vertex satisfy one or more conditions before they are adjusted. For example, the one or more conditions may include being at specific LOD(s) (e.g., being at LOD 0 corresponding to the base mesh).
[0239] In some embodiments, based on the directional alignment indicating that the first and second vertex normals are not aligned, the coder determines that the first and second vertices are not to be adjusted.
[0240] Embodiments of the present disclosure may be implemented in hardware using analog and / or digital circuits, in software, through the execution of instructions by one or more general purpose or special-purpose processors, or as a combination of hardware and software. Consequently, embodiments of the disclosure may be implemented in the environment of a computer system or other processing system. An example of such a computer system 1600 is shown in FIG. 16. Blocks depicted in the figures above, such as the blocks in FIG. 1 , may execute on one or more computer systems 1600. Furthermore, each of the steps of the flowcharts depicted in this disclosure may be implemented on one or more computer systems 1600. When more than one computer system 1600 is used to implement embodiments of the present disclosure, the computer systems 1600 may be interconnected by one or more networks to form a cluster of computer systems that may act as a single pool of seamless resources. The interconnected computer systems 1600 may form a "cloud” of computers.
[0241] Computer system 1600 includes one or more processors, such as processor 1604. Processor 1604 may be, for example, a special purpose processor, general purpose processor, microprocessor, or digital signal processor. Processor 1604 may be connected to a communication infrastructure 1602 (for example, a bus or network). Computer system 1600 may also include a main memory 1606, such as random access memory (RAM), and may also include a secondary memory 1608.
[0242] Secondary memory 1608 may include, for example, a hard disk drive 1610 and / or a removable storage drive 1612, representing a magnetic tape drive, an optical disk drive, or the like. Removable storage drive 1612 may read from and / or write to a removable storage unit 1616 in a well-known manner. Removable storage unit 1616 represents a magnetic tape, optical disk, or the like, which is read by and written to by removable storage drive 1612. As will be appreciated by persons skilled in the relevant art(s), removable storage unit 1616 includes a computer usable storage medium having stored therein computer software and / or data.
[0243] In alternative implementations, secondary memory 1608 may include other similar means for allowing computer programs or other instructions to be loaded into computer system 1600. Such means may include, for example, a removable storage unit 1618 and an interface 1614. Examples of such means may include a program cartridge and cartridge interface (such as that found in video game devices), a removable memory chip (such as an EPROM or PROM) and associated socket, a thumb drive and USB port, and other removable storage units 1618 and interfaces 1614 which allow software and data to be transferred from removable storage unit 1618 to computer system 1600.
[0244] Computer system 1600 may also include a communications interface 1620 Communications interface 1620 allows software and data to be transferred between computer system 1600 and external devices. Examples of communications interface 1620 may include a modem, a network interface (such as an Ethernet card), a communications port, etc. Software and data transferred via communications interface 1620 are in the form of signals which may be electronic, electromagnetic, optical, or other signals capable of being received by communications interface 1620. These signals are provided to communications interface 1620 via a communications path 1622. Communications path 1622 carries signals and may be implemented using wire or cable, fiber optics, a phone line, a cellular phone link, an RF link, and other communications channels.
[0245] Computer system 1600 may also include one or more sensor(s) 1624. Sensor(s) 1624 may measure or detect one or more physical quantities and convert the measured or detected physical quantities into an electrical signal in digital and / or analog form. For example, sensor(s) 1624 may include an eye tracking sensor to track the eye movement of a user. Based on the eye movement of a user, a display of a 3D mesh may be updated. In another example, sensor(s) 1624 may include a head tracking sensor to the track the head movement of a user. Based on the head movement of a user, a display of a 3D mesh may be updated. In yet another example, sensor(s) 1624 may include a camera sensor fortaking photographs and / or a 3D scanning device, like a laser scanning, structured light scanning, and / or modulated light scanning device 3D scanning devices may obtain geometry information by moving one or more laser heads, structured light, and / or modulated light cameras relative to the object or scene being scanned. The geometry information may be used to construct a 3D mesh.
[0246] As used herein, the terms "computer program medium” and “computer readable medium” are used to refer to tangible storage media, such as removable storage units 1616 and 1618 or a hard disk installed in hard disk drive 1610. These computer program products are means for providing software to computer system 1600. Computer programs (also called computer control logic) may be stored in main memory 1606 and / or secondary memory 1608. Computer programs may also be received via communications interface 1620. Such computer programs, when executed, enable the computer system 1600 to implement the present disclosure as discussed herein In particular, the computer programs, when executed, enable processor 1604 to implement the processes of the present disclosure, such as any of the methods described herein. Accordingly, such computer programs represent controllers of the computer system 1600.
[0247] In another embodiment, features of the disclosure may be implemented in hardware using, for example, hardware components such as application-specific integrated circuits (ASICs) and gate arrays. Implementation of a hardware state machine to perform the functions described herein will also be apparent to persons skilled in the relevant art(s).
Claims
CLAIMS1. A method comprising: obtaining, from a bitstream, a base mesh for a 3-dimensional (3D) mesh; subdividing an edge, formed by a first and a second vertex from a first mesh associated with the base mesh, to determine a vertex, wherein the subdividing the edge comprises: determining a directional alignment, between a first vertex normal of the first vertex and a second vertex normal of the second vertex, based on a dot product between the first and second vertex normals; determining, based on the directional alignment, whether to refine a midpoint of the edge with a refinement vector that is based on a linear combination of the first and second vertex normals; and determining the vertex based on the midpoint and the determining of whether to refine the midpoint; obtaining, from the bitstream, displacements of second vertices of a second mesh comprising vertices of the first mesh and the vertex mesh; and reconstructing the 3D mesh based on updating the second vertices with the displacements.
2. A method comprising: obtaining a base mesh for a 3-dimensional (3D) mesh; subdividing an edge, formed by a first and a second vertex from a first mesh associated with the base mesh, to determine a vertex, wherein the subdividing the edge comprises: determining a directional alignment, between a first vertex normal of the first vertex and a second vertex normal of the second vertex, based on a dot product between the first and second vertex normals; determining, based on the directional alignment, whether to refine a midpoint of the edge with a refinement vector that is based on a linear combination of the first and second vertex normals; and determining the vertex based on the midpoint and the determining of whether to refine the midpoint; obtaining displacements for second vertices of a second mesh, comprising vertices of the first mesh and the vertex mesh, based on differences between the second mesh and the 3D mesh; and encoding the 3D mesh based on signaling, in a bitstream, the obtained displacements of the second mesh.
3. A method comprising: obtaining a base mesh for a 3-dimensional (3D) mesh; subdividing an edge, formed by a first and a second vertex from a first mesh associated with the base mesh, to determine a vertex, wherein the subdividing the edge comprises:determining a directional alignment between a first vertex normal of the first vertex and a second vertex normal of the second vertex; and determining the vertex based on a point along the edge and a refinement vector that is based on the directional alignment, the first vertex normal, and the second vertex normal; and encoding the 3D mesh based on a second mesh comprising vertices of the first mesh and the vertex.
4. The method of claim 3, wherein the encoding the 3D mesh comprises: obtaining displacements, for second vertices of the second mesh, based on differences between the second mesh and the 3D mesh; and signaling, in a bitstream, the obtained displacements of the second mesh.
5. The method of claim 2 or 4, wherein the differences are between the second mesh and a representation of the 3D mesh.
6. The method of claim 5, wherein the representation is a subdivided mesh fitted to the 3D mesh.
7. The method of any one of claims 2-6, further comprising: signaling, in the bitstream, a mode indication of a subdivision scheme, and wherein the subdividing the edge is based on the subdivision scheme being associated with normal-based subdivision.
8. A method comprising: obtaining, from a bitstream, a base mesh for a 3-dimensional (3D) mesh; subdividing an edge, formed by a first and a second vertex from a first mesh associated with the base mesh, to determine a vertex, wherein the subdividing the edge comprises: determining a directional alignment between a first vertex normal of the first vertex and a second vertex normal of the second vertex; and determining the vertex based on a point along the edge and a refinement vector that is based on the directional alignment, the first vertex normal, and the second vertex normal; and decoding the 3D mesh based on a second mesh comprising vertices of the first mesh and the vertex.
9. The method of any one of claims 3-8, wherein the point is a midpoint point of the edge.
10. The method of any one of claims 3-9, wherein the determining the directional alignment comprises determining a dot product between the first and second vertex normals.11 . The method of any one of claims 3-10, wherein the refinement vector is based on a linear combination of the first and second vertex normals.
12. The method of any one of claims 3-11 , wherein the determining the vertex comprises: determining, based on the directional alignment, whether to refine the point with the refinement vector, wherein the vertex is determined based on the point and the determining of whether to refine the point.
13. The method of any one of claims 1-2 or 12, wherein the refinement vector is enabled for refining the point based on the determination of directional alignment indicating the first and second vertex normals are aligned.
14. The method of any one of claims 1-2 or 12-13, wherein the vertex is determined as the point refined with the refinement vector based on the determined directional alignment indicating the first and second vertex normals are aligned.
15. The method of any one of claims 1-2 or 12-14, wherein the vertex is determined as the point without refinement by the refinement vector based on the determined directional alignment indicating the first and second vertex normals are not aligned.
16. The method of any one of claims 13-15, wherein the dot product being positive indicates the first and second vertex normals are aligned and the dot product being negative or zero indicates the first and second vertex normals are not aligned.
17. The method of any one of claims 1-16, further comprising: determining, based on the directional alignment, whether to update positions of the first and second vertex along and in the opposite directions of the first and second vertex normals, respectively.
18. The method of claim 17, wherein the positions of the first and second vertices are updated based on the determined directional alignment indicating the first and second vertex normals are aligned.
19. The method of any one of claims 17-18, wherein a position of the first vertex is updated based on subtracting the first vertex normal, after scaling, from the first vertex, and wherein a position of the second vertex is updated based on subtracting the second vertex normal, after scaling, from the second vertex.
20. The method of any one of claims 1-19, wherein the point being refined with the refinement vector comprises adding the refinement vector to the point to determine the vertex.21 . The method of any one of claims 1 -2 or 11-20, wherein the first and second vertex normals in the linear combination are weighted by a first weight and a second weight, respectively, wherein the first weight is determined based on a dot product between the first vertex normal and a first vector indicated by a displacement from the second vertex to the first vertex, and wherein the second weight is determined based on a dot product between the second vertex normal and a second vector indicated by a displacement from the first vertex to the second vertex.
22. The method of any one of claims 1-2 or 11-21 , wherein the refinement vector is weighted by a vector weight associated with a level of detail (LOD) of the vertex.
23. The method of any one of claims 21-22, wherein, based on the determined directional alignment, one or both of the first and second weights are adjusted.
24. The method of claim 23, wherein the adjusting the one or both of the first and second weights comprises inverting a sign of the one or both of the first and second weights, setting the one or both ofthe first and second weights to zero, setting the one or both of the first and second weights to a fixed value, or scaling the one or both of the first and second weights.
25. The method of any one of claims 1 -24, wherein the first mesh is a resulting mesh of one or more iterations of a subdivision scheme applied to the base mesh.
26. The method of any one of claims 1-25, wherein the first mesh is the base mesh.
27. The method of any one of claims 1 or 8-26, further comprising: obtaining, from the bitstream, a mode indication of a subdivision scheme, and wherein the subdividing the edge is based on the subdivision scheme being associated with normal-based subdivision.
28. The method of any one of claims 8-27, wherein the decoding the 3D mesh comprises: obtaining, from the bitstream, displacements of second vertices of the second mesh; and reconstructing the 3D mesh based on updating the second vertices with the displacements.
29. The method of any one of claims 1 or 28, wherein the obtained displacements are displacement vectors, and wherein updating the second vertices comprises each vertex of the second vertices being added to a respective displacement vector of the displacement vectors to determine a reconstructed vertex of vertices of the reconstructed 3D mesh.
30. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of an apparatus, cause the apparatus to perform the method of any one of claims 1- 29.31 . An encoder comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the encoder to perform the method of any one of claims 2-7 or 9-26.
32. A non-transitory computer-readable recording medium storing a bitstream generated by the method for encoding a video according to any one of claims 2-7 or 9-26.
33. A decoder comprising: one or more processors; and memory storing instructions that, when executed by the one or more processors, cause the decoder to perform the method of any one of claims 1 or 8-29.
34. A non-transitory computer readable medium storing a bitstream, which, when decoded by a decoder, causes the decoder to perform the method according to any one of claims 1 or 8-29.
35. A bitstream generated according to any one of claims 2-7 or 9-26.