Improved dual-degree-based coding algorithm for polygon mesh compression

Through the double-degree encoding algorithm, the problem of inefficient encoding efficiency of 3D polygon mesh is solved, efficient vertex and facet encoding is achieved, and encoding efficiency and decoding performance are improved.

CN120077411APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004474.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2024-06-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When encoding a 3D polygon mesh, it is difficult to efficiently process the double-degree encoding of surfaces and vertices, resulting in low encoding efficiency.

Method used

A double-degree-based encoding algorithm is proposed, by determining the first coded region and the second coded region of the polygon mesh, selecting a reference vertex, adding a polygonal face to the surface coding degree, and explicitly signaling the segmented vertices, linking the vertices to the polygonal face, to achieve efficient entropy coding and decoding.

Benefits of technology

The encoding efficiency of vertex and facets of 3D polygon mesh is improved, the signaling requirements for segmentation offsets and positions are reduced, and the encoding efficiency and decoding performance are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077411A_ABST
    Figure CN120077411A_ABST
Patent Text Reader

Abstract

The present disclosure relates generally to encoding and decoding of three-dimensional (3D) meshes, in particular to efficient encoding of vertices and planarity of 3D polygon meshes. In some implementations, when a face is encoded, vertices of the face may be linked after the divided vertices of the face, rather than signaled. In some other example implementations, a context may be selected for entropy encoding / decoding vertices in verticality or planes in planarity based on intersection features in planarity or verticality.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference

[0001] This application claims the benefit of priority of U.S. Non - provisional patent application Ser. No. 18 / 756,835, filed Jun. 27, 2024, and U.S. Provisional patent application Ser. No. 63 / 524,552, filed Jun. 30, 2023, both entitled "Improved Dual - degree based Coding Algorithm for Polygon Mesh Compression", which are hereby incorporated by reference in their entirety. Technical Field

[0002] The present disclosure generally relates to the encoding and decoding of three - dimensional (3D) meshes, and more particularly to efficient dual - degree coding of vertex degrees and face degrees of 3D polygon meshes. Background Art

[0003] The background description provided herein is intended to present the background of the present application as a whole. The extent to which the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of the filing of the present application, and it has never been expressly or implicitly admitted as prior art to the present application.

[0004] A variety of techniques have been developed to capture, represent, and simulate real - world objects, environments, etc. in 3D space. 3D representations of the world enable more immersive forms of interactive communication. Example 3D representations of objects and environments include, but are not limited to, point clouds and meshes. A sequence of 3D representations of objects and environments can form a video sequence. Redundancy and correlation within a sequence of 3D representations of objects and environments can be used to compress and encode such a video sequence into a more compact digital form. Summary of the Invention

[0005] The present disclosure generally relates to the encoding and decoding of three - dimensional (3D) meshes, and more particularly to efficient dual - degree coding of vertex degrees and face degrees of 3D polygon meshes. In some implementations, when encoding faces in a 3D mesh, the split vertices of the face are linked instead of being signaled. In some other example implementations, contexts for entropy encoding / decoding vertex degrees or face degrees can be selected based on cross - features in face degrees or vertex degrees.

[0006] In some example implementations, a method for encoding a polygon mesh is disclosed. The method may include determining a first encoded region and a second encoded region of the polygon mesh, the polygon mesh including a set of vertices and a set of faces that are separately encoded by vertex encoding degrees and face encoding degrees of the polygon mesh; selecting a reference vertex from the set of vertices in the vertex encoding degrees that belong to the first encoded region to add a polygon face to the set of faces in the face encoding degrees; for the polygon face, determining one or more vertices of the first encoded region in the set of vertices; linking the one or more vertices to the polygon face in the face encoding degrees of the polygon mesh; identifying a splitting vertex of a polygon face in the second encoded region in the set of vertices; explicitly signaling the splitting vertex in the vertex encoding degrees and linking the splitting vertex to the polygon face in the face encoding degrees; identifying vertices adjacent to the splitting vertex in the second encoded region to link the vertices to the polygon face in the face encoding degrees.

[0007] In the above example implementation, the first encoded region and the second encoded region are not contiguous.

[0008] In any of the above example implementations, the reference vertex for the polygon face to be added in the first encoded region is selected from the set of vertices by selecting a candidate vertex with the maximum concavity from the first encoded region.

[0009] In any of the above example implementations, the concavity of each vertex is quantified by the sum of all angles of the faces adjacent to the each vertex in the first encoded region.

[0010] In any of the above example implementations, the reference vertex for the polygon face to be added in the first encoded region is selected by selecting a candidate node with the largest set of connected unencoded faces from the first encoded region.

[0011] In any of the above example implementations, the set of vertices includes a first symbol sequence, the set of faces includes a second symbol sequence, and the encoded splitting offsets and positions are separated from other offsets and positions within the first symbol sequence.

[0012] In any of the above example implementations, the set of vertices includes a first symbol sequence, the set of faces includes a second symbol sequence, and before entropy encoding, the values of at least one symbol subset in at least one of the first symbol sequence and the second symbol sequence are shifted downward by a minimum value associated with the at least one symbol subset.

[0013] In any of the above example implementations, the at least one symbol subset includes the number of edges of at least two faces, and the offset is 2.

[0014] In any of the above example implementations, the initial face among the set of faces of the initial mesh component of the polygon mesh to be encoded and the initial vertex among the set of vertices are determined as the face and vertex of the initial mesh component closest to the centroid of the initial mesh component or closest to the origin of the polygon mesh, and the next initial face among the set of faces of the next mesh component of the polygon mesh to be encoded and the next initial vertex among the set of vertices are determined as the face and vertex of the next mesh component closest to the centroid of the next mesh component or closest to the last encoded vertex of the initial mesh component, whichever provides a lower encoding cost.

[0015] In any of the above example implementations, the initial face among the set of faces for encoding the mesh component of the polygon mesh and / or the initial vertex among the set of vertices are determined as the face and vertex that provide the lowest encoding cost together with the N next faces and / or vertices of the mesh component among the other faces and / or vertices of the mesh component.

[0016] In some other example implementations, a method for decoding a bitstream of a polygon mesh is disclosed. The method may include receiving the bitstream; extracting from the bitstream a first sub-bitstream including the encoded vertices of the polygon mesh and a second sub-bitstream including the encoded faces of the polygon mesh; selecting a context to entropy-decode the encoded vertices in the first sub-bitstream with at least one subset of the encoded faces in the second sub-bitstream, or entropy-decode the encoded faces in the second sub-bitstream based on at least one subset of the encoded vertices in the first sub-bitstream; and entropy-decoding the encoded vertices or the encoded faces using the selected context.

[0017] In any of the above example implementations, the method includes selecting the context to entropy-decode the encoded vertices in the first sub-bitstream by: identifying the number of faces including the encoded vertices; and selecting the context for decoding the encoded vertices from a set of contexts according to the number of faces.

[0018] In any of the above example implementations, the set of contexts is mapped to a set of numbers of faces.

[0019] In any of the above example implementations, multiple highest numbers of faces are mapped to one context in the set of contexts.

[0020] In any of the above example implementations, the method includes selecting the context for entropy decoding the encoded faces in the second sub-bitstream by: identifying the number of vertices of the encoded faces; and selecting the context for decoding the encoded faces from a set of contexts based on the number of vertices.

[0021] In any of the above example implementations, the set of contexts is mapped to a set of numbers of vertices.

[0022] In any of the above example implementations, multiple highest numbers of vertices are mapped to one context in the set of contexts.

[0023] In some other example implementations, a method for encoding a polygon mesh is disclosed, the method including: generating a first sub-bitstream of a bitstream including vertices of the polygon mesh; generating a second sub-bitstream separated from the first sub-bitstream and including faces of the polygon mesh; selecting a context for entropy encoding the vertices in the first sub-bitstream based on at least one subset of the faces in the second sub-bitstream, or for entropy encoding the faces in the second sub-bitstream based on at least one subset of the vertices in the first sub-bitstream; and entropy encoding the vertices or the faces using the selected context.

[0024] Aspects of the present disclosure also provide an electronic device or apparatus including circuitry configured to perform any of the above method implementations.

[0025] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing computer instructions that, when executed by a computer for 3D mesh processing, cause the computer to perform any of the above method implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0027] Figure 1 is a simplified block diagram schematic of an example communication system according to an embodiment of the present disclosure;

[0028] Figure 2 is a simplified block diagram schematic of an example streaming system according to an embodiment of the present disclosure;

[0029] Figure 3 illustrates the data flow in the encoding and decoding of a 3D mesh frame or a point cloud frame according to some embodiments of the present disclosure;

[0030] Figure 4Shows a block diagram of an encoder for encoding a 3D mesh frame or a point cloud frame according to some embodiments of the present disclosure;

[0031] Figure 5 Shows a block diagram of a decoder for decoding a compressed bitstream corresponding to a 3D mesh frame and a point cloud frame according to some embodiments of the present disclosure;

[0032] Figure 6 Is a simplified block diagram schematic of a video decoder according to an embodiment of the present disclosure;

[0033] Figure 7 Is a simplified block diagram schematic of a video encoder according to an embodiment of the present disclosure;

[0034] Figure 8 Illustrates an example face degree and vertex degree of a polygon mesh;

[0035] Figure 9 Illustrates an example implementation for maximizing the connection of vertices;

[0036] Figure 10 Illustrates an example implementation for determining a reference vertex;

[0037] Figure 11 Illustrates another example implementation for determining a reference vertex;

[0038] Figure 12 Shows an example flowchart outlining an example process according to some embodiments of the present disclosure;

[0039] Figure 13 Shows another example flowchart outlining an example process according to some embodiments of the present disclosure;

[0040] Figure 14 Shows yet another example flowchart outlining an example process according to some embodiments of the present disclosure; and

[0041] Figure 15 Is a schematic diagram of an example computer system according to an embodiment. Detailed Description

[0042] Throughout the specification and claims, terms may have nuanced meanings that are implied or implicit in a context beyond their explicitly stated meaning. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" does not necessarily refer to different embodiments. Similarly, as used herein, the phrase "in one implementation" or "in some implementations" does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" does not necessarily refer to different implementations. For example, the claimed subject matter is intended to include combinations of all or part of the exemplary embodiments / implementations.

[0043] Generally, terms can be understood, at least in part, based on their use in context. For example, as used herein, terms such as "and," "or," or "and / or" can include a variety of meanings that may depend, at least in part, on the context in which such terms are used. Generally, "or" if used in connection with a list (such as A, B, or C) is intended to mean A, B, and C (used herein in an inclusive sense), as well as A, B, or C (used herein in an exclusive sense). Additionally, at least in part depending on the context, the term "one or more" or "at least one" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, at least in part depending on the context, terms such as "a," "an," or "the" can likewise be understood to convey a singular usage or to convey a plural usage. Further, the terms "based on" or "determined by" can be understood to not necessarily be intended to convey a set of exclusive factors, but can allow for the existence of additional factors that are not necessarily explicitly described, again, at least in part depending on the context.

[0044] Technological developments in 3D media processing (such as advancements in 3D capture, 3D modeling, and 3D rendering) have driven the widespread creation of 3D content across several platforms and devices. Such 3D content contains information that can be processed to generate various forms of media to provide, for example, immersive viewing / rendering and interactive experiences. The applications of 3D content are diverse and include, but are not limited to, virtual reality, augmented reality, metaverse interactions, gaming, immersive video conferencing, robotics, computer-aided design (CAD), and the like. According to one aspect of the present disclosure, to improve the immersive experience, 3D models are becoming increasingly complex, and the creation and consumption of 3D models require substantial data resources, such as data storage resources, data transmission resources, and data processing resources.

[0045] Compared to traditional two-dimensional (2D) content, which is typically represented by a dataset in the form of a 2D pixel array (such as an image), 3D content with three-dimensional full-resolution pixelation can be prohibitively resource-intensive and unnecessary in many (if not most) practical applications. In most 3D immersive applications, according to some aspects of the present disclosure, a less data-intensive representation of 3D content can be employed. For example, in most applications, only the terrain information of objects in a 3D scene (a real-world scene captured by a sensor such as a LIDAR device or an animated 3D scene generated by a software tool) rather than the volume information may be required. Thus, a more efficient form of dataset can be used to represent 3D objects and 3D scenes. For example, a 3D mesh can be used as a type of 3D model to represent immersive 3D content, such as 3D objects in a 3D scene.

[0046] The mesh of one or more objects (or referred to as a mesh model) can include a set of vertices. These vertices can be connected to each other to form edges. These edges can be further connected to form faces. These faces can further form polygons. These polygons can describe the surface of a volumetric object. Each polygon can be defined by its vertices in 3D space and information on how these vertices are connected. Thus, the 3D surfaces of individual objects can be decomposed into, for example, faces and polygons. Each of the vertices, edges, faces, polygons, or surfaces can be associated with various attributes (such as color, surface normal, texture, etc.). The normal of a surface can be referred to as a surface normal; and / or the normal of a vertex can be referred to as a vertex normal. Attributes can also be associated with the surface of the mesh through mapping information that parameterizes the mesh using a 2D attribute map. This mapping is typically described by a set of parametric coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. The 2D attribute map is used to store high-resolution attribute information, such as texture, normal, displacement, etc. Such information can be used for various purposes, such as texture mapping, shading, and mesh reconstruction. The information on how these vertices are connected into edges, faces, or polygons can be referred to as connectivity information. Connectivity information is important for uniquely defining the components of the mesh because the same set of vertices can form different faces, surfaces, and polygons. Generally, the position of a vertex in 3D space can be represented by its 3D coordinates. A face can be represented by a set of sequentially connected vertices, each vertex associated with a set of 3D coordinates. Similarly, an edge can be represented by two vertices each associated with their 3D coordinates. Vertices, edges, and faces can be indexed in a 3D mesh dataset.

[0047] A mesh can be defined and described by a collection of one or more of these basic element types. However, not all types of elements described above are necessary to fully describe a mesh. For example, a mesh can be fully described by using only vertices and their connectivity. For another example, a mesh can be fully described by using only a list of faces and the common vertices of these faces. Thus, a mesh can have various alternative types consisting of and described by alternative data sets. Example mesh types include, but are not limited to, face-vertex meshes, winged-edge meshes, half-edge meshes, quad-edge meshes, corner table meshes, vertex-vertex meshes, etc. Correspondingly, mesh data sets can be stored together with information conforming to alternative file formats having file extensions including, but not limited to, .raw, .blend, .fbx, .3ds, .dae, .dng, 3dm, .dsf, .dwg, .obj, .ply, .pmd, .stl, amf, .wrl, .wrz, .x3d, .x3db, .x3dv, .x3dz, .x3dbz, .x3dvz, .c4d, .lwo, .smb, .msh, .mesh, .veg, .z3d, .vtk, .l4d, etc. The attributes of these elements, such as color, surface normal, texture, etc., can be included in the mesh data set in various ways.

[0048] In some embodiments, the vertices of a mesh can be mapped into a pixelated 2D space (referred to as the UV space). Thus, each vertex of the mesh can be mapped to a pixel in the UV space. In some embodiments, a vertex can be mapped to more than one pixel in the UV space. For example, a vertex at a boundary can be mapped to two or three pixels in the UV space. Similarly, the faces or surfaces in a mesh can be sampled as a plurality of 3D points, which may or may not be among the recorded vertices in the mesh, and these plurality of 3D points can also be mapped to pixels in the two-dimensional UV space. Mapping the vertices and the sampled 3D points of the faces or surfaces in a mesh into the UV space, and subsequent data analysis and processing in the UV space can facilitate data storage, compression, and encoding of the 3D data set of the mesh or mesh sequence, as described in further detail below. The mapped UV space data set can be referred to as a UV image, or a 2D map, or a 2D image of the mesh.

[0049] After mapping the vertices and sampled surface points in the 3D mesh into the 2D UV space, some pixels may be mapped to the vertices and sampled surface points of the 3D mesh, while other pixels may not have been mapped (or are unmapped). Each mapped pixel in the 2D image of the mesh can be associated with the information of the corresponding mapped vertex or surface point in the 3D mesh. Depending on the type of information included in the pixels in the UV space, various 2D images or 2D maps of the mesh can be constructed. A collection of multiple 2D maps can be used as an alternative representation and / or combined representation of the mesh.

[0050] For example, the simplest 2D map for the mesh can be constructed as an occupancy map. This occupancy map can indicate the pixels in the UV space that are mapped to the 3D vertices or sampled surface points of the mesh. The indication of occupancy can be represented by a binary indicator at each 2D pixel, where, for example, the binary value "1" indicates mapping or occupancy, and the binary value "0" indicates non - mapping or non - occupancy (unmapped). Thus, the occupancy map can be constructed as a 2D image. While a normal 2D image contains an array of three channels (RGB, YUV, YCrCb, etc.) with a color depth of, for example, 8 bits, this 2D occupancy map of the mesh only requires a single - bit binary channel. The pixels with the value "1" in the 2D image are mapped and occupied; otherwise, the pixel is unmapped and unoccupied.

[0051] For another example, a 2D geometry map can be constructed for the mesh. The 2D geometry map will be a full three - channel image instead of containing a single binary channel, where each of the three - color channels at each occupied pixel will correspond to the three 3D coordinates of the corresponding mapped vertex or sampled 3D point in the mesh.

[0052] In some embodiments, other 2D maps can be constructed for the mesh. For example, a set of attributes for each of the vertices and sampled 3D points of the mesh can be extracted from the mesh dataset, and this set of attributes can be encoded into the three color channels of the 2D map image. Such a 2D map can be called an attribute map of the mesh. A particular attribute map can contain the three - channel color of each occupied pixel in the UV space. For another example, the texture attributes associated with each mapped vertex or sampled 3D point of the mesh can be parameterized as three - channel values and encoded as a 2D attribute map. For another example, the normal attributes associated with each mapped vertex or sampled 3D point of the mesh can be parameterized as three - channel values and encoded as a 2D attribute map. In some example embodiments, multiple 2D attribute maps can be constructed in order to preserve all the necessary attribute information of the vertices and sampled surface points of the mesh.

[0053] The above 2D maps are merely examples. Other types of 2D maps of the mesh can be constructed. Additionally, other datasets can be extracted from the 3D mesh to be associated with the above 2D Figure 1The combined representation represents the original 3D mesh. For example, the connectivity or connectivity information between vertices can be grouped and organized separately in a list, table, etc., independently of the 2D graph. The connectivity information can refer to vertices using vertex indices, for example. The vertex indices can be mapped to their corresponding pixel positions in the 2D graph. For another example, surface textures, colors, normals, displacements, and other information can be extracted and organized separately, independently of the 2D graph, rather than being represented as a 2D graph. Other metadata can be further extracted from the 3D mesh to represent the 3D mesh in combination with the above 2D graph and other data sets.

[0054] Although the above example embodiments focus on static meshes, according to one aspect of the present disclosure, the 3D mesh can be dynamic. For example, a dynamic mesh can refer to a mesh in which at least one of the components (geometric information, connectivity information, mapping information, vertex attributes, and attribute graphs) changes over time. Thus, a dynamic mesh can be described by a sequence of meshes or meshes (also referred to as mesh frames), similar to the timed sequence of 2D image frames that form a video.

[0055] In some example embodiments, the dynamic mesh can have constant connectivity information, time-varying geometry, and time-varying vertex attributes. In some other examples, the dynamic mesh can have time-varying connectivity information. In some examples, digital 3D content creation tools can be used to generate dynamic meshes with time-varying attribute graphs and time-varying connectivity information. In some other examples, volumetric acquisition / detection / sensing techniques are used to generate dynamic meshes. Volumetric acquisition techniques can generate dynamic meshes with time-varying connectivity information, especially under real-time constraints.

[0056] Dynamic meshes can require a large amount of data because the dynamic mesh may include a large amount of information that changes over time. However, compression can be performed to take advantage of the redundancy within the mesh frames (intra-frame compression) and between the mesh frames (inter-frame compression). Various mesh compression processes can be implemented to allow for efficient storage and transmission of media content in the mesh representation, especially for mesh sequences.

[0057] Aspects of the present disclosure provide example architectures and techniques for mesh compression. These techniques can be used for various mesh compressions, including but not limited to static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying attribute graphs, etc. These techniques can be used for lossy and lossless compression in various applications such as real-time immersive communication, storage, free viewpoint video, augmented reality (AR), and virtual reality (VR). These applications can include features such as random access and scalable / progressive coding.

[0058] While the present disclosure clearly describes techniques and implementations applicable to 3D meshes, the basic principles of the various implementations described herein are applicable to other types of 3D data structures, including but not limited to point cloud (PC) data structures. For simplicity, the following references to 3D meshes are intended to be general and include other types of 3D representations, such as point clouds and other 3D volume datasets.

[0059] Turning first to an example architecture-level implementation, Figure 1 FIG. illustrates a simplified block diagram of a communication system (100) according to an example embodiment of the present disclosure. The communication system (100) may include a plurality of terminal devices that may communicate with each other via, for example, a communication network (150) (alternatively referred to as a network). For example, the communication system (100) may include a pair of terminal devices (110) and (120) interconnected via a network (150). In Figure 1 an example, the first pair of terminal devices (110) and (120) may perform a one-way transmission of a 3D mesh. For example, the terminal device (110) may compress a 3D mesh or a sequence of 3D meshes, which may be generated by the terminal device (110), obtained from a storage device, or captured by a 3D sensor (105) connected to the terminal device (110). The compressed 3D mesh or sequence of 3D meshes may be transmitted via the network (150) to another terminal device (120), for example, in the form of a bitstream (also referred to as an encoded bitstream). The terminal device (120) may receive the compressed 3D mesh or sequence of 3D meshes from the network (150), decompress the bitstream to reconstruct the original 3D mesh or sequence of 3D meshes, and appropriately process the reconstructed 3D mesh or sequence of 3D meshes for display or other purposes / uses. One-way data transmission may be common in media service applications and the like.

[0060] In Figure 1 an example, one or both of the terminal devices (110) and (120) may be implemented as a server, a fixed or mobile personal computer, a laptop computer, a tablet computer, a smartphone, a gaming terminal, a media player, and / or a dedicated three-dimensional (3D) device, etc., but the principles of the present disclosure are not limited thereto. The network (150) may represent any type of network or combination of networks that transmits the compressed 3D mesh between the terminal devices (110) and (120). The network (150) may include, for example, a wired (wired) communication network and / or a wireless communication network. The network (150) may exchange data in a circuit-switched channel and / or a packet-switched channel. Representative networks include long-distance telecommunication networks, local area networks, wide area networks, cellular networks, and / or the Internet. For the purposes of the present disclosure, unless otherwise explained hereinafter, the architecture and topology of the network (150) may be immaterial to the operation of the present disclosure.

[0061] Figure 2 FIG. illustrates a simplified block diagram example of a streaming system (200) in accordance with an embodiment of the present disclosure. Figure 2 FIG. illustrates an example application of the disclosed embodiments related to 3D meshes and compressed 3D meshes. The disclosed subject matter may equally apply to other 3D mesh or point cloud enabled applications such as 3D telepresence applications, virtual reality applications, etc.

[0062] The streaming system (200) may include a capture or storage subsystem (213). The capture or storage subsystem (213) may include a 3D mesh generator or storage medium (201), such as a 3D mesh or point cloud generation tool / software, a graphics generation component, or a point cloud sensor, such as a light detection and ranging (LIDAR) system, a 3D camera, a 3D scanner, a 3D mesh memory, etc., that generates or provides an uncompressed 3D mesh (202) or point cloud (202). In some example embodiments, the 3D mesh (202) includes vertices of the 3D mesh or 3D points of the point cloud (both referred to as 3D meshes). The 3D mesh (202) is depicted as a thick line to emphasize the corresponding high data volume when compared to the compressed 3D mesh (204) (the bitstream of the compressed 3D mesh). The compressed 3D mesh (204) may be generated by an electronic device (220) including an encoder (203) coupled to the 3D mesh (202). The encoder (203) may include hardware, software, or a combination thereof to implement aspects of the disclosed subject matter, as described in more detail below. The compressed 3D mesh (204) (or the bitstream of the compressed 3D mesh (204)) (depicted as a thin line to emphasize the lower data volume when compared to the stream of the uncompressed 3D mesh (202)) may be stored in the streaming server (205) for future use. One or more streaming client subsystems, such as Figure 2 client subsystems (206) and (208) therein, may access the streaming server (205) to retrieve copies (207) and (209) of the compressed 3D mesh (204). The client subsystem (206) may include, for example, a decoder (210) in an electronic device (230). The decoder (210) may be configured to decode an incoming copy (207) of the compressed 3D mesh and create an outgoing stream of a reconstructed 3D mesh (211) that can be rendered on a rendering device (212) or used for other purposes.

[0063] It should be noted that the electronic device (220) and the electronic device (230) may include other components (not shown). For example, the electronic device (220) may include a video decoder (not shown), and the electronic device (230) may further include a video encoder (not shown).

[0064] In some streaming systems, compressed 3D meshes (204), (207), and (209) (e.g., the bitstreams of the compressed 3D meshes) can be compressed according to certain criteria. In some examples, as described in further detail below, video coding standards are used to exploit the redundancy and correlation in 3D mesh compression after first projecting the 3D mesh to map it to a 2D representation suitable for video compression. Non-limiting examples of such standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), etc., as described in further detail below.

[0065] A compressed 3D mesh or sequence of 3D meshes can be generated by an encoder, and a decoder can be configured to decompress the compressed or encoded 3D mesh. Figure 3 Illustrated is a high-level example data flow of 3D meshes in such an encoder (301) and decoder (303). As Figure 3As shown, the original input 3D mesh or 3D mesh sequence (302) can be pre - processed by trajectory re - meshing, parameterization, and / or voxelization to generate input data for mapping units that map the 3D mesh to a 2D UV space (304). In some embodiments, the 2D UV space can include a mesh with a UV atlas. The 3D mesh can be sampled to include 3D surface points that may not be among the vertices, and these sampled 3D surface points are added to the mapping in the UV space. Various 2D maps can be generated in the encoder 301, including but not limited to Occupancy Maps (310), Geometry Maps (312), and Attribute Maps (314). Maps of these image types can be compressed by the encoder 301 using, for example, video encoding / compression techniques. For example, a video encoder can use intra - prediction techniques and inter - prediction through other 3D mesh reference frames to assist in compressing 3D mesh frames. For non - limiting examples, other non - image or non - map data or metadata (316) can also be encoded in various ways to remove redundancy, thereby generating compressed non - map data via entropy encoding. Then, the encoder 301 can combine or multiplex the compressed 2D maps and non - map data, and further encode the combined data to generate an encoded bitstream (or referred to as an encoded bitstream). Then this encoded bitstream can be stored or transmitted for use by the decoder 303. The decoder can be configured to decode the bitstream, demultiplex the decoded bitstream to obtain the compressed 2D maps and non - map data, and perform decompression to generate a decoded occupancy map (320), a decoded geometry map (322), a decoded attribute map (324), and decoded non - map data and metadata (326). Then, the decoder 303 can be further configured to reconstruct the 3D mesh or 3D mesh sequence (330) based on the decoded 2D maps (320, 322, and 324) and the decoded non - map data (326).

[0066] More specifically, Figure 4 A block diagram of an example 3D mesh encoder (400) for encoding 3D mesh frames according to some embodiments of the present disclosure is shown. In some example embodiments, the mesh encoder (400) can be used in a communication system (100) and a streaming system (200). For example, the encoder (203) can be configured and operate in a similar manner to the mesh encoder (400).

[0067] The mesh encoder (400) can receive a 3D mesh frame as an uncompressed input and generate a bitstream corresponding to the compressed 3D mesh frame. In some example embodiments, the mesh encoder (400) can receive a 3D mesh frame from any source, such as Figure 2 a mesh or point cloud source (201), etc.

[0068] In Figure 4 an example of, the mesh encoder (400) may include a patch generation module (406) (or referred to as a chart generation module), a patch packing module (408), a geometry image generation module (410), a texture image generation module (412), a patch information module (404), an occupancy map module (414), a smoothing module (436), an image filling module (416) and (418), a grouped dilation module (420), a video compression module (422), (423) and (432), an auxiliary patch information compression module (438), an entropy compression module (434), and a multiplexer (424).

[0069] In various embodiments of the present disclosure, a module may refer to a software module, a hardware module, or a combination thereof. A software module may include a computer program or a part of the computer program, which has a predefined function and works together with other related parts to achieve a predefined goal, such as those functions described in the present disclosure. A hardware module may be implemented using a processing circuit system and / or a memory configured to execute the functions described in the present disclosure. Each module may be implemented using one or more processors (or multiple processors and a memory). Similarly, a processor (or multiple processors and a memory) may be used to implement one or more modules. In addition, each module may be a part of an overall module including the functions of the module. The description herein may also apply to the terms module and other equivalent terms (e.g., unit).

[0070] According to one aspect of the present disclosure, and as described above, the mesh encoder (400) converts a 3D mesh frame into an image-based representation (e.g., a 2D chart), and some non-primitive metadata (e.g., patch or chart information) for helping to convert the compressed 3D mesh back to the decompressed 3D mesh. In some examples, the mesh encoder (400) may convert a 3D mesh frame into a 2D geometry chart or image, a texture chart or image, and an occupancy chart or image, and then use video coding techniques to encode the geometry image, the texture image, and the occupancy chart together with the metadata and other compressed non-graphical data into a bitstream. Generally, and as described above, a 2D geometry image is a 2D image in which 2D pixels are filled with geometric values associated with 3D points projected (the term "projected" is used to mean "mapped") onto these 2D pixels, and the 2D pixels filled with geometric values may be referred to as geometric samples. A texture image is a 2D image in which pixels are filled with texture values associated with 3D points projected onto 2D pixels, and the 2D pixels filled with texture values may be referred to as texture samples. An occupancy chart is a 2D image in which 2D pixels are filled with values indicating whether a 3D point is occupied or unoccupied.

[0071] The patch generation module (406) divides the 3D mesh into a set of charts or patches (e.g., patches are defined as continuous subsets of the surface described by the 3D mesh or point cloud), which may or may not overlap, such that each patch can be described by a depth field relative to a plane in 2D space (e.g., flattening the surface such that deeper 3D points on the surface are further from the center of the corresponding 2D map). In some embodiments, the patch generation module (406) aims to decompose the 3D mesh into the minimum number of patches with smooth boundaries while also minimizing the reconstruction error.

[0072] The patch information module (404) can collect patch information indicating the size and shape of the patches. In some examples, the patch information can be packed into a data frame and then encoded by an auxiliary patch information compression module (438) to generate compressed auxiliary patch information. Auxiliary patch compression can be implemented in various forms, including but not limited to various types of arithmetic coding.

[0073] The patch or chart packing module (408) can be configured to map the extracted patches onto a 2D grid of the UV space while minimizing the unused space. In some example implementations, the pixels of the 2D UV space can be granulated into pixel blocks to map the patches or charts. The block size can be predefined. For example, the block size can be M×M (e.g., 16×16). With this granularity, it can be ensured that each M×M block of the 2D UV grid is associated with a unique patch. In other words, each patch is mapped to a 2D UV space with a granularity of M×M. Efficient patch packing can directly affect the compression efficiency by minimizing the unused space or ensuring temporal consistency. Example implementations of packing patches or charts into the 2D UV space are given in further detail below.

[0074] The geometric image generation module (410) can generate a 2D geometric image associated with the geometry of the 3D mesh at a given patch location in the 2D grid. The texture image generation module (412) can generate a 2D texture image associated with the texture of the 3D mesh at a given patch location in the 2D grid. The geometric image generation module (410) and the texture image generation module (412) basically utilize the 3D-to-2D mapping calculated during the above packing process to store the geometry and texture of the 3D mesh as 2D images, as described above. In some embodiments, to better handle the case where multiple points are projected onto the same sample (e.g., patches overlap in the 3D space of the mesh), the 2D image can be layered. In other words, each patch can be projected onto, for example, two images (referred to as layers) such that multiple points can be projected onto the same point in different layers.

[0075] In some example embodiments, a geometric image can be represented by a monochromatic frame of size width × height (W×H). Thus, three geometric images with three luminance channels or chrominance channels can be used to represent 3D coordinates. In some example embodiments, a geometric image can be represented by a 2D image having three channels (such as RGB, YUV, YCrCb, etc.), with a specific color depth (e.g., 8 bits, 12 bits, 16 bits, etc.). Thus, one geometric image with three color channels can be used to represent 3D coordinates.

[0076] To generate a texture image, a texture generation program uses the reconstructed / smoothed geometry to calculate the color associated with the sampled points from the original 3D mesh (see Figure 3 “Sampling” which, for example, will generate 3D surface points that are not among the vertices of the original 3D mesh).

[0077] The occupancy map module (414) can be configured to generate an occupancy map that describes the occupancy information at each cell. For example, as described above, the occupancy image can include a binary map that indicates for each cell of a 2D grid whether the cell belongs to empty space or to the 3D mesh. In some example embodiments, the occupancy map can use binary information to describe for each pixel whether the pixel is occupied. In some other example embodiments, the occupancy map can use binary information to describe for each pixel block (e.g., each M×M block) whether the pixel block is occupied.

[0078] The occupancy map generated by the occupancy map module (414) can be compressed using lossless coding or lossy coding. When using lossless coding, the entropy compression module (434) can be used to compress the occupancy map. When using lossy coding, the video compression module (432) can be used to compress the occupancy map.

[0079] Note that the patch packing module (408) may leave some empty space between the 2D patches packed in an image frame. The image padding modules (416) and (418) can fill the empty space (referred to as padding) in order to generate an image frame that may be suitable for 2D video and image codecs. Image padding is also known as background padding, which can fill the unused space with redundant information. In some examples, well-implemented background padding increases the bit rate minimally while avoiding introducing significant coding artifacts around the patch boundaries.

[0080] Video compression modules (422), (423), and (432) can encode 2D images (such as filled geometric images, filled texture images, and occupancy maps) based on suitable video coding standards (such as HEVC, VVC, etc.). In some example embodiments, the video compression modules (422), (423), and (432) are separate components that operate independently. Note that in some other example embodiments, the video compression modules (422), (423), and (432) can be implemented as a single component.

[0081] In some example embodiments, the smoothing module (436) can be configured to generate a smoothed image of the reconstructed geometric image. This smoothed image can be provided to the texture image generation (412). Then, the texture image generation (412) can adjust the generation of the texture image based on the reconstructed geometric image. For example, when the patch shape (e.g., geometric) undergoes slight deformation during encoding and decoding, this deformation can be considered during the generation of the texture image to correct the deformation of the patch shape.

[0082] In some embodiments, the grouped dilation (420) is configured to fill the pixels around the object boundary with redundant low-frequency content in order to improve the coding gain and the visual quality of the reconstructed 3D mesh.

[0083] The multiplexer (424) can be configured to multiplex the compressed geometric image, compressed texture image, compressed occupancy map, and compressed auxiliary patch information into the compressed bitstream.

[0084] Figure 5 A block diagram of an example mesh decoder (500) for decoding a compressed bitstream corresponding to a 3D mesh frame according to some embodiments of the present disclosure is shown. In some example embodiments, the mesh decoder (500) can be used in a communication system (100) and a streaming system (200). For example, the decoder (210) can be configured to operate in a similar manner to the mesh decoder (500). The mesh decoder (500) receives the compressed bitstream and generates a reconstructed 3D mesh based on the compressed bitstream including, for example, the compressed geometric image, compressed texture image, compressed occupancy map, and compressed auxiliary patch information.

[0085] In Figure 5 the example of, the mesh decoder (500) can include a demultiplexer (532), video decompression modules (534) and (536), an occupancy map decompression module (538), an auxiliary patch information decompression module (542), a geometric reconstruction module (544), a smoothing module (546), a texture reconstruction module (548), and a color smoothing module (552).

[0086] The demultiplexer (532) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometry image, a compressed occupancy map, and compressed auxiliary patch information.

[0087] Video decompression modules (534) and (536) can decode the compressed images according to suitable standards (e.g., HEVC, VVC, etc.) and output the decompressed images. For example, video decompression module (534) can decode the compressed texture image and output the decompressed texture image. Video decompression module (536) can further decode the compressed geometry image and output the decompressed geometry image.

[0088] The occupancy map decompression module (538) can be configured to decode the compressed occupancy map according to a suitable standard (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.

[0089] The auxiliary patch information decompression module (542) can be configured to decode the compressed auxiliary patch information according to a suitable decoding algorithm and output the decompressed auxiliary patch information.

[0090] The geometry reconstruction module (544) can be configured to receive the decompressed geometry image and generate a reconstructed 3D mesh geometry based on the decompressed occupancy map and the decompressed auxiliary patch information.

[0091] The smoothing module (546) can be configured to smooth the inconsistencies at the patch edges. The smoothing procedure can aim to mitigate potential discontinuities that may occur at the patch boundaries due to compression artifacts. In some example embodiments, a smoothing filter can be applied to the pixels located on the patch boundary to mitigate the distortion that may be caused by compression / decompression.

[0092] The texture reconstruction module (548) can be configured to determine the texture information of the points in the 3D mesh based on the decompressed texture image and the smoothed geometry.

[0093] The color smoothing module (552) can be configured to smooth the coloring inconsistencies. Non - adjacent patches in 3D space are typically packed next to each other in 2D video. In some examples, the pixel values from non - adjacent patches may be confounded by a block - based video codec. The goal of color smoothing may be to reduce the visible artifacts that appear at the patch boundaries.

[0094] Figure 6A block diagram of an example video decoder (610) in accordance with an embodiment of the present disclosure is shown. The video decoder (610) may be used in a mesh decoder (500). For example, the video decompression modules (534) and (536), and the occupancy map decompression module (538) may be similarly configured as the video decoder (610).

[0095] The video decoder (610) may include a parser (620) to reconstruct symbols (621) from a compressed image, such as an encoded video sequence. The categories of these symbols may include information for managing the operation of the video decoder (610). The parser (620) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be according to a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (620) may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the set. The subgroup may include a group of pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (620) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0096] The parser (620) may perform entropy decoding / parsing operations on the image sequence received from the buffer memory to create symbols (621).

[0097] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (621) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (620) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (620) and the multiple units below are not described.

[0098] In addition to the functional blocks already mentioned, the video decoder (610) may be conceptually divided into many functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. The conceptual division into functional units below is only for the purpose of describing the disclosed subject matter.

[0099] The video decoder (610) may include a scaler / inverse transform unit (651). The scaler / inverse transform unit (651) may receive quantized transform coefficients and control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) as one or more symbols (621) from a parser (620). The scaler / inverse transform unit (651) may output a block including sample values, and these sample values may be input into an aggregator (655).

[0100] In some cases, the output samples of the scaler / inverse transform unit (651) may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (652). In some cases, the intra-picture prediction unit (652) may generate a surrounding block having the same size and shape as the block being reconstructed using the reconstructed information extracted from a current picture buffer (658). For example, the current picture buffer (658) may buffer a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (655) may add the predictive information generated by the intra-prediction unit (652) to the output sample information provided by the scaler / inverse transform unit (651) based on each sample.

[0101] In other cases, the output samples of the scaler / inverse transform unit (651) may belong to an inter-coded and potentially motion-compensated block. In this case, a motion compensation prediction unit (653) may access a reference picture memory (657) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (621), these samples may be added by the aggregator (655) to the output of the scaler / inverse transform unit (651) (referred to as residual samples or a residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (653) obtaining prediction samples from an address within the reference picture memory (657) may be controlled by a motion vector, and the motion vector is in the form of the symbol (621) for use by the motion compensation prediction unit (653), and the symbol (621) may include, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values extracted from the reference picture memory (657), a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.

[0102] The output samples of the aggregator (655) can be adopted by various loop filtering techniques in the loop filter unit (656). Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream), and the parameters can be used in the loop filter unit (656) as symbols (621) from the parser (620). However, in other embodiments, the video compression techniques can also respond to meta-information obtained during the decoding of the previous (in decoding order) part of the encoded picture or the encoded video sequence, and respond to the previously reconstructed and loop-filtered sample values.

[0103] The output of the loop filter unit (656) can be a sample stream, which can be output to the display device (612) and stored in the reference picture memory (657) for subsequent inter-picture prediction.

[0104] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (620)), the current picture buffer (658) can become part of the reference picture memory (657), and a new current picture buffer can be reallocated before starting to reconstruct the subsequent encoded pictures.

[0105] The video decoder (610) can perform decoding operations according to, for example, the predetermined video compression techniques in the ITU-T H.265 standard. In the sense that the encoded video sequence conforms to the syntax specified by the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can comply with the syntax specified by the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, for example, megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0106] Figure 7FIG. shows a block diagram of a video encoder (703) according to an embodiment of the present disclosure. The video encoder (703) may be used in a mesh encoder (400) for compressing 3D meshes or point clouds. In some example embodiments, the video compression modules (422) and (423) and the video compression module (432) are configured similarly to the encoder (703).

[0107] The video encoder (703) may receive 2D images (such as filled geometric images, filled texture images, etc.) and generate compressed images.

[0108] According to an embodiment of the present disclosure, the video encoder (703) may encode and compress pictures of a source video sequence (image) into an encoded video sequence (compressed image) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller (750). In some embodiments, the controller (750) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the couplings are not labeled in the figure. The parameters set by the controller (750) may include rate control related parameters (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (750) may be used for other suitable functions that relate to optimizing the video encoder (703) for a particular system design.

[0109] In some embodiments, the video encoder (703) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (730) (e.g., responsible for creating symbols based on an input picture to be encoded and reference pictures, such as a symbol stream) and a (local) decoder (733) embedded in the video encoder (703). The decoder (733) may reconstruct the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques considered in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (734). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (734) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs in cases where synchronization cannot be maintained due to, for example, channel errors) is also used in some related technologies.

[0110] The operation of the "local" decoder (733) can be the same as that of the "remote" decoder that has been described in detail above, for example, in connection with Figure 6 the video decoder (610). However, briefly referring also to Figure 6 , when symbols are available and the entropy encoder (745) and the parser (620) are able to encode / decode the symbols losslessly into an encoded video sequence, the entropy decoding part of the video decoder (610), including the parser (620), may not be fully implemented in the local decoder (733).

[0111] In various embodiments of the present disclosure, any decoder technology other than the parsing / entropy decoding present in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, this application focuses on decoder operations. The description of encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology described in detail. A more detailed description is only needed in certain areas and is provided below.

[0112] During operation, in some embodiments, the source encoder (730) can perform motion-compensated predictive coding. Referring to one or more previously encoded pictures designated as "reference pictures" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (732) can encode the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.

[0113] The local video decoder (733) can decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (730). The operation of the coding engine (732) can be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 7 not shown), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (733) replicates the decoding process that can be performed by the video decoder on the reference picture and can store the reconstructed reference picture in the reference picture cache (734). In this way, the video encoder (703) can locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by the remote video decoder.

[0114] The predictor (735) may perform a prediction search for the encoding engine (732). That is, for a new picture to be encoded, the predictor (735) may search in the reference picture memory (734) for sample data (as candidate reference pixel blocks) or certain metadata that can serve as an appropriate prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (735) may operate block by block based on sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (735), it may be determined that the input picture may have a prediction reference obtained from multiple reference pictures stored in the reference picture memory (734).

[0115] The controller (750) may manage the encoding operations of the source encoder (730), including, for example, setting parameters and subgroup parameters for encoding video data.

[0116] The outputs of all the above functional units may be entropy encoded in the entropy encoder (745). The entropy encoder (745) may perform lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0117] The controller (750) may manage the operations of the video encoder (703). During encoding, the controller (750) may assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures may generally be assigned to any of the following picture types:

[0118] An intra picture (I picture), which may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variants of I pictures and their corresponding applications and characteristics.

[0119] A predictive picture (P picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, and the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0120] A bi - predictive picture (B picture), which may be a picture that can be encoded and decoded using intra prediction or inter prediction, and the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for reconstructing a single block.

[0121] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0122] The video encoder (703) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (703) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0123] Video can be in the form of multiple source pictures (images) in a time series. Intra-picture prediction (often simplified to intra-frame prediction) utilizes the spatial correlation within a given picture, while inter-picture prediction utilizes the (temporal or other) correlation between pictures. In an embodiment, the specific picture being encoded / decoded is segmented into blocks, and the specific picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0124] In some embodiments, bidirectional prediction techniques can be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.

[0125] In addition, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.

[0126] According to some embodiments disclosed in the present application, the execution of predictions such as inter-picture prediction and intra-picture prediction is performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression, and the CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In an embodiment, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. Additionally, depending on the temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operations in encoding (encoding / decoding) are performed in units of prediction blocks. Taking the luminance prediction block as the prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0127] In various embodiments, the above grid encoder (400) and grid decoder (500) can be implemented in hardware, software, or a combination thereof. For example, the grid encoder (400) and grid decoder (500) can be implemented using processing circuitry such as one or more integrated circuits (ICs), which operate with or without software, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. In another example, the grid encoder (400) and grid decoder (500) can be implemented as software or firmware including instructions stored in a non-volatile (or non-transitory) computer-readable storage medium. When these instructions are executed by processing circuitry such as one or more processors, the processing circuitry is caused to perform the functions of the grid encoder (400) and / or grid decoder (500).

[0128] In some example embodiments, a bi-degree based method can be used to efficiently encode the connectivity of a polygon mesh with arbitrary face degrees or vertex degrees. This method can be configured to utilize the duality between the original mesh and the dual mesh to encode the connectivity by generating two symbol sequences of vertex degrees (or vertex dimensions) and face degrees (or face dimensions), as Figure 8 shown. Specifically, Figure 8 shows a polygon mesh object containing faces and vertices (vertex / vertice) with various degrees (the number of incident edges of the face or vertex). Faces of degree 5 and vertices of degree 4 are highlighted and pointed to by the Figure 8 arrows in. The vertex degree can be referred to as the valence of the vertex.

[0129] The encoding performance of the bi-degree based method may highly depend on the mesh regularity, which measures the degree of variation of the vertex degrees and / or face degrees of the mesh. The smaller the vertex degree / face degree variation, the more regular the mesh, and the higher the connectivity encoding efficiency. It can be shown that for the worst-case mesh, the bi-degree based method can be close to optimal, where the entropy of the two symbol sequences can reach the Tutte entropy bound of the planar graph, e.g., 2 bits per edge.

[0130] The example bi-degree based method for polygon mesh encoding may be close to optimal in the worst case, but for the general case with large variations in vertex and / or face degrees, polygon mesh compression remains a challenge, and thus more efficient methods for encoding connectivity, geometry, and other mesh properties may be needed. In the further disclosures below, many example embodiments are described to improve the bi-degree based method for polygon mesh compression. These example embodiments can be applied individually or in any form of combination.

[0131] A 3D polygon mesh object can include multiple mesh components. Each mesh component can be a continuous 3D volume entity. The 3D polygon mesh object can be encoded by component under the bi-degree method. These components can be connected to form the object. The mesh can contain multiple objects. Each component can include multiple vertices and multiple faces. Each face can be defined by multiple vertices (with corresponding valence numbers or edge numbers) and multiple edges (with corresponding face degrees). Encoding the connectivity of vertices and faces can form separate symbol sequences. The vertices or faces can be encoded following various scan orders, and the initial vertex or initial face to be encoded can be determined in various example ways.

[0132] In some example embodiments of the dual-degree based mesh coding algorithm, an initial face and / or an initial vertex can be selected for each connected component first. For each connected component, its initial face and initial vertex to be encoded in a separate symbol sequence can be selected as the initial face and initial vertex closest to a pre-specified location. In one example embodiment, the initial face and / or initial vertex of the first or initial connected component to be encoded can be selected as the initial face and / or initial vertex closest to the origin or the centroid of the first component. The initial face and initial vertex of the following connected components can be selected as the initial face and initial vertex closest to the centroid of the corresponding component or the first vertex or the last vertex in the previously or previously encoded connected components.

[0133] In some example embodiments for selecting an initial vertex and / or face for each connected component during encoding, in addition to considering the encoding cost of the first vertex / face of each connected component, the cost of encoding the second vertex / face and the third vertex / face by incremental encoding given the initial vertex / face can also be considered. The incremental encoding of a vertex or a face can be based on the difference from the previous vertex or face, or the difference from other predefined or configured vertices or faces. For example, the initial face / vertex with the minimum total encoding cost among the first 3 faces / vertices can be selected as the initial face / vertex of this connected component. Generally speaking, the initial face / vertex with the minimum total encoding cost among the first N faces / vertices can be selected as the initial face / vertex of this connected component. N can be a positive integer. Here, any suitable metric (such as the L p norm or entropy-coded bits) can be used to calculate the proximity / distance or cost.

[0134] The encoding process of a mesh component or object may involve adding faces into the face symbol sequence one by one. The added faces and their vertices before the current face to be added can be collectively referred to as the visited region. The next face to be added will expand this visited region. In some example embodiments, when a new face is added to the visited region during encoding, the vertices of this face that can be determined according to the relationship between the face to be added and the visited region can be linked to this face, and then the remaining undetermined vertices of this face can be linked as new vertices or split vertices. Figure 9 An example is shown in Figure 9In it, the shaded surface and the vertices therein form visited regions (e.g., two separate visited regions 903 and 905). Next, a 5-degree surface 901 is to be added. The encoding process for adding surface 901 can start with determining a reference vertex 902 among the vertices of the first visited region 903. The vertices of the new surface 901 can be encoded / specified in a specific predefined order. For example, they can be encoded in a counterclockwise scan order. Using this example scan order, the next vertex after the reference vertex 902 will be 904, which can be determined based on the relationship between the new surface 901 and the visited region 903 and linked to the vertex 904 of the visited region 903. In this example, the next vertex of the new surface will be 906, which constitutes a splitting vertex and is not part of the visited region 903 that contains the reference vertex 902. Thus, this vertex 904 of the new surface can be explicitly signaled. The next vertex 908 of the new surface is in the visited region 905, which is also a vertex split from the visited region 903. However, since 908 is in the same visited region as the vertex 906, it can be linked to this vertex of the visited region 905 instead of explicitly signaling 908 for the new surface 901. The next and last vertex 910 of the new surface is in the first visited region that contains the reference vertex 902 and can be linked instead of explicitly signaling 910 for the new surface 901.

[0135] In other words, in this example embodiment, the reference vertex of the new surface can be determined first, and some remaining vertices can still be determined based on the relationship between the surface and the visited regions. In this case, some splitting vertices (such as 908) may still not need to be signaled but can be linked. By checking whether the next vertex can be inferred, the explicit signaling requirement for vertices can be reduced, and more links can be used, thus improving the encoding efficiency (since the cost of a link is lower than that of implicit signaling).

[0136] In dual-degree-based mesh coding, the signaling of split offsets and positions (offsets and positions of splitting vertices) may consume a large amount of connectivity code stream. To reduce the number of splits, new surfaces can be added during the encoding process such that the extended visited regions are as convex as possible. The reference vertex for adding the next surface can also be selected among the vertices of the current visited region with the aim of maintaining a high convexity of this visited region.

[0137] In some example embodiments for selecting the next reference vertex to maintain a high convexity of the visited region, the discrete concavity / curvature of the visited region at each active vertex can be calculated, and the vertex with the maximum concavity can be selected as the next reference vertex for adding the next surface. For example, the discrete concavity / curvature at each vertex can be calculated as the sum of all angles of the visited surfaces adjacent to that vertex, asFigure 10 As shown. In Figure 10 , the currently visited area is again shown by the shaded surface. The vertex shown as 1002 has the maximum concavity (compared to other vertices such as vertex 1006 with a concavity of 120 degrees, the sum of all angles of the surfaces of the visited area at vertex 1002 is ~300 degrees), and thus can be selected as the next reference vertex for adding the next face 1004. For example, after adding face 1004, the next maximum concavity can be 180 degrees, such as at vertex 1006.

[0138] In some alternative embodiments for selecting the next reference vertex, at each vertex of the visited area, a number of connected visited faces (which is equivalent to a number of connected unvisited faces) can be calculated, and then the vertex with the highest number of connected components of visited faces can be selected as the next reference vertex, as Figure 11 shown. In Figure 11 , once again, the visited area is shown as a shaded surface containing three connected components. Vertex 1102 is associated with three sets of connected visited faces 1104, 1106, and 1106. The set of visited connected faces 1104 includes one face, while the sets of faces 1106 and 1108 each include two faces. Among other vertices, the three sets of connected visited faces (or three components) of vertex 1102 are the largest. Thus, vertex 1102 can be selected as the reference vertex for adding the next face.

[0139] Figure 10 and Figure 11 The above example embodiments in Figure 10 or Figure 11 can be combined, or determined based on which provides better coding efficiency. In some examples, when there are multiple vertex candidates for the next reference vertex after applying one of the methods in Figure 10 and Figure 11 , the other method in

[0140] The two-degree based algorithm encodes connectivity by generating two symbol sequences (vertex degree and face degree). In some embodiments, the segmentation offset and position are encoded by the entropy codec together with the remaining vertex degree symbols. To achieve higher coding efficiency, the segmentation offset and position can be encoded separately. Additionally, before encoding the symbol sequence using the entropy codec, the smallest possible number can be subtracted from all symbols in the sequence before encoding. For example, in the face degree sequence, if 2 is used to indicate a virtual face that fills a hole in the mesh and integers greater than 2 are used to signal the face degree, the smallest possible number of the face degree sequence is 2, and 2 can be subtracted (downshifted) from all face degree symbols. In this way, fewer bits may need to be encoded. Then, the decoder will upshift to obtain the original degree.

[0141] In two-degree mesh coding, since vertices with large degrees may have associated faces with small degrees, this correlation can be exploited to use the vertex degree to predict the associated face degree. In some example embodiments, the vertex degree can be used to select the entropy coding context for encoding the associated face degree. For example, if the degree of the current vertex is 3, then the third context (or the corresponding mapped context from multiple contexts) can be used to encode the face degree to encode the degree of its associated face. And in some embodiments, if some vertex degrees are too large, the number of entropy coding contexts used for encoding the face degree can be restricted to reduce the computational cost. For example, if the maximum vertex degree is 20, then it may not be necessary to use 20 contexts to encode the face degree, and the number of contexts used for encoding the face degree can be restricted to 10 because all vertices with degrees greater than 10 may have triangles as associated faces. Alternatively, the face degrees can be grouped, such as {0 - 3}, {4 - 5},...... And the degrees of each group can correspond to one context in the set of contexts. The same principle can be applied to predict other symbols, such as using the face degree to predict the segmentation vertex position.

[0142] Similarly, in some example embodiments, the face degree can be used to select an entropy coding context for encoding the associated vertex degree. For example, if the degree of the current face is 3, then we can use the third context (or the corresponding mapped context from a plurality of contexts) to encode the vertex degree. And in some embodiments, if some face degrees are too large, the number of entropy coding contexts used for encoding the vertex degree can be restricted to reduce the computational cost. For example, if the maximum face degree is 20, then it may not be necessary to use 20 contexts to encode the vertex degree, and the number of contexts used for encoding the vertex degree may be restricted to 10 for all faces. Alternatively, the vertex degrees can be grouped, such as {0-3}, {4-5}, ...... And the degrees of each group can correspond to one context in the set of contexts. The same principle can be applied to predict other symbols, such as using the face degree to predict the split vertex position.

[0143] The above example embodiments are at least partially reflected in the following example code.

[0144] Figure 12A flowchart of an example process (1200) according to an embodiment of the present disclosure is shown. The process (1200) starts at step (S1201). In step (S1210), a first encoded region and a second encoded region of a polygon mesh are determined, the polygon mesh including a set of vertices and a set of faces separately encoded by vertex encoding degrees and face encoding degrees of the polygon mesh. In step (S1220), a pivot vertex is selected among the set of vertices in the vertex encoding degrees belonging to the first encoded region to add a polygon face to the set of faces in the face encoding degrees. In step (S1230), for the polygon face, one or more vertices of the first encoded region among the set of vertices are determined. In step (S1240), the one or more vertices are linked to the polygon face in the face encoding degrees of the polygon mesh. In step (S1250), a split vertex for the polygon face in the second encoded region is identified among the set of vertices. In step (S1250), the split vertex is explicitly signaled in the vertex encoding degrees and the split vertex is linked to the polygon face in the face encoding degrees. In step (S1260), a vertex adjacent to the split vertex in the second encoded region is identified to link the vertex to the polygon face in the face encoding degrees. The process (1200) stops at (S1299).

[0145] Figure 13 A flowchart of an example process (1300) according to an embodiment of the present disclosure is shown. The process (1300) starts at step (S1301). In step (S1310), a bitstream of a polygon mesh is received. In step (S1320), a first sub-bitstream including the encoded vertices of the polygon mesh and a second sub-bitstream including the encoded faces of the polygon mesh are extracted from the bitstream. In step (S1330), a context for entropy decoding the encoded vertices in the first sub-bitstream is selected based on at least one subset of the encoded faces in the second sub-bitstream, or a context for entropy decoding the encoded faces in the second sub-bitstream is selected based on at least one subset of the encoded vertices in the first sub-bitstream. In step (S1340), the encoded vertices or the encoded faces are entropy decoded using the selected context. The process (1300) stops at (S1399).

[0146] Figure 14A flowchart of an example process (1400) in accordance with an embodiment of the present disclosure is shown. The process (1400) begins at step (S1401). In step (S1410), a first sub-bitstream of a bitstream including vertices of a polygon mesh is generated. In step (S1420), a second sub-bitstream separated from the first sub-bitstream and including faces of the polygon mesh is generated. In step (S1430), a context for entropy decoding the encoded vertices in the first sub-bitstream is selected based on at least one subset of the encoded faces in the second sub-bitstream, or a context for entropy decoding the encoded faces in the second sub-bitstream is selected based on at least one subset of the encoded vertices in the first sub-bitstream. In step (S1440), the vertices or faces are entropy encoded using the selected context. The process (1400) stops at (S1499).

[0147] The techniques disclosed in the present disclosure can be used alone or in any combination in any order. Additionally, each of the techniques (e.g., methods, embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In some examples, the one or more processors execute a program stored in a non-volatile computer-readable medium.

[0148] The above techniques can be implemented as computer software by computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 15 A computer system (1500) is shown that is adapted to implement certain embodiments of the disclosed subject matter.

[0149] The computer software can be encoded in any suitable machine code or computer language, creating code including instructions through mechanisms such as assembly, compilation, linking, etc., which can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.

[0150] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0151] Figure 15 The components shown for the computer system (1500) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependency on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system (1500).

[0152] A computer system (1500) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs from one or more human users through tactile inputs (such as keyboard inputs, swipes, data glove movements), audio inputs (such as sounds, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0153] The human-machine interface input device may include one or more of the following (only one is shown): keyboard (1501), mouse (1502), touchpad (1503), touch screen (1510), data glove (not shown), joystick (1505), microphone (1506), scanner (1507), camera (1508).

[0154] The computer system (1500) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback through the touch screen (1510), data glove (not shown), or joystick (1505), but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1509), headphones (not shown)), visual output devices (such as screens (1510) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touch screen input functionality and may or may not have tactile feedback functionality - some of which may output two-dimensional visual output or output beyond three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0155] The computer system (1500) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) (1520) or similar media (1521), thumb drives (1522), removable hard disk drives or solid-state drives (1523), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0156] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0157] The computer system (1500) may also include an interface to one or more communication networks. For example, the network may be wireless, wired, optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. The network also includes local area networks such as Ethernet, wireless local area network, cellular network (GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital network (including cable television, satellite television, and terrestrial television broadcasting), vehicular and industrial networks (including CANBus), and so on. Some networks typically require an external network interface adapter for connection to some common data ports or peripheral buses (1549) (e.g., the USB port of the computer system (1500)); other systems are typically integrated into the core of the computer system (1500) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (1500) can communicate with other entities. The communication may be one-way, only for receiving (e.g., wireless television), one-way only for sending (e.g., CAN bus to some CAN bus devices), or two-way, such as through a local or wide area digital network to other computer systems. Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0158] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be connected to the core (1540) of the computer system (1500).

[0159] The core (1540) may include one or more central processing units (CPUs) (1541), a graphics processing unit (GPU) (1542), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (1543), a hardware accelerator (1544) for specific tasks, etc. These devices, as well as a read-only memory (ROM) (1545), a random access memory (1546), an internal mass storage (e.g., an internal non-user-accessible hard disk drive, a solid-state drive, etc.) (1547), etc. may be connected through a system bus (1548). In some computer systems, the system bus (1548) may be accessed in the form of one or more physical plugs so as to be expandable by additional central processing units, graphics processing units, etc. Peripheral devices may be directly attached to the system bus (1548) of the core or connected through a peripheral bus (1549). The architecture of the peripheral bus includes an external controller interface PCI, a universal serial bus USB, etc.

[0160] The CPU (1541), GPU (1542), FPGA (1543), and accelerator (1544) can execute certain instructions, which, when combined, can constitute the above-mentioned computer code. This computer code can be stored in the ROM (1545) or RAM (1546). Transitional data can also be stored in the RAM (1546), while permanent data can be stored in, for example, the internal mass storage (1547). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1541), GPUs (1542), mass storage (1547), ROM (1545), RAM (1546), etc.

[0161] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this application, or they can be media and code well-known and available to those skilled in the art of computer software.

[0162] By way of example and not limitation, a computer system having an architecture (1500), particularly a core (1540), can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such a computer-readable medium can be a medium associated with the above-mentioned user-accessible mass storage, as well as a specific memory of the non-volatile core (1540), such as the core internal mass storage (1547) or ROM (1545). The software implementing the various embodiments of this application can be stored in such a device and executed by the core (1540). Depending on specific needs, the computer-readable medium can include one or more storage devices or chips. The software can cause the core (1540), particularly the processors therein (including the CPU, GPU, FPGA, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in the RAM (1546) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise incorporated in a circuit (e.g., the accelerator (1544)), which can operate in place of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to the computer-readable medium can include a circuit (such as an integrated circuit (IC)) that stores and executes the software, a circuit that contains the execution logic, or both. This application encompasses any suitable combination of hardware and software.

[0163] Although the present application has described multiple exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments all fall within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and thus fall within the spirit and scope of the present application.

Claims

1. A method for encoding a polygonal mesh, characterized in that include: Determining a first encoded region and a second encoded region of the polygonal mesh, the polygonal mesh comprising a set of vertices and a set of faces that are separately encoded at a vertex encoding degree and a face encoding degree, respectively, of the polygonal mesh; selecting a reference vertex from among the set of vertices in the vertex encoding degree belonging to the first encoded region to add a polygonal face to the set of faces in the face encoding degree; For the polygonal face, determining at least one vertex of the first encoded region in the vertex set; linking the at least one vertex to the polygonal face in the face encoding of the polygonal mesh; identifying, among the set of vertices, splitting vertices for the polygonal face in the second encoded region; explicitly signaling the split vertex in the vertex-encoded degree and linking the split vertex to the polygonal face in the face-encoded degree; as well as Vertices in the second encoded region that are immediately adjacent to the segmentation vertex are identified to link the vertices to the polygonal face in the face encoding.

2. The method according to claim 1, characterized in that The first encoded region and the second encoded region are discontinuous.

3. The method according to claim 1, characterized in that The reference vertex in the first encoded region for the polygonal face to be added is selected from the set of vertices by selecting a candidate vertex with a maximum concavity from the first encoded region.

4. The method according to claim 3, characterized in that The concavity of each vertex is quantified by the sum of all angles of faces adjacent to each vertex in the first encoded region.

5. The method according to any one of claims 1 to 4, characterized in that The reference vertex in the first encoded region for the polygonal face to be added is selected from the set of vertices by selecting a candidate vertex of a set of unencoded faces having a maximum number of connections from the first encoded region.

6. The method according to any one of claims 1 to 4, characterized in that The set of vertices comprises a first sequence of symbols, the set of faces comprises a second sequence of symbols, and wherein the encoded segmentation offsets and positions are separated from other offsets and positions within the first sequence of symbols.

7. The method according to any one of claims 1 to 4, characterized in that The vertex set includes a first symbol sequence, the face set includes a second symbol sequence, and wherein, prior to entropy encoding, values ​​of at least one subset of symbols in at least one of the first symbol sequence and the second symbol sequence are offset downward by a minimum value, the minimum value being associated with the at least one subset of symbols.

8. The method according to claim 7, characterized in that The at least one subset of symbols includes the number of edges of at least two faces, and the offset is 2.

9. The method according to any one of claims 1 to 7, characterized in that The initial face among the face set and the initial vertex among the vertex set of an initial mesh component of the polygonal mesh to be encoded are determined as the face and vertex of the initial mesh component that are closest to the centroid of the initial mesh component or closest to the origin of the polygonal mesh, and the next initial face among the face set and the next initial vertex among the vertex set of a next mesh component of the polygonal mesh to be encoded are determined as the face and vertex of the next mesh component that are closest to the centroid of the next mesh component or closest to the last encoded vertex of the initial mesh component, whichever provides a smaller encoding cost.

10. The method according to any one of claims 1 to 7, characterized in that An initial face from the set of faces and / or an initial vertex from the set of vertices used to encode the mesh component of the polygonal mesh is determined as the face and vertex that provides the minimum encoding cost together with the N next faces and / or vertices of the mesh component among the other faces and / or vertices of the mesh component.

11. A method for decoding a code stream of a polygonal mesh, characterized in that: include: Receiving the code stream; Extracting from the code stream a first sub-code stream including encoded vertices of the polygon mesh and a second sub-code stream including encoded surfaces of the polygon mesh; Selecting a context to entropy decode the encoded vertices in the first substream based on at least a subset of the encoded surfaces in the second substream, or to entropy decode the encoded surfaces in the second substream based on at least a subset of the encoded vertices in the first substream; as well as The encoded vertex or the encoded surface is entropy decoded using the selected context.

12. The method according to claim 11, characterized in that The method comprises selecting the context to entropy decode the encoded vertices in the first sub-stream by: identifying the number of faces that include the encoded vertex; and The context for decoding the encoded vertex is selected from a context set according to the number of faces.

13. The method according to claim 12, characterized in that The context set is mapped to a set of numbers of faces.

14. The method according to claim 13, characterized in that The highest numbers of faces are mapped to a context in the set of contexts.

15. The method according to any one of claims 11 to 14, characterized in that The method comprises selecting the context to entropy decode the coded surface in the second substream by: identifying the number of vertices of the encoded face; and The context for decoding the encoded surface is selected from a context set according to the number of the vertices.

16. The method according to claim 15, characterized in that A set mapping the context set to a number of vertices.

17. The method according to claim 16, characterized in that The plurality of highest numbers of vertices are mapped to a context in the set of contexts.

18. A method for encoding a polygonal mesh, characterized in that include: generating a first sub-codestream of a codestream including vertices of the polygonal mesh; generating a second sub-stream separated from the first sub-stream and including the surface of the polygonal mesh; Selecting a context to entropy encode vertices in the first substream based on at least a subset of the surfaces in the second substream, or to entropy encode surfaces in the second substream based on at least a subset of the vertices in the first substream; as well as The vertex or the face is entropy encoded using the selected context.

19. An encoder comprising at least one processor and a memory for storing instructions, wherein the at least one processor is configured to execute the instructions to perform the method of any one of claims 1-10 and 18.

20. A decoder comprising at least one processor and a memory for storing instructions, wherein: The at least one memory is configured to execute the instructions to perform the method of any one of claims 11 to 17.