Competition-based displacement skipping for grid compression

By generating a set of candidate predicted values ​​and selecting the best candidate predicted value index, the problem of low dynamic grid compression efficiency in the prior art is solved, and more efficient displacement coding and finer grid reconstruction are achieved, reducing transmission costs.

CN120266170APending Publication Date: 2025-07-04TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005017.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-09-20
Filing Date
2024-09-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing grid compression standards cannot effectively handle dynamic grids with time-varying connection information and optional time-varying attribute diagrams. Especially under real-time constraints, it is difficult to achieve efficient lossy and lossless compression, which affects the widespread application of 3D content on multiple platforms and devices.

Method used

By generating a set of candidate prediction values, selecting the best candidate prediction value index, and sending displacement residuals and indexes in the code stream, the displacement signaling cost is reduced, and the competition-based displacement skip method is used for encoding and decoding.

Benefits of technology

It realizes the reduction of displacement encoding bit amount while maintaining visual quality, improves encoding efficiency, can rebuild the grid more finely, and reduces transmission costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005415287900000011
    Figure HDA0005415287900000011
  • Figure HDA0005415287900000021
    Figure HDA0005415287900000021
  • Figure HDA0005415287900000031
    Figure HDA0005415287900000031
Patent Text Reader

Abstract

A method includes receiving a polygon mesh including a plurality of vertices; generating a candidate predicted value set, wherein each candidate predicted value in the candidate predicted value set corresponds to a corresponding displacement vector between a vertex from the plurality of vertexes and a vertex to be encoded; selecting a candidate predicted value from the candidate predicted value set; and generating a code stream including at least one candidate predictor index, the at least one candidate predictor index corresponding to the selected candidate predictor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Application No. 63 / 539,780, filed on September 21, 2023, U.S. Provisional Application No. 63 / 539,779, filed on September 21, 2023, and U.S. Application No. 18 / 891,201, filed on September 20, 2024, the disclosures of each of which are hereby incorporated by reference in their entireties. Technical Field

[0003] This disclosure relates to a set of advanced video coding techniques. More specifically, this disclosure relates to a method and apparatus for competition - based displacement skipping for network compression. Background Art

[0004] Advances in three - dimensional (3D) acquisition, modeling, and rendering technologies have led to the widespread presence of 3D content across multiple platforms and devices. Today, a baby's first steps can be captured in one location and viewed (and perhaps interacted with) by grandparents in another location, enjoying a fully immersive experience with the child. However, to achieve this realism, the models are becoming increasingly complex, and the creation and consumption of these models are associated with large amounts of data. 3D meshes are commonly used to represent such immersive content.

[0005] A mesh is composed of multiple polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in three - dimensional space and information on how the vertices are connected (referred to as connectivity information). Optionally, vertex attributes (such as color, normal, etc.) can be associated with the mesh vertices. Attributes can also be associated with the surface of the mesh through mapping information that parameterizes the mesh using two - dimensional attribute maps (2D attribute maps). Such mappings are typically described by a set of parametric coordinates (referred to as UV coordinates or texture coordinates) associated with the mesh vertices. 2D attribute maps are used to store high - resolution attribute information such as texture, normal, displacement, etc., which can be used for various purposes such as texture mapping and shading.

[0006] Since a dynamic mesh sequence may consist of a large amount of time-varying information, a dynamic mesh sequence may require a large amount of data. Therefore, efficient compression techniques are needed to store and transmit such content. Previously, the mesh compression standards IC, MESHGRID, and FAMC developed by the Moving Picture Experts Group (MPEG) were used to process dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards did not consider time-varying attribute maps and connectivity information. Digital Content Creation (DCC) tools typically generate such dynamic meshes. On the other hand, especially under real-time constraints, it is challenging for volumetric acquisition techniques to generate dynamic meshes with constant connectivity. Existing standards do not support this type of content. MPEG plans to develop a new mesh compression standard to directly handle dynamic meshes with time-varying connectivity information and optionally time-varying attribute maps. The standard targets lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, Augmented Reality (AR), and Virtual Reality (VR). The standard also considers features such as random access and scalable / progressive coding. Summary of the Invention

[0007] According to one aspect of the present disclosure, an encoding method, executed by at least one processor, the encoding method comprising: receiving a polygon mesh including a plurality of vertices; generating a set of candidate prediction values, each candidate prediction value in the set of candidate prediction values corresponding to a respective displacement vector between a vertex from the plurality of vertices and a vertex to be encoded; selecting a candidate prediction value from the set of candidate prediction values; and generating a bitstream including at least one candidate prediction value index corresponding to the selected candidate prediction value.

[0008] According to one aspect of the present disclosure, a decoding method, executed by at least one processor, the decoding method comprising: receiving a bitstream including an encoded polygon mesh; generating a set of candidate prediction values associated with the encoded polygon mesh; selecting a candidate prediction value from the set of candidate prediction values; generating a set of candidate prediction values, each candidate prediction value in the set of candidate prediction values corresponding to a respective displacement vector between a vertex from the plurality of vertices and a vertex to be encoded; selecting a candidate prediction value from the set of candidate prediction values; and decoding one or more vertices of the encoded polygon mesh using the selected candidate prediction value.

[0009] According to one aspect of the present disclosure, a coding method is executed by at least one processor. The coding method includes: generating a bitstream of a polygon mesh including a plurality of vertices, wherein a set of candidate prediction values is generated, each candidate prediction value in the set of candidate prediction values corresponding to a respective displacement vector between a vertex from the plurality of vertices and a vertex to be coded, wherein a candidate prediction value is selected from the set of candidate prediction values, and wherein a candidate prediction value index corresponding to the selected candidate prediction value is included in the bitstream. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0011] Figure 1 is a schematic diagram of a block diagram of a communication system according to an embodiment of the present disclosure.

[0012] Figure 2 is a schematic diagram of a block diagram of a streaming system according to an embodiment of the present disclosure.

[0013] Figure 3 illustrates vertex prediction using edge-based interpolation according to an embodiment of the present disclosure.

[0014] Figure 4 illustrates exemplary candidate prediction values and associated displacements according to an embodiment of the present disclosure.

[0015] Figure 5 illustrates a flowchart of an exemplary coding process according to an embodiment of the present disclosure.

[0016] Figure 6 illustrates a flowchart of an exemplary decoding process according to an embodiment of the present disclosure.

[0017] Figure 7 illustrates an example diagram of a computer system according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0018] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.

[0019] The foregoing disclosure provides illustration and description, but is not exhaustive or intended to limit the implementation to the precise forms disclosed. Modifications and variations are possible in light of the above disclosure, or may be acquired from practice of the embodiments. Additionally, one or more features or components of one embodiment may be incorporated into another embodiment (or one or more features of another embodiment) or combined with another embodiment (or one or more features of another embodiment). Further, in the flowcharts and descriptions of operations provided below, it should be understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least in part), and the order of one or more operations may be switched.

[0020] It is clear that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods does not limit the embodiments. Accordingly, the operations and behavior of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware may be designed based on the description herein to implement the systems and / or methods.

[0021] Even though specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible embodiments includes the combination of each dependent claim in the claim set with every other claim.

[0022] Unless explicitly described, any element, action, or instruction used herein should not be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be interchangeable with "one or more". If only one item is intended, the term "one" or similar language is used. Further, as used herein, the terms "having", "have", "possess", "include", "comprise", etc. are intended to be open - ended terms. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on". Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B simultaneously.

[0023] References to "one embodiment", "an embodiment", or similar language throughout this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present solution. Thus, the phrases "in one embodiment", "in an embodiment", and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0024] In addition, in one or more embodiments, the features, advantages, and characteristics described in this disclosure may be combined in any suitable manner. Those skilled in the relevant art will recognize, based on the description herein, that the present disclosure may be practiced without one or more specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be found in certain embodiments that may not be present in all embodiments of the present disclosure.

[0025] Reference Figures 1 to 2 , one or more embodiments of the present disclosure for implementing the encoding structure and decoding structure of the present disclosure are described.

[0026] Figure 1 FIG. shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 may include at least two terminals 110, 120 interconnected by a network 150. For unidirectional transmission of data, the first terminal 110 may encode video data at a local location for transmission over the network 150 to another terminal 120, and the video data may include mesh data. The second terminal 120 may receive the encoded video data of another terminal from the network 150, decode the encoded video data, and display the recovered video data. Unidirectional data transmission is relatively common in applications such as media services.

[0027] Figure 1 FIG. shows a second pair of terminals 130, 140 provided to support bidirectional transmission of encoded video, which may occur, for example, during a video conference. For bidirectional data transmission, each terminal 130, 140 may encode video data collected at a local location for transmission over the network 150 to another terminal. Each terminal 130, 140 may also receive the encoded video data sent by another terminal, may decode the encoded data, and may display the recovered video data on a local display device.

[0028] In Figure 1Among them, terminals 110 to 140 may be, for example, servers, personal computers, smart phones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 150 represents any number of networks that transfer encoded video data between terminals 110 to 140, including, for example, wired and / or wireless communication networks. Communication network 150 may exchange data in circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise explained below, the architecture and topology of network 150 may be irrelevant to the operation of the present disclosure.

[0029] Figure 2 Illustrated is the setup of a video encoder and a video decoder in a streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter may be used with other video-supported applications, including, for example, video conferencing, digital television, storing compressed video on digital media including CDs, DVDs, memory sticks, and so on.

[0030] As Figure 2 shown, the streaming system 200 may include an acquisition subsystem 213, which includes a video source 201 and an encoder 203. The streaming system 200 may also include at least one streaming server 205 and / or at least one streaming client 206.

[0031] For example, the video source 201 may create a sample stream 202, which includes a 3D mesh and metadata associated with the 3D mesh. For example, the video source 201 may include a 3D sensor (such as a depth sensor) or 3D imaging technology (such as one or more digital cameras), and a computing device configured to generate a 3D mesh using data received from the 3D sensor or 3D imaging technology. Compared with the encoded video bitstream, the sample stream 202 with a high data volume may be processed by an encoder 203 coupled to the video source 201. The encoder 203 may include hardware, software, or a combination of both to implement or enforce various aspects of the disclosed subject matter described in more detail below. The encoder 203 may also generate an encoded video bitstream 204. Compared with the uncompressed stream 202, the encoded video bitstream 204 with a lower data volume may be stored on the streaming server 205 for future use. One or more streaming clients 206 and 207 may access the streaming server 205 to retrieve video bitstreams 208 and 209, respectively, which may be copies of the encoded video bitstream 204.

[0032] The streaming client 207 may include a video decoder 210 and a display 212. For example, the video decoder 210 may decode a video bitstream 209 (which is an input copy of the encoded video bitstream 204) and create an output video sample stream 211 that can be rendered on the display 212 or another display device (not depicted). In some streaming systems, the video bitstreams 204, 208, and 209 may be encoded according to certain video coding / compression standards.

[0033] In one or more examples, to efficiently represent a mesh signal, a subset of mesh vertices may first be encoded along with connectivity information between the mesh vertices. In the original mesh, since these vertices are subsampled from the original mesh, there may be no connections between these vertices. There are different ways to generate the connectivity information between the vertices. Therefore, this subset of mesh vertices is referred to as a base mesh or base vertices.

[0034] However, other vertices may be predicted by applying interpolation between two or more decoded mesh vertices. A predictor vertex will have a geometric position along the edge of two connected existing vertices, and thus the geometric information of the predicted value can be calculated based on the adjacent decoded vertices. In some cases, after decoding the base vertices (e.g., Figure 3 the solid triangle 300 in Figure 3 ), the displacement vector or prediction error from the vertex to be encoded to the vertex prediction value is further encoded. These vertices may be referred to as layer 0, and interpolation can be performed between these base vertices along the connected edges. For example, the midpoint of each edge can be generated as the prediction value. Therefore, the geometric position of these interpolation points is the (weighted) average of two adjacent decoded vertices ( Figure 3 the dotted points in the reference mesh 302 in Figure 3 ). When there is more than one midpoint between two already decoded vertices, it can also be done in a similar manner. Therefore, the actual vertex to be encoded (

[0035] Figure 3 layer 1 vertex) can be reconstructed by adding the displacement vector to the prediction value. After decoding these additional vertices, the connection can still be maintained between the new decoded vertices and the existing base vertices. In addition, the connection between the new decoded vertices can be further established. Together with the base vertices, by connecting these new decoded vertices and the base vertices, more intermediate vertex prediction values can be generated along the new edges ( Figure 3 reference mesh 304, layer 2 vertices). Therefore, there may be more actual vertices to be decoded with the relevant displacement vectors. Such interpolation-based vertex prediction can continue until the total number of layers is reached and then stop.

[0035] All displacement vectors for all layers (including displacement vectors for layer 0 or base vertices) can be processed using various methods, such as grouping these displacement vectors together to perform transformation and entropy encoding.

[0036] The displacement is not limited by the subdivision of the base mesh. The displacement can be generated by other mesh coding tools, such as symmetry coding, etc. In all cases, these displacements need to be predicted to achieve improved compression. A typical method of displacement coding is the so-called "lifting" scheme. In this scheme, first, the displacement field is updated using the reconstructed quantized base mesh to generate an updated displacement field. This process takes into account the difference between the reconstructed base mesh and the original base mesh. Then a wavelet transform is applied to generate a set of wavelet coefficients. Then the wavelet coefficients are quantized and packed into a 2D image / video or directly encoded using an arithmetic encoder.

[0037] Mesh compression can be a lossy process: reducing the size of the bitstream (the information to be transmitted) at the cost of introducing distortion in the decoded mesh. The trade-off between bitrate savings and the resulting distortion is a typical problem in mesh compression. Although creating displacements through coding tools enables better mesh reconstruction at the decoder, the cost of displacement information in a typical bitstream is high. Reducing this cost while maintaining quality is the main problem addressed in the present disclosure.

[0038] Embodiments of the present disclosure directly propose methods and systems for predicting displacements in 3D compression (such as mesh compression). The embodiments can be applied individually or in any form of combination. The disclosed methods and systems are not limited to mesh compression.

[0039] The following operations can be performed for encoding.

[0040] A set of displacement candidate prediction values (see candidate prediction values) containing one or more candidate prediction values (see candidate prediction value sets) can be created. For each displacement to be predicted: (A) For each candidate prediction value in the set, (i) calculate the candidate prediction value: calculate the candidate prediction value according to adjacent information (see candidate prediction value), (ii) calculate the displacement residual (see residual calculation), and (iii) evaluate the displacement cost associated with the candidate prediction value (see selection criteria); (B) select the best candidate prediction value (see selection criteria); (C) signal the displacement residual and the index of the selected candidate prediction value in the bitstream (see signaling); alternatively, in the absence of signaling, infer the selected prediction value candidate.

[0041] The following operations can be performed for decoding.

[0042] For each displacement: (A) Parse the bitstream to retrieve the displacement residual and retrieve the selected candidate prediction value index from the set of generated candidate prediction values; alternatively, in the case where the index is not signaled, infer the selected prediction value candidate from the set of generated candidate prediction values; (B) Calculate the selected candidate prediction value (see candidate prediction value); (C) Calculate the reconstructed displacement (see inverse residual calculation) or assign the selected candidate prediction value to the displacement.

[0043] Embodiments of the present disclosure provide many advantages. First, predicting the displacement and sending the residual and index reduces the signaling cost of the displacement. This enables the use of fewer bits for a mesh with equivalent visual quality, or the allocation of more bits to other information to improve visual quality. Additionally, using fewer bits to encode the displacement enables better subdivision (more vertices), resulting in finer fidelity relative to the original mesh for the same displacement cost. Further, skipping the transmission of the displacement and predicting the displacement due to the transmission of the index reduces the signaling cost of the displacement. Compared with competition-based displacement prediction, embodiments of the present disclosure will achieve a lower bitrate at the cost of higher transmission.

[0044] According to one or more embodiments, candidate prediction values can be defined to predict the displacement. Figure 4 A non-exhaustive list of candidate prediction values is shown.

[0045] In Figure 4 , vi (where i is an index representing different positions) can represent vertices in a local neighborhood, and d(vi) can represent the associated displacement. In one or more examples, i.e., in the preferred embodiment, the candidate prediction value can be calculated by the following formula:

[0046] Formula (1): cand_pred = 1 / SUM(wi)*SUM(wi*d(vi)), i = 1 to n,

[0047] where wi is a weight associated with some local features, such as the distance between vertices and displacements, etc.; SUM(*) is the sum of a set of variables in the parentheses. wi can be recovered at the decoder (no need to send).

[0048] Figure 4 The faces in the mesh 400 or mesh 402 of are faces of the base mesh or sub-faces obtained after subdividing the base mesh.

[0049] Formula (1) includes the following cases:

[0050] - Taking 0 displacements as the candidate prediction value (all wi are equal to 0).

[0051] - Only use one displacement as the candidate prediction value (all wi are equal to 0 except one wi). This enables considering only one adjacent vertex.

[0052] - The displacement is the average of the displacements of two edge vertices (all wi are equal to 0 except the two wi associated with the edge vertices which are equal to 1 / 2).

[0053] - In particular, if a vertex is generated by interpolating from two adjacent vertices sharing the same edge, the displacement of the vertex can be predicted by the weighted average of the two "parent" vertices.

[0054] According to one or more embodiments, the candidate prediction value can be obtained by a machine learning-based algorithm, where, given features consisting of local information or general grid information, training is performed to learn the most suitable candidate prediction value. In one or more examples, the candidate prediction value can be a curve fitting algorithm (considering curvature based on multiple vertex positions), the equation of the surface can be calculated, and the prediction value can be derived.

[0055] In one or more examples, the displacement d(vi) is the decoded displacement available at both the encoder and the decoder to avoid any drift caused by quantization during decoder reconstruction of the displacement. In one or more examples, the displacement d(vi) is the original displacement, thus improving the prediction accuracy at the cost of the drift generated by displacement quantization.

[0056] In one or more examples, the set of candidate prediction values contains one or more candidate prediction values, and subsequently, a best prediction value will be selected from the set of candidate prediction values. The candidate prediction values forming the set can vary as follows: (i) depending on local grid features (curvature, etc.), some candidate prediction values are more relevant than others; (ii) depending on the geometry of the face (size, distance between vertices), some candidate prediction values are more relevant than others; or (iii) depending on the result of a multi-pass approach that applies the first encoding and derives statistics, features of relevant candidate prediction values are extracted.

[0057] When the set varies, the candidate and order selected in the set can be automatically selected based on information available at the decoder or information signaled in the bitstream (an index identifies each set). The number of candidate prediction values in the set can vary. Depending on the complexity of the prediction, whether a smaller or larger set of candidate prediction values is more relevant.

[0058] There may be changes in the set of candidate prediction values at the grid level, sub-grid level, or any differently defined grid partition. As an example, in the set of candidate prediction values, the grid portion representing the human face requires more prediction values than the grid portion representing the human flatback. As another example, in the case of grid recursive partitioning or context level of detail (LOD), different sets can be used at each level. The number of candidates in the set can be signaled in the mesh frame header and applied to all vertices in the mesh frame. Or alternatively, the number of candidates in the set can be signaled and applied at each sub-grid level.

[0059] Having a varying set of prediction values provides many advantages. For example, different parts of the mesh have different characteristics, so different sets of candidate prediction values can be utilized. Thus, the displacement cost is locally optimized. Additionally, while a large set of prediction values will improve prediction efficiency, it will reduce signaling efficiency (the cost of selecting the best candidate prediction value index in a large set of prediction values is high). A compromise solution between prediction efficiency and signaling efficiency is to have a varying set of prediction values.

[0060] In one or more examples, the residual calculation includes calculating the difference (residual_displacement) between the displacement to be predicted and the selected candidate prediction value sel_cand_pred:

[0061] Equation (2): residual_displacement = displacement – sel_cand_pred

[0062] Wherein, the difference represented by the symbol "-" can be an operator that is used to calculate the difference between two objects. In one or more examples, the "-" operator can be the Euclidean distance calculated separately on each of the three components of the displacement. The difference can be a reversible operator.

[0063] In one or more examples, the inverse residual calculation performed on the decoder can include calculating the reconstructed displacement and the selected candidate prediction value based on the transmitted residual displacement, and the selected candidate prediction value is retrieved by the transmitted prediction value index.

[0064] In one or more examples, a residual calculation is performed for each candidate prediction value. The candidate prediction value that minimizes the cost criterion is selected as the best candidate prediction value (sel_cand_pred). This provides an optimal trade-off between the cost of signaling information in the bitstream and the degradation in quality. In one or more examples, the cost can be expressed as:

[0065] Equation (3) J = D + lambda * R

[0066] In Equation (3), lambda is the Lagrange parameter. D is the distortion, which measures the difference between the source displacement and the reconstructed displacement after prediction using the candidate prediction value. For example, the distortion can be approximated by an objective metric such as D2 or Peak Signal-to-Noise Ratio (PSNR). R is the rate associated with all the information that needs to be signaled, or an approximation of that rate. R includes the rate of the residual displacement and the rate of the best candidate prediction value index.

[0067] In one or more examples, the residual displacement is sent in the bitstream. A classical entropy coding algorithm / arithmetic coding algorithm can be used. In one or more examples, the index corresponding to the selected candidate prediction value is sent in the bitstream. The index is associated with each candidate prediction value. In typical coding, the number of bits used for the index will be low for the most frequently selected candidate prediction value. In one or more examples, the index of the prediction value is not sent in the bitstream. The index can be inferred from the available information at the encoder and decoder.

[0068] In one or more examples, multi-face prediction can be used. In multi-face prediction, the same selected candidate prediction value can be used for a group of faces. The result of this is lower prediction efficiency but fewer indices to transmit. The group of faces can be consecutive faces in a given coding order, or faces collected according to some predefined criteria (applicable to the decoder). In one or more examples, the selection criterion can be modified such that the distortion is averaged over all the faces in the group and weighted according to some criteria available at the decoder. For example, larger faces can have a larger weight because they contribute more to the overall distortion. In this multi-face method, the rate can be the sum of the rates for each face in the group.

[0069] In one or more examples, a recursive / hierarchical-based method can be used. In this method, the competition applies as Figure 3The hierarchical method shown. Candidate displacements are provided by the face being currently encoded. Candidate prediction values are calculated based on the displacements obtained from a given level of the hierarchy and the candidate prediction values are used for the next level. While this method can have some scalability (the decomposition of levels can be stopped according to a given criterion), there is a delay in the prediction process and fewer prediction values are available.

[0070] In one or more examples, a non - recursive method can be used. In this method, competition applies to the non - recursive method. Candidate prediction values are calculated based on the displacements obtained from the adjacent faces that have been decoded. In this case, not only the displacements from the vertices of the face can be obtained, but also the displacements of all the sub - faces used for subdivision. All these displacements and their functions can be used to calculate the candidate prediction values.

[0071] In one or more examples, instead of selecting and signaling a candidate prediction value for each face or selecting a candidate prediction value for a group of faces, the following process is applied to a given face or set of faces: (i) independently select the best candidate prediction value for each displacement; or (ii) use the most frequently selected candidate prediction value to predict all the prediction values of the given face or given set of faces and signal the index of the most frequently selected candidate prediction value.

[0072] In one or more examples, both competition - based displacement skipping and competition - based displacement skipping are used jointly to give the opportunity to skip residuals or provide residuals in various situations with different candidate prediction values. The choice of one method or the other is signaled in the bitstream or inferred from the information available on the encoder and decoder.

[0073] Figure 5 An exemplary encoding process 500 according to one or more embodiments is shown. The encoding process 500 can be performed by the encoder 203( Figure 2 ).

[0074] The encoding process 500 begins at operation S502, in which a polygon mesh is received. The polygon mesh can correspond to one or more surfaces of an object in a picture.

[0075] The encoding process proceeds to operation S504, in which a set of candidate prediction values is generated. Any of the methods described above can be used to generate the set of candidate prediction values.

[0076] The encoding process proceeds to operation S506, in which a candidate prediction value is selected from the set of candidate prediction values. The selected candidate prediction value can be the one that minimizes a cost function.

[0077] The encoding process proceeds to operation S508, in which a bitstream including a candidate prediction value index corresponding to a selected candidate prediction value is included.

[0078] Figure 6 An exemplary decoding process 600 according to one or more embodiments is shown. The decoding process 600 can be performed by the decoder 210( Figure 2 ).

[0079] The decoding process begins at operation S602, in which a bitstream including an encoded polygon mesh is received.

[0080] The decoding process begins at operation S604, in which a set of candidate prediction values is generated. The set of candidate prediction values can be generated by the decoder 210 according to any of the above methods.

[0081] The decoding process begins at operation S606, in which a candidate prediction value is selected from the set of candidate prediction values. The selected candidate prediction value can be the candidate prediction value that minimizes a cost function.

[0082] The decoding process begins at operation S608, in which one or more vertices of the encoded polygon mesh are decoded using the selected candidate prediction value.

[0083] The above techniques can be implemented using computer-readable instructions as computer software and physically stored on one or more computer-readable media. For example, Figure 7 A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0084] Any suitable machine code or computer language can be used to encode the computer software, and any suitable machine code or computer language can be assembled, compiled, linked, or otherwise processed to create code including instructions that can be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode, or the like.

[0085] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0086] Figure 7The components of the computer system 700 shown are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement related to any one or combination of components shown in the non-limiting embodiments of the computer system 700.

[0087] The computer system 700 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as: tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as voice, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0088] The human-machine interface input devices may include one or more of the following (only one of each is shown): keyboard 701, mouse 702, touchpad 703, touch screen 710, data glove (not shown), joystick 705, microphone 706, scanner 707, camera 708.

[0089] The computer system 700 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (such as tactile feedback through the touch screen 710, data glove, or joystick 705, but may also be tactile feedback devices that are not input devices). For example, such human-machine interface output devices may be audio output devices (e.g., speaker 709, headphones (not shown)), visual output devices (e.g., screen 710 including a cathode ray tube (CRT) screen, liquid-crystal display (LCD) screen, plasma screen, organic light-emitting diode (OLED) screen, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, some of which are capable of outputting two-dimensional visual output or output beyond three dimensions through means such as stereoscopic picture output, virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0090] The computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW 720 with CD / DVD or similar media 721, thumb drives 722, removable hard disk drives or solid state drives 723, traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and so on.

[0091] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0092] The computer system 700 may also include an interface to one or more communication networks. For example, the network may be a wireless network, a wired network, an optical network. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a low-latency network, and so on. Examples of networks include local area networks such as Ethernet, wireless LAN (Wireless Local Area Network, WLAN), cellular networks including Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc., television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial television broadcasting, vehicle and industrial networks including CAN bus, and so on. Some networks typically require an external network interface adapter connected to certain general-purpose data ports or peripheral buses 749 (e.g., the Universal Serial Bus (USB) port of the computer system 700, for example); other network interface adapters are typically integrated into the core of the computer system 700 by connecting to the system bus described below (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). The computer system 700 may communicate with other entities using any of these networks. Such communication may be one-way reception only (e.g., television broadcasting), one-way transmission only (e.g., CAN bus to certain CAN Bus devices), or two-way, such as connecting to other computer systems using a local area network or a wide area digital network. Such communication may include communication to a cloud computing environment 755. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.

[0093] The above-described human-machine interface devices, human-accessible storage devices, and network interfaces may be connected to the core 740 of the computer system 700.

[0094] The kernel 740 may include one or more CPUs 741, GPUs 742, dedicated programmable processing units 743 in the form of Field Programmable Gate Areas (FPGAs), hardware accelerators 744 for certain tasks, and so on. These devices, together with a Read-only Memory (ROM) 745, a Random Access Memory 746, and internal mass storage 747 such as an internal hard disk drive that is not user-accessible, a Solid State Drive (SSD), etc., may be connected via a system bus 748. In some computer systems, the system bus 748 may be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus 748 of the kernel or connected to the system bus 748 of the kernel via a peripheral bus 749. The architecture of the peripheral bus includes Peripheral Component Interconnect (PCI), USB, etc. A graphics adapter 750 may be included in the kernel 740.

[0095] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute certain instructions that, when combined, may constitute the aforementioned computer code. The computer code may be stored in the ROM 745 or a Random Access Memory (RAM) 746. Transitional data may also be stored in the RAM 746, while permanent data may be stored in, for example, the internal mass storage 747. Fast storage and retrieval of any memory device may be achieved by using a cache memory that may be closely associated with one or more CPUs 741, GPUs 742, mass storage 747, ROM 745, RAM 746, etc.

[0096] Computer-readable media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code that are specially designed and constructed for the purposes of this disclosure, or the media and the computer code may be of the type well-known and available to those skilled in the art of computer software.

[0097] By way of example and not limitation, a computer system having architecture 700, particularly core 740, can provide functionality as one or more processors (including a CPU, GPU, FPGA, accelerator, etc.) execute software included in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of core 740, such as on-core mass storage 747 or ROM 745. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by core 740. Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause core 740, particularly the processors therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described in the present disclosure, including defining data structures stored in RAM 746 and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality that is hardwired logically or otherwise included in circuitry (e.g., accelerator 744) that can operate instead of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate instances, portions that refer to software can include logic and vice versa. In appropriate instances, portions that refer to computer-readable media can include circuitry (e.g., an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0098] Although the present disclosure has described multiple exemplary embodiments, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design various systems and methods that, although not explicitly shown or described in the present disclosure, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A coding method, executed by at least one processor, the method comprising: Receiving a polygon mesh including a plurality of vertices; Generating a set of candidate prediction values, each candidate prediction value in the set of candidate prediction values corresponding to a respective displacement vector between a vertex from the plurality of vertices and a vertex to be coded; Selecting a candidate prediction value from the set of candidate prediction values; And Generating a bitstream including at least one candidate prediction value index corresponding to the selected candidate prediction value.

2. The method according to claim 1, wherein Calculating the at least one candidate prediction value using a weighted sum of one or more displacement vectors of vertices adjacent to at least one candidate prediction value in the set of candidate prediction values.

3. The method according to claim 1, wherein, Based on one or more characteristics of the polygon mesh, providing a higher priority to a first candidate prediction value in the set of candidate prediction values than a second candidate prediction value in the set of candidate prediction values.

4. The method according to claim 3, wherein The one or more characteristics include: the curvature of the polygon mesh, the size of the faces of the polygon mesh, or the distance between vertices.

5. The method according to claim 1, wherein The candidate prediction value selected from the set of candidate prediction values is the candidate prediction value that minimizes a cost function.

6. The method according to claim 5, wherein, The cost function depends on the distortion after prediction using the candidate prediction value, the distortion being associated with the reconstructed displacement.

7. The method according to claim 5, wherein, The cost function depends on the rate associated with signaling information, the information being associated with the bitstream.

8. The method according to claim 1, further comprising: Determining a residual displacement corresponding to the difference between the displacement to be predicted and the selected candidate prediction value, wherein the bitstream includes the residual displacement.

9. The method according to claim 1, wherein The selected candidate prediction value is associated with a plurality of faces in the polygon mesh.

10. The method according to claim 9, wherein, The plurality of faces are consecutive faces in the coding order.

11. The method according to claim 9, wherein, The plurality of faces are weighted according to the size of the faces, wherein larger faces have higher weights.

12. A decoding method, executed by at least one processor, the method comprising: Receiving a bitstream including an encoded polygon mesh; Generating a set of candidate prediction values associated with the encoded polygon mesh; Selecting a candidate prediction value from the set of candidate prediction values; And Decoding one or more vertices of the encoded polygon mesh using the selected candidate prediction value.

13. The method according to claim 12, further comprising: Parsing the bitstream to extract an index corresponding to the selected candidate prediction value.

14. The method according to claim 12, wherein Inferring an index corresponding to the selected candidate prediction value based on one or more characteristics of the polygon mesh.

15. The method according to claim 12, further comprising: Parsing the bitstream to extract a residual displacement; And Using the extracted residual displacement and the selected candidate prediction value to determine a reconstructed displacement.

16. The method according to claim 12, wherein, Calculating the at least one candidate prediction value using a weighted sum of one or more displacement vectors of vertices adjacent to at least one candidate prediction value in the set of candidate prediction values.

17. The method according to claim 12, wherein Based on one or more characteristics of the polygon mesh, providing a higher priority to a first candidate prediction value in the set of candidate prediction values than a second candidate prediction value in the set of candidate prediction values.

18. The method according to claim 3, wherein The one or more characteristics include: the curvature of the polygonal mesh, the size of the faces of the polygonal mesh, or the distance between vertices.

19. The method according to claim 12, wherein, The candidate prediction value selected from the set of candidate prediction values is the candidate prediction value that minimizes the cost function.

20. An encoding method, executed by at least one processor, the method comprising: generating a bitstream for a polygonal mesh including a plurality of vertices, wherein a set of candidate prediction values is generated, each candidate prediction value in the set of candidate prediction values corresponding to a respective displacement vector between a vertex from the plurality of vertices and a vertex to be encoded, wherein a candidate prediction value is selected from the set of candidate prediction values, wherein a candidate prediction value index corresponding to the selected candidate prediction value is included in the bitstream.