Optimal subgrid coding order for initial vertex selection in position coding

The vertex pair with the minimum prediction cost is selected as the predictor vertex by the two-degree encoder and decoder, and the sub-grid is encoded and decoded, and the encoding sequence is optimized. The problem of inefficient processing of time-varying attribute diagrams and connectivity information in the prior art is solved, and data transmission efficiency is improved.

CN120283266APending Publication Date: 2025-07-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005110.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-23
Filing Date
2024-07-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing grid compression standards fail to effectively handle time-varying attribute graphs and connectivity information in dynamic grids, especially the challenges of generating constant connectivity dynamic grids under real-time constraints, resulting in inefficiencies in data storage and transmission.

Method used

Using a two-degree encoder and decoder, the sub-grid is encoded and decoded by selecting the vertex pair with the minimum predicted cost as the predictor vertex, optimized the encoding order to reduce the bit rate, and accelerated optimization algorithms using new vertex symbols and simplified cost estimation methods.

Benefits of technology

It improves the encoding efficiency of dynamic grids and reduces the bit rate required for data transmission. It is suitable for real-time communication, storage, free viewpoint video, AR and VR applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283266A_ABST
    Figure CN120283266A_ABST
Patent Text Reader

Abstract

Methods and apparatus, including computer code for mesh attribute coding and decoding using a two-degree encoder or decoder configured to cause one or more processors to receive at least two sub-meshes and determine pairs of vertices having a minimum prediction cost, a first vertex of the vertex pair is part of the first subgrid, and a second vertex of the vertex pair is part of the second subgrid. The method may cause one or more processors to set a first vertex of the vertex pair as a predictor vertex for a second subgrid to be encoded. The method may cause the one or more processors to encode a second sub-grid in which a second vertex of the vertex pair is first encoded in the first encoding order.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims priority to U.S. Provisional Application Nos. 63 / 528,635 and 63 / 528,636, filed on July 24, 2023, and U.S. Application No. 18 / 781,135, filed on July 23, 2024. The entire disclosures of the above - mentioned U.S. provisional applications and U.S. applications are incorporated herein by reference. Technical field

[0003] Embodiments of the present disclosure relate to video encoding and decoding. Specifically, embodiments of the present disclosure encode and decode multiple sub - grids, including position encoding in grid motion vector encoding. Background art

[0004] Advanced three - dimensional (3D) representations of the world are enabling more immersive forms of interaction and communication. To achieve photorealism in 3D representations, 3D models are becoming increasingly complex, and a large amount of data is associated with the creation and consumption of these 3D models. 3D meshes are widely used for 3D model immersive content.

[0005] A 3D mesh can include a number of polygons that describe the surface of a volumetric object. A dynamic mesh sequence may require a large amount of data because it may have a large amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content.

[0006] Although previous grid compression standards such as IC, MESHGRID, and FAMC have been developed to handle dynamic meshes with constant connectivity, time - varying geometry, and vertex attributes. However, these standards do not consider time - varying attribute graphs and connectivity information.

[0007] In addition, for volume acquisition techniques, generating a dynamic mesh with constant connectivity, especially under real - time constraints, is also challenging. Existing standards do not support this type of dynamic mesh content.

[0008] As another example, glTF (GL Transmission Format) is a standard developed by the Khronos Group for efficiently transmitting and loading 3D scenes and models through an application. glTF aims to minimize both the size of 3D assets and the runtime processing required for unpacking. A geometric compression extension of glTF 2.0 using Google Draco technology is being developed to reduce the size of glTF models and scenes. Summary of the invention

[0009] According to an embodiment, a method, an apparatus, and a non-transitory medium for encoding and decoding mesh attributes using a bi-degree encoder can be provided. The method can include: receiving at least two sub-meshes, wherein a first sub-mesh of the at least two sub-meshes has been encoded, and wherein a second sub-mesh of the at least two sub-meshes is to be encoded; determining a vertex pair having a minimum prediction cost, wherein a first vertex of the vertex pair is a part of the first sub-mesh, and wherein a second vertex of the vertex pair is a part of the second sub-mesh; setting the first vertex of the vertex pair as a predictor vertex for the second sub-mesh to be encoded; signaling the first vertex of the vertex pair to a bi-degree decoder; and encoding the second sub-mesh in a first encoding order, wherein the second vertex of the vertex pair is first encoded in the second sub-mesh.

[0010] According to an embodiment, a method, an apparatus, and a non-transitory medium for encoding and decoding mesh attributes using a bi-degree decoder can be provided. The apparatus for encoding and decoding mesh attributes using a bi-degree decoder can include: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code can include: a first receiving code configured to cause the at least one processor to receive at least two encoded sub-meshes; a second receiving code configured to cause the at least one processor to receive a predictor vertex from a bi-degree mesh encoder, wherein the predictor vertex is a first vertex of a vertex pair having a minimum prediction cost; a first determining code configured to cause the at least one processor to determine a second vertex based on the predictor vertex, wherein the first vertex of the vertex pair is a part of a first sub-mesh of the at least two encoded sub-meshes, and wherein the second vertex of the vertex pair is a part of a second sub-mesh of the at least two encoded sub-meshes; and a first decoding code configured to cause the at least one processor to decode the second sub-mesh after decoding the first sub-mesh using the second vertex of the vertex pair.

[0011] According to an embodiment, a method, an apparatus, and a non-transitory medium for grid property encoding and decoding can be provided. The non-transitory computer-readable medium storing instructions can include one or more instructions that, when executed by one or more processors of a device for grid property encoding and decoding, cause the one or more processors to perform the following operations: perform a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the bitstream includes at least two sub-grids, a first sub-grid of the at least two sub-grids has been converted, wherein the bitstream includes predictor vertices for a second sub-grid to be converted, wherein a predictor vertex is a first vertex in a vertex pair having a minimum prediction cost, wherein the first vertex in the vertex pair is part of the first sub-grid of the at least two sub-grids, and wherein the second vertex in the vertex pair is part of the second sub-grid of the at least two sub-grids; and encode or decode the second sub-grid using the second vertex in the vertex pair. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0013] Figure 1 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment of the present disclosure.

[0014] Figure 2 is a schematic illustration of a simplified block diagram of a streaming system according to an embodiment of the present disclosure.

[0015] Figure 3 is a schematic illustration of a simplified block diagram of a video encoder and decoder according to an embodiment of the present disclosure.

[0016] Figure 4A is a schematic illustration of a double-degree traversal according to an embodiment of the present disclosure.

[0017] Figure 4B is a schematic illustration of prediction of the first three positions according to an embodiment of the present disclosure.

[0018] Figure 5A is an exemplary diagram of a bidirectional graph showing a sub-grid encoding order according to an embodiment of the present disclosure.

[0019] Figure 5B is an exemplary diagram of a directed graph showing a sub-grid encoding order according to an embodiment of the present disclosure.

[0020] Figure 6 is an exemplary illustration of positions included in a candidate list according to an embodiment of the present disclosure.

[0021] Figure 7 is an exemplary process for encoding a plurality of meshes using predictor vertices belonging to an already encoded mesh according to an embodiment of the present disclosure.

[0022] Figure 8 is an exemplary diagram of a computer system suitable for implementing an embodiment. DETAILED DESCRIPTION

[0023] The proposed features discussed below can be used alone or in any combination. Additionally, embodiments can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0024] Figure 1 FIG. 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The communication system 100 can include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional transmission of data, a first terminal 103 can encode video data at a local location for transmission via the network 105 to another terminal 102. The second terminal 102 can receive the encoded video data of the other terminal from the network 105, decode the encoded data, and display the recovered video data. Unidirectional data transmission can be common in media service applications and the like.

[0025] Figure 1 FIG. 2 shows a second pair of terminals 101 and 104 configured to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal 101 and 104 can encode video data captured at a local location for transmission via the network 105 to the other terminal. Each terminal 101 and 104 can also receive the encoded video data sent by the other terminal, can decode the encoded data, and can display the recovered video data at a local display device.

[0026] In Figure 1Among them, terminals 101, 102, 103, and 104 can be shown as servers, personal computers, and smart phones, but the principles of the present disclosure are not limited thereto. Embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 105 represents any number of networks that transmit encoded video data between terminals 101, 102, 103, and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, unless otherwise stated below, the architecture and topology of network 105 may be immaterial to the operation of the present disclosure.

[0027] As an example of an application of the disclosed subject matter, Figure 2 the placement of video encoders and decoders in a streaming environment is shown. The disclosed subject matter can be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CD (Compact Disc), DVD (Digital Versatile Disc), memory sticks, etc.

[0028] A streaming system can include a capture subsystem 203, which can include a video source 201, such as a digital camera device, that creates, for example, an uncompressed video sample stream 213. The sample stream 213 can be emphasized as having a high data volume when compared with an encoded video bitstream and can be processed by an encoder 202 coupled to the video source 201, which can be, for example, a camera device as discussed above. Encoder 202 can include hardware, software, or a combination thereof to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video bitstream 204 can be emphasized as having a lower data volume when compared with the sample stream, and the encoded video bitstream 204 can be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 can access the streaming server 205 to retrieve copies 208 and 206 of the encoded video bitstream 204. Client 212 can include a video decoder 211 that decodes an incoming copy of the encoded video bitstream 208 and creates an outgoing video sample stream 210 that can be rendered on a display 209 or other rendering device (not depicted). In some streaming systems, the video bitstreams 204, 206, and 208 can be encoded according to certain video coding / compression standards. Examples of these standards are mentioned above and further described herein.

[0029] According to the exemplary embodiments described further below, the term "mesh" indicates the composition of one or more polygons that describe the surface of a volumetric object. Each polygon is defined by its vertices in 3D space and information on how the vertices are connected (referred to as connectivity information). Optionally, vertex attributes such as color, normal, etc. can be associated with the mesh vertices. The attributes can also be associated with the surface of the mesh by using mapping information that parameterizes the mesh with a 2D attribute map. Such a mapping can be described by a set of parameter coordinates, which are referred to as UV coordinates or texture coordinates and are associated with the mesh vertices. The 2D attribute map is used to store high-resolution attribute information such as texture, normal, displacement, etc. According to the exemplary embodiments, such information can be used for various purposes such as texture mapping and shading.

[0030] However, a dynamic mesh sequence may require a large amount of data because a dynamic mesh sequence may include a large amount of information that changes over time. For example, in contrast to a "static mesh" or a "static mesh sequence" where the information of the mesh cannot change from one frame to another, a "dynamic mesh" or a "dynamic mesh sequence" indicates the movement of which of the vertices represented by the mesh change from one frame to another. Therefore, efficient compression techniques are needed to store and transmit such content. Mesh compression standards IC, MESHGRID, FAMC were previously developed by MPEG (Moving Picture Experts Group, MPEG) to handle dynamic meshes with constant connectivity and time-varying geometry and vertex attributes. However, these standards do not consider time-varying attribute maps and connectivity information. Digital Content Creation (DCC) tools typically generate such dynamic meshes. Correspondingly, for volumetric acquisition techniques, generating a dynamic mesh with constant connectivity, especially under real-time constraints, is challenging. Existing standards do not support this type of content. According to the exemplary embodiments herein, aspects of a new mesh compression standard for directly handling dynamic meshes with time-varying connectivity information and optionally time-varying attribute maps are described, which is for lossy and lossless compression for various applications such as real-time communication, storage, free viewpoint video, AR (Augmented Reality, AR), and VR (Virtual Reality, VR). Functions such as random access and scalable / progressive coding are also considered.

[0031] Figure 3Shows an example framework 300 for dynamic mesh compression for methods based on 2D atlas sampling, for example. Each frame of the input mesh 301 can be preprocessed through a series of operations (e.g., tracing, remeshing, parameterization, voxelization). Note that these operations can be encoder-only, meaning that these operations may not be part of the decoding process, and this possibility can be signaled in the metadata through the following flags: for example, for encoder-only, the flag indicates 0 and for others, the flag indicates 1. Thereafter, a mesh 302 with a 2D UV atlas can be obtained, where each vertex of the mesh has one or more associated UV coordinates on the 2D atlas. Then, by sampling the 2D atlas, the mesh can be converted into multiple graphs, including a geometry graph and an attribute graph. Then, these 2D graphs can be encoded by a video / image codec such as HEVC (High-Efficiency Video Coding, HEVC), VVC (Versatile Video Coding, VVC), AV1 (AOMedia Video 1, AV1), AVS3 (Audio Video Coding Standard 3, AVS3), etc. On the decoder 303 side, the mesh can be reconstructed based on the decoded 2D graphs. Any post-processing and filtering can also be applied to the reconstructed mesh 304. Note that for the purpose of 3D mesh reconstruction, other metadata can be signaled to the decoder side. Note that the chart boundary information including uv coordinates and xyz coordinates of the boundary vertices can be predicted, quantized, and entropy-coded in the bitstream. The quantization step size can be configured on the encoder side to trade off quality and bitrate.

[0032] In some implementations, a 3D mesh can be divided into several segments (or patches / charts), and according to an exemplary embodiment, one or more 3D mesh segments can be regarded as a "3D mesh". Each segment consists of a set of connected vertices associated with its geometric, attribute, and connectivity information.

[0033] Dual-degree mesh encoding and decoding is a specialized technique aimed at efficiently encoding and decoding the connectivity of polygon meshes. Through the principle of duality, dual-degree mesh encoding and decoding can encode and decode the connectivity data of a sub-mesh by constructing the following two separate sequences: one characterizing the valence of vertices and the other depicting the degree of faces.

[0034] Figure 4A Is a graph 400 showing a dual-degree traversal according to an embodiment of the present disclosure.

[0035] As shown in FIG. 400, the encoding process involves simultaneous traversal around both faces and vertices. Specifically, the traversal starts from an arbitrary seed face, and from that seed face, the degree of the face is recorded. Subsequently, the valence of adjacent vertices is also marked as a VERTEX symbol with valence 5, such as V5. Then, the so-called "pivot" vertex is identified by having the minimum degree of freedom, where the degree of freedom represents the count of the non-traversed adjacent faces of that vertex. The traversal continues around that pivot vertex, appending new faces and vertices and recording the degrees of these faces and vertices respectively. The performance of the dual-degree encoding depends largely on the regularity of the face degrees and vertex valences.

[0036] The attributes of the mesh include vertex positions, texture coordinates, normal vectors, and associated texture maps. The geometric attributes such as vertex positions are encoded following the traversal order and implementing a predictive coding scheme. That is, the residual vector between the current position and the predicted position is encoded into the bitstream as follows:

[0037] r = v - p... Equation 1

[0038] Figure 4B FIG. 450 is a diagram of the processing of predicting the first three positions of a mesh according to an embodiment of the present disclosure.

[0039] As shown in FIG. 450, parallelogram prediction, which is typically associated with quadrilateral meshes, is initially usually difficult, especially for the first three vertices, where these advanced predictions are not available because there are not enough reference vertices (i.e., at least 3 previously encoded vertices).

[0040] For the first position, since there is no reference position, the initial predictor can be the center of the input mesh. In an embodiment, for a quantized mesh, without knowing the bounding box, the initial predictor will instead use P0 = (2 QP-1 , 2 QP-1 , 2 QP-1 ). QP represents the bit depth of the position attribute.

[0041] For the second and third positions, the previously reconstructed positions are used as predictors. Compared with the parallelogram predictor, the initial and final predictors are generally the worst and result in large residual vectors.

[0042] In practice, a mesh typically includes multiple connected components such as sub-meshes, which result from modeling a complex mesh from multiple simpler meshes. This modular design is common in both computer-generated meshes and 3D scanned meshes. Unfortunately, position encoding generally follows the traversal order, which poses challenges when encoding the initial vertices of each sub-mesh.

[0043] Accordingly, embodiments of the present disclosure enable efficient processing of predictions for the first three vertices for each connected component. The methods disclosed herein can be applied to any positional attribute compression algorithm, regardless of the polygon structure. By introducing new vertex symbols, the methods disclosed herein are also applicable to all traversal methods.

[0044] The methods disclosed herein enable selection of the best initial vertex for a given connected component and then optimization of the prediction order for a mesh with multiple connected components. Additionally, the methods disclosed herein include methods for signaling the best prediction mode for the first three vertices of each sub-mesh and optimizing the encoding order of the sub-mesh to further reduce the bitrate.

[0045] Given two non-connected sub-meshes Mi and Mj, where Mi has been encoded and the object of the present disclosure is to find a first position in Mi to encode Mj with the best predictor for that first position. In an embodiment, the minimum cost is the distance between the two closest vertices belonging to each sub-mesh. Assume that the vertex pair with the minimum distortion is and Then, the strategy according to the embodiment for position encoding would be to select as the initial predictor and select as the first vertex to be encoded.

[0046] Multiple methods can be used to signal the initial predictor. In one embodiment, signaling can be performed via the encoding order in sub-mesh Mi. In the same or another embodiment, the selection of the best predictor vertex can be restricted to an offset n from the last encoded vertex, thus minimizing bit signaling. However, for sub-meshes with a large number of vertices, signaling the index may not be efficient. While truncating the offset can help reduce the bit overhead, it may also result in a sub-optimal predictor.

[0047] According to an embodiment of the present disclosure, another signaling method can be to introduce a new vertex symbol in the dual-degree coding. A new vertex PRED (but not limited to this) is introduced adjacent to the vertex to mark it as the predictor for the next sub-mesh. In the worst case, this symbol may occur once for each vertex. Encoding this PRED symbol can be based on an offset that is based on the closest distance to the current pivot. This distance can be the L1 norm distance or the L2 norm distance. Since the number of vertices added for each pivot is relatively small, this significantly reduces the offset value for signaling the last vertex.

[0048] In an embodiment, a PRED symbol may be inserted after a bidirectional traversal only when the cost with the PRED symbol is significantly better than using the last encoded vertex, thereby significantly reducing the number of PRED symbols in the bitstream. On the decoder side, on the other hand, the last encoded vertex from the previous submesh may be used by default, and only the vertex with the PRED symbol is used as its predictor (if available).

[0049] Methods in the related art handle two submeshes in a greedy manner. However, if there are a large number of submeshes, the problem of the optimal submesh coding order becomes highly relevant. In an embodiment, an input mesh may be divided into multiple submeshes, an initial prediction cost between each pair of submeshes may be calculated, and a bidirectional graph may be formed.

[0050] Figure 5A FIG. 500 is an exemplary diagram showing a bidirectional graph of a submesh coding order according to an embodiment of the present disclosure.

[0051] Each direction in the figure indicates the cost of encoding the initial vertex and the cost of signaling the PRED symbol. This cost is denoted as Cost(Mi, Mj), which is different in direction Cost(Mi, Mj)≠Cost(Mj, Mi), thus highlighting the importance of the optimal coding order. In one embodiment, the optimized coding order may be determined as a Hamiltonian Path Problem (HPP) of a bidirectional complete graph as shown in FIG. 500. The numbers in FIG. 500 are used to show the cost of using the starting submesh as the initial predictor of the target submesh.

[0052] In another embodiment, the optimized coding order may be determined as a Hamiltonian Path Problem (HPP) of an incomplete bidirectional diagonal graph, where the starting node is the first initial predictor having only a forward direction. This method may be more complex, but provides higher performance due to considering the first initial predictor S as shown Figure 5B as shown.

[0053] Cost(Mi,Mj) is the gap between the cost of encoding the first three vertices in Mi using the nearest vertex in Mj and the cost of signaling the PRED symbol. However, calculating this cost is computationally complex because it involves comparing many vertices. Therefore, several simplified cost estimation methods are introduced in the present disclosure.

[0054] According to the present disclosure, the input mesh may be simplified before performing a search for the nearest minimum pair to reduce complexity. Then, once the above order is completed, a refinement search may be performed to find the minimum distance pair from the two selected submeshes.

[0055] In an embodiment, the cost is defined as the distance between the centroids of two selected sub - grids. For example:

[0056] Cost(M i , M j ) = |C i - C j |... Equation 2

[0057] In this embodiment, the dynamic graph becomes an "undirected weighted graph", which accelerates the optimization algorithm.

[0058] In another embodiment, instead of using the centroid distance, an axis - aligned bounding box is determined, and then the minimum distance between the two bounding boxes is found. If the bounding boxes overlap, the distance is set to zero. Otherwise, the distance is the minimum distance between the surfaces of the bounding boxes of the two sub - grids.

[0059] In another embodiment, the bounding box difference can be enhanced by incorporating the orientation of the vertices. This can result in a more accurate approximation of the rate estimate.

[0060] In the above - mentioned embodiment, the selection of the initial vertex for encoding or decoding can be based on the minimum cost. However, in an embodiment, the initial predictor, which is a subsampled position of the bounding box, can be determined from a list of predictor candidates. The predictor can be a predetermined subsampled position within the bounding box as Figure 6 shown. Depending on the amount of bits b to be signaled, only predictors of order less than 2 b - 1 can be used. In one embodiment, the candidates in the list are given in order of priority. An example priority order can include the order in Table 1.

[0061] Table 1: List of candidates for the initial predictor of the first sub - grid in order of priority

[0062] Sequence Point 1 <![CDATA[(2 QP_1 ,2 QP_1 ,2 QP_1 )]]> 2 (0,0,0) 3 <![CDATA[(2 QP_2 , 2 QP_2 , 2 QP_2 )]]> 4 <![CDATA[3*(2 QP_2 ,2 QP_2 ,2 QP_2 )]]> 5 …

[0063] For each predictor candidate, the best vertex position V with the minimum cost for the selected predictor v is found from the set of uncoded vertices Vu as follows: p r

[0064]

[0065] r i = v i - v p ... Equation 4

[0066] In an embodiment, the Sum of Absolute Different (SAD) can be used as a cost function. In another embodiment, rate estimation can be used. In this case, the search space is the number n of vertices v .

[0067] In one embodiment, it also includes the cost of encoding the second and third vertices as well. That is, for each vertex Vi, the initial cost of encoding it is calculated as follows:

[0068]

[0069] Then, for each neighbor face f of v i , calculate the cost of encoding j and in the counterclockwise order of the face:

[0070]

[0071] where is the reconstructed version of v in the case of lossy encoding of the vertex. It can be written as such a formula:

[0072]

[0073] Then, record the best prediction mode of the initial predictor.

[0074] Starting from the second sub-grid, there are additional predictors to choose from, including the centroid of the previously encoded sub-grid (referred to as Vcentroid), the first (Vfirst) and last (Vlast) encoded vertices of the previously encoded sub-grid, the central vertex (vcenter), and the previously encoded last vertex, the previously k-th last encoded vertex (Vlast-k; k = 1, 2). In an embodiment, the candidate list can include the candidates defined in Table 2.

[0075] Table 2: Candidate list of initial predictors for the remaining sub-grid in priority order

[0076] Sequence Point l <![CDATA[V last > 2 <![CDATA[V first > 3 <![CDATA[V entroid > 4 <![CDATA[V center > 5 <![CDATA[V ast-1 > … n <![CDATA[V last-k >

[0077] ​According to an embodiment, the mechanism is used to signal the prediction mode of the initial vertex prediction. It includes signaling the number of sub-grids and the prediction mode of each sub-grid: b(.) represents boolean, i(.) represents signed, u(.) represents unsigned bit, and S represents the number of sub-grids. In an embodiment, b1 bits can be used to signal the number of sub-grids, b2 bits can be used to signal the prediction mode of the first sub-grid, and b3 can be used to signal the remaining sub-grids. In an embodiment, b2 can be greater than b3 because b3 is a poorer predictor.

[0078] Table 3: Signaling the Initial Prediction Mode

[0079]

[0080]

[0081] According to an embodiment, only one reference determined by both the encoder and the decoder can be used,

[0082] such that no signaling is required.

[0083] Figure 7 A is an exemplary process 700 for encoding a plurality of grids using predictor vertices belonging to an already encoded grid according to an embodiment of the present disclosure.

[0084] At operation 705, at least two sub-grids can be received. In an embodiment, the first sub-grid among the at least two sub-grids can already be encoded, and the second sub-grid among the at least two sub-grids is to be encoded.

[0085] At operation 710, a vertex pair with the minimum prediction cost can be determined. The first vertex in the vertex pair can belong to the first sub-grid, and the second vertex in the vertex pair can belong to the second sub-grid. In an embodiment, the minimum prediction cost includes selecting the first vertex from a candidate list, where the candidate list includes predetermined sub-sampling positions having a bounding box associated with the first sub-grid.

[0086] At operation 715, the first vertex in the vertex pair can be set as the predictor vertex of the second sub-grid to be encoded.

[0087] At operation 720, the first vertex in the vertex pair can be signaled to the bi-degree decoder, and then the second sub-grid can be encoded in a first encoding order, where the second vertex in the vertex pair is encoded first in the second sub-grid.

[0088] In an embodiment, signaling may include: signaling a first vertex via an encoding order of the first vertex in a first sub-grid; or signaling the first vertex as an offset from a last encoded vertex. In an embodiment, a new predictor vertex may be set adjacent to the first vertex; and the new predictor vertex may be signaled together with the first vertex in the vertex pair.

[0089] In an embodiment, the new predictor vertex may be signaled as an offset based on a distance to a current pivot of the first sub-grid or the second sub-grid.

[0090] In an embodiment, a cost function based on a prediction cost between grid pairs and a cost of signaling the sign of the new predictor vertex may be used to select a first sub-grid among at least two sub-grids to be encoded. The cost function may be based on a centroid distance between respective grid pairs and / or a minimum distance between respective bounding boxes associated with respective grid pairs.

[0091] In an embodiment, the processing 700 may further include determining remaining predictor vertices of a remaining grid among at least two sub-grids. The remaining predictor vertices may be determined from a predetermined list, where the predetermined list includes a first encoded vertex of a previously encoded sub-grid, a last encoded vertex of the previously encoded sub-grid, a centroid of the previously encoded sub-grid, and a center vertex of the previously encoded sub-grid. The remaining grid may be encoded based on the remaining predictor vertices.

[0092] In an embodiment, the processing 700 may further include signaling the number of grids among at least two sub-grids using a first number of bits; signaling a first prediction mode of the first sub-grid using a second number of bits; and signaling a remaining prediction mode of the remaining grid using a third number of bits, where the third number of bits is less than the second number of bits.

[0093] It will be appreciated that the processing 700 may describe an encoding process, but those skilled in the art will know that for a decoding process, similar operations may be performed in a modified order.

[0094] The proposed methods may be used alone or in any order combination. The proposed methods may be used for any polygon mesh, but even so, only triangular meshes may be used for the demonstration of various embodiments. As described above, it will be assumed that the input mesh may contain one or more instances, a sub-grid is a part of the input mesh having one or more instances, and multiple instances may be grouped to form a sub-grid.

[0095] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media, or implemented by one or more specially configured hardware processors. For example, Figure 8 FIG. shows a computer system 800 suitable for implementing certain embodiments of the disclosed subject matter.

[0096] The computer software may be encoded using any suitable machine code or computer language, which may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.

[0097] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0098] Figure 8 The components shown in FIG. for the computer system 800 are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system 800.

[0099] The computer system 800 may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface device may also be used to capture certain media that is not necessarily directly related to conscious input made by a human, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera device), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0100] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard 801, mouse 802, touchpad 803, touch screen 810, joystick 805, microphone 806, scanner 808, camera device 807.

[0101] The computer system 800 may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen 810 or the joystick 805, but there may also be tactile feedback devices that do not function as input devices); audio output devices (e.g., the speaker 809, headphones (not depicted)); visual output devices (e.g., the screen 810, including a CRT (Cathode Ray Tube) screen, an LCD (Liquid Crystal Display) screen, a plasma screen, an OLED (Organic Light Emitting Diode) screen, each screen having or not having touch screen input capabilities, each screen having or not having tactile feedback capabilities - some of the screens may be able to output two-dimensional visual output or more than three-dimensional output in ways such as stereoscopic graphics output; virtual reality glasses (not depicted); holographic displays and smoke cans (not depicted)); and printers (not depicted).

[0102] The computer system 800 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory) / RW 820 with media such as CD / DVD 811, thumb drives 822, removable hard disk drives or solid state drives 823, traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC (Application Specific Integrated Circuit) / PLD (Programmable Logic Device) such as security dongles (not depicted), etc.

[0103] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0104] The computer system 800 may also include an interface 899 to one or more communication networks 898. The network 898 may be, for example, wireless, wired, or optical. The network 898 may also be local area, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, etc. Examples of the network 898 include: local area networks such as Ethernet, wireless LAN (Local Area Network, LAN); cellular networks including GSM (Global System for Mobile Communications, GSM), 3G (the Third Generation, 3G), 4G (the Fourth Generation, 4G), 5G (the Fifth Generation, 5G), LTE (Long Term Evolution, LTE), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CAN Bus (Controller Area Network Bus, CAN Bus), etc. Certain networks 898 typically require an external network interface adapter attached to certain common data ports or peripheral buses (750 and 851) (such as, for example, the USB (Universal Serial Bus, USB) port of the computer system 800); other networks are typically integrated into the core of the computer system 800 by attaching to a system bus as described below (e.g., an Ethernet interface in a PC computer system or a cellular network interface in a smartphone computer system). Using any of these networks 898, the computer system 800 can communicate with other entities. Such communication can be one-way reception only (e.g., broadcast TV), one-way transmission only (e.g., CAN Bus to certain CANbus devices), or two-way, such as to other computer systems using local area digital networks or wide area digital networks. Certain protocols and protocol stacks can be used on each of these networks and network interfaces as described above.

[0105] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core 840 of the computer system 800.

[0106] The core 840 may include one or more central processing units (CPUs) 841, a graphics processing unit (GPU) 842, a graphics adapter 817, a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) 843, a hardware accelerator 844 for certain tasks, etc. These devices, together with a read-only memory (ROM) 845, a random access memory 846, and an internal mass storage device 847 such as an internal hard disk drive, SSD (Solid State Drive, SSD) that is not accessible to users, can be connected via a system bus 848. In some computer systems, the system bus 848 can be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus 849 to the system bus 848 of the core. The architecture of the peripheral bus includes PCI (Peripheral Component Interconnect / Interface, PCI), USB, etc.

[0107] The CPU 841, GPU 842, FPGA 843, and accelerator 844 can execute certain instructions, which, when combined, can constitute the computer code mentioned above. This computer code can be stored in the ROM 845 or the RAM (Random Access Memory, RAM) 846. Transient data can also be stored in the RAM 846, while permanent data can be stored in, for example, the internal mass storage device 847. Fast storage and retrieval of any memory device in the memory devices can be achieved by using a cache memory that can be closely associated with one or more CPUs 841, GPUs 842, mass storage device 847, ROM 845, RAM 846, etc.

[0108] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specifically designed and constructed for the purposes of this disclosure, or the medium and the computer code can be of the types known and available to those skilled in the art of computer software.

[0109] By way of example and not limitation, a computer system having architecture 800 and in particular core 840 can provide functionality due to software executed by a processor (including a CPU, GPU, FPGA, accelerator, etc.) included in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of core 840 having a non-transitory nature such as mass storage device 847 internal to the core or ROM 845. Software implementing various embodiments of the present disclosure can be stored in such devices and executed by core 840. Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause core 840 and in particular the processors therein (including the CPU, GPU, FPGA, etc.) to perform specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM 846 and modifying such data structures in accordance with processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to being logically hardwired or otherwise embodied in circuitry (e.g., accelerator 844) that can operate in place of or in conjunction with the software to perform specific processes or specific portions of specific processes described herein. In appropriate instances, references to software can include logic, and references to logic can also include software. In appropriate instances, references to computer-readable media can include circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry implementing logic for execution, or both of the above. The present disclosure encompasses any suitable combination of hardware and software.

[0110] Although the present disclosure has described several exemplary embodiments, there are variations, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to envision many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for encoding and decoding grid attributes using a bi-degree grid encoder, the method being executed by at least one processor, the method comprising: Receiving at least two sub-grids, wherein a first sub-grid of the at least two sub-grids was previously encoded, and wherein a second sub-grid of the at least two sub-grids is to be encoded; Determining a vertex pair having a minimum prediction cost, wherein a first vertex of the vertex pair is part of the first sub-grid, and wherein a second vertex of the vertex pair is part of the second sub-grid; and Setting the first vertex of the vertex pair as a predictor vertex corresponding to the second sub-grid to be encoded; Signaling the first vertex of the vertex pair to a bi-degree grid decoder; and Encoding the second sub-grid by first encoding the second vertex of the vertex pair in the second sub-grid.

2. The method according to claim 1, wherein, The signaling comprises: Signaling the first vertex via the encoding order of the first vertex in the first sub-grid; or Signaling the first vertex as an offset from the last encoded vertex.

3. The method according to claim 1, wherein, The signaling comprises: Setting a new predictor vertex adjacent to the first vertex; and Signaling the new predictor vertex together with the first vertex of the vertex pair.

4. The method according to claim 3, wherein Signaling the new predictor vertex comprises: Signaling an offset based on the distance to the current pivot of the first sub-grid or the second sub-grid.

5. The method according to claim 1, wherein, The method further comprises: Selecting the first sub-grid of the at least two sub-grids for encoding using a cost function based on the prediction cost between grid pairs and the cost of signaling the symbol of the new predictor vertex.

6. The method according to claim 5, wherein, The prediction cost between the grid pairs is based on one of the following: The centroid distance between individual grid pairs; and The minimum distance between individual bounding boxes associated with the individual grid pairs.

7. The method according to claim 1, wherein The minimum prediction cost includes selecting the first vertex from a candidate list, wherein the candidate list includes predetermined subsampling positions having a bounding box associated with the first sub-grid.

8. The method according to claim 1, wherein The method further comprises: Determining remaining predictor vertices of the remaining grid of the at least two sub-grids, wherein the remaining predictor vertices are determined from a predetermined list, wherein the predetermined list includes a first encoded vertex of a previously encoded sub-grid, a last encoded vertex of the previously encoded sub-grid, a centroid of the previously encoded sub-grid, and a center vertex of the previously encoded sub-grid; and Encoding the remaining grid based on the remaining predictor vertices.

9. According to the method of claim 8, wherein The method further comprises: Signaling the number of grids in the at least two sub-grids using a first number of bits; Signaling a first prediction mode of the first sub-grid using a second number of bits; and Signaling a remaining prediction mode of the remaining grid using a third number of bits, wherein the third number of bits is less than the second number of bits.

10. An apparatus for encoding and decoding grid attributes using a bi-degree grid decoder, the apparatus comprising: At least one memory configured to store program code; And At least one processor configured to read the program code and operate in accordance with the program code, the program code including: First receiving code configured to cause the at least one processor to receive at least two encoded sub-grids; Second receiving code configured to cause the at least one processor to receive predictor vertices from a bi-degree grid encoder, wherein the predictor vertices are the first vertices in vertex pairs having a minimum prediction cost; First determination code configured to cause the at least one processor to determine a second vertex based on the predictor vertices, wherein the first vertex in the vertex pair is part of a first sub-grid of the at least two encoded sub-grids, and wherein the second vertex in the vertex pair is part of a second sub-grid of the at least two encoded sub-grids; and First decoding code configured to cause the at least one processor to decode the second sub-grid by first decoding the second vertex in the vertex pair in the second sub-grid.

11. The apparatus according to claim 10, wherein, The predictor vertices are received as one of the following: The encoding order of the predictor vertices in the first sub-grid; or The offset of the last encoded vertex.

12. The apparatus according to claim 10, wherein, New predictor vertices are received from the bi-degree grid encoder together with the predictor vertices.

13. The apparatus according to claim 12, wherein, The new predictor vertices are received as an offset based on the distance to the current pivot of the first sub-grid or the second sub-grid.

14. The apparatus according to claim 10, wherein, The program code further includes: Third receiving code configured to cause the at least one processor to receive remaining predictor vertices of a remaining grid of the at least two encoded sub-grids, wherein the remaining predictor vertices are determined from a predetermined list including a first encoded vertex of a previously encoded sub-grid, a last encoded vertex of the previously encoded sub-grid, a centroid of the previously encoded sub-grid, and a central vertex of the previously encoded sub-grid; and Second decoding code configured to cause the at least one processor to decode the remaining grid based on the remaining predictor vertices.

15. The apparatus according to claim 14, wherein, The program code further includes: Fourth receiving code configured to cause the at least one processor to receive the number of grids in the at least two encoded sub-grids using a first number of bits; Fifth third receiving code configured to cause the at least one processor to receive a first prediction mode of the first sub-grid using a second number of bits; and Sixth receiving code configured to cause the at least one processor to receive a remaining prediction mode of the remaining grid using a third number of bits, wherein the third number of bits is less than the second number of bits.

16. A non-transitory computer-readable medium storing instructions, the instructions including one or more instructions that, when executed by one or more processors of a device for grid property encoding and decoding, cause the one or more processors to perform the following operations: Perform a conversion between a visual media file and a bitstream of visual media data according to format rules, Among them, The bitstream including at least two sub-grids, a first sub-grid of the at least two sub-grids having been previously converted, Wherein the bitstream includes a predictor vertex for a second sub-grid to be converted, wherein the predictor vertex is the first vertex in a pair of vertices having a minimum prediction cost, wherein the first vertex in the pair of vertices is part of the first sub-grid of the at least two sub-grids, and wherein the second vertex in the pair of vertices is part of the second sub-grid of the at least two sub-grids; and Encode or decode the second sub-grid using the second vertex in the pair of vertices.

17. The non-transitory computer-readable medium according to claim 16, wherein, The bitstream includes the predictor vertex as one of the following: The encoding order of the predictor vertex in the first sub-grid; or The offset of the last encoded vertex.

18. The non-transitory computer-readable medium according to claim 16, wherein, The bitstream includes a new predictor vertex together with the predictor vertex.

19. The non-transitory computer-readable medium according to claim 18, wherein, The new predictor vertex is an offset based on the distance to the current pivot of the first sub-grid or the second sub-grid.

20. The non-transitory computer-readable medium according to claim 16, wherein, The bitstream further includes remaining predictor vertices of remaining grids of the at least two sub-grids, wherein the remaining predictor vertices are determined from a predetermined list including a first encoded vertex of a previously converted sub-grid, a last encoded vertex of the previously converted sub-grid, a centroid of the previously converted sub-grid, and a central vertex of the previously converted sub-grid; and Wherein the encoding or decoding further includes encoding or decoding the remaining grids based on the remaining predictor vertices.