Improvement of Coding of Boundary UV2XYZ Index for Mesh Compression
The method addresses inefficiencies in existing mesh compression by encoding UV2XYZ indices to efficiently compress and reconstruct dynamic meshes with time-varying connectivity, supporting real-time applications like AR and VR.
Patent Information
- Application Number
- JP2024527229
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-03-30
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2043-03-30
AI Technical Summary
Existing mesh compression standards fail to efficiently handle dynamic meshes with time-varying connectivity information and attribute maps, particularly under real-time constraints, and do not support volume acquisition techniques for continuously connected meshes.
A method for coding UV2XYZ indices of boundary vertices using a tuple format that includes parameters for the start index, length, and direction of runs in the UV2XYZ array, enabling efficient compression and reconstruction of 3D meshes from 2D meshes.
Enables efficient compression and reconstruction of dynamic meshes, supporting real-time applications such as AR and VR by reducing data volume and maintaining high-quality 3D mesh representation.
Smart Images

Figure 0007715943000003 
Figure 0007715943000004 
Figure 0007715943000005
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 331,699, filed Apr. 15, 2022, and U.S. Patent Application No. 18 / 191,457, filed Mar. 28, 2023, the disclosures of which are hereby incorporated by reference in their entireties.
[0002] This disclosure is directed to a set of advanced video coding techniques. More specifically, this disclosure is directed to video - based mesh compression, including a method for coding UV2XYZ indices of boundary vertices for efficient mesh compression.
Background Art
[0003] The world's advanced three - dimensional (3D) representations enable more immersive interactions and communications. To achieve the sense of presence of 3D representations, 3D models have become more sophisticated than ever, and a significant amount of data is associated with the creation and consumption of these 3D models. 3D meshes are widely used in 3D model immersive content.
[0004] A 3D mesh can be composed of several polygons that describe the surface of a volumetric object. A dynamic mesh sequence can require a large amount of data because it can have a significant amount of information that changes over time. Therefore, efficient compression techniques are needed to store and transmit such content.
[0005] Mesh compression standards such as IC, MESHGRID, and FAMC were previously developed to handle dynamic meshes with constant connectivity and time - varying geometry and vertex attributes. However, these standards do not take into account time - varying attribute maps and connectivity information.
[0006] Furthermore, especially under real-time constraints, it is also difficult for volume acquisition techniques to generate a continuously connected dynamic mesh. This type of dynamic mesh content is not supported by existing standards.
SUMMARY OF THE INVENTION
MEANS FOR SOLVING THE PROBLEM
[0007] According to one or more embodiments, a method implemented by at least one processor in a decoder includes receiving a coded video bitstream that includes (i) one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh and (ii) a 2D-3D index array that maps each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh. The method further includes reconstructing the 3D mesh using the 2D-3D index array and mapping each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh. The 2D-3D index array is encoded in a tuple format that includes a first parameter that specifies a start index of a run of consecutive integers for each tuple in the 2D-3D index array, a second parameter that specifies a length of the run, and a third parameter that specifies a direction of the run.
[0008] According to one or more embodiments, a decoder comprises at least one memory configured to store program code and at least one processor configured to read the program code and operate as instructed by the program code. The program code includes receive code configured to cause the at least one processor to receive a coded video bitstream including (i) one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh and (ii) a 2D-3D index array mapping each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh. The program code includes reconstruction code configured to cause the at least one processor to reconstruct the 3D mesh using the 2D-3D index array and map each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh. The 2D-3D index array is encoded in a tuple format including a first parameter specifying a start index of a run for each tuple in the 2D-3D index array, a second parameter specifying a length of a run of consecutive integers, and a third parameter specifying a direction of the run.
[0009] According to one or more embodiments, a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor in a decoder, cause the processor to: receive a coded video bitstream including one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh, and (ii) a 2D-3D index array that maps each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh; and reconstruct the 3D mesh using the 2D-3D index array, mapping each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh. The 2D-3D index array is encoded in a tuple format, where each tuple in the 2D-3D index array includes a first parameter that specifies a starting index of a run of consecutive integers, a second parameter that specifies a length of the run, and a third parameter that specifies a direction of the run.
[0010] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
DETAILED DESCRIPTION OF THE INVENTION
[0012] The following detailed description of the exemplary embodiments refers to the accompanying drawings. The same reference numbers in different drawings may identify the same or similar elements.
[0013] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the disclosed embodiments to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure, or may be obtained from the practice of the embodiments. Additionally, one or more features or components of one embodiment may be incorporated into or combined with another embodiment (or one or more features of another embodiment). In addition, in the flowcharts and descriptions of operations provided below, one or more operations may be omitted, one or more operations may be added, one or more operations may be performed simultaneously (at least partially), and the order of one or more operations may be interchanged.
[0014] It will be apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit the implementation form. Thus, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code, and it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0015] Certain combinations of features are recited in the claims and / or disclosed herein, but these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may not be specifically recited in the claims and / or may be combined in ways not disclosed herein. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.
[0016] Elements, operations, or instructions used in this specification should not be construed as important or essential unless explicitly described as such. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more". When only one item is intended, the term "one" or similar language is used. Also, as used in this specification, terms such as "has", "have", "having", "include", "including", etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0017] Throughout this specification, references to "one embodiment", "an embodiment", or similar language mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the solution. Thus, the phrases "in one embodiment", "in an embodiment", and similar language throughout this specification may, but do not necessarily, refer to the same embodiment.
[0018] Furthermore, the described features, advantages, and characteristics of the present disclosure may be combined in any suitable manner in one or more embodiments. One of ordinary skill in the art will recognize that the present disclosure may be practiced without one or more of the specific features or advantages of a particular embodiment in light of the description herein. In other instances, additional features and advantages may be recognized in particular embodiments that are not necessarily present in all embodiments of the present disclosure.
[0019] Embodiments of the present disclosure are directed to compressing a mesh. The mesh can be composed of several polygons that describe the surface of a volume object. Information about the vertices of the mesh in 3D space and how the vertices are connected can define each polygon and can be referred to as connectivity information. Optionally, vertex attributes such as color, normal, etc. can be associated with the mesh vertices. The attributes may also be associated with the surface of the mesh by utilizing mapping information that parameterizes the mesh with a 2D attribute map. Such a mapping is called UV coordinates or texture coordinates and can be defined using a set of parametric coordinates associated with the mesh vertices. A 2D attribute map can be used to store high-resolution attribute information such as texture, normal, displacement, etc. The high-resolution attribute information can be used for various purposes such as texture mapping and shading.
[0020] As described above, a 3D mesh or dynamic mesh can consist of a significant amount of information that changes over time and thus may require a large amount of data. Existing standards do not consider time-varying attribute maps and connectivity information. Existing standards also do not support volume acquisition techniques for generating always-connected dynamic meshes, especially under real-time conditions.
[0021] Therefore, a new mesh compression standard is needed for directly handling dynamic meshes with time-varying connectivity information and optionally time-varying attribute maps. Embodiments of the present disclosure enable efficient compression techniques for storing and transmitting such dynamic meshes. Embodiments of the present disclosure enable irreversible compression and / or reversible compression for various applications such as real-time communication, storage, free viewpoint video, AR, and VR.
[0022] According to one or more embodiments of the present disclosure, a method, a system, and a non-transitory storage medium for dynamic mesh compression are provided. Embodiments of the present disclosure can also be applied to a static mesh where only one frame of the mesh or the mesh content does not change over time.
[0023] Referring to FIGS. 1 to 2, one or more embodiments of the present disclosure for implementing the encoding and decoding structures of the present disclosure are described.
[0024] FIG. 1 shows a simplified block diagram of a communication system 100 according to an embodiment of the present disclosure. The system 100 can include at least two terminals 110, 120 interconnected via a network 150. In the case of unidirectional data transmission, the first terminal 110 can encode video data that may include mesh data at a local location in order to transmit it to the other terminal 120 via the network 150. The second terminal 120 can receive the encoded video data of the other terminal from the network 150, decode the encoded data, and display the restored video data. Unidirectional data transmission can be common in media serving applications and the like.
[0025] FIG. 1 shows, for example, a second pair of terminals 130, 140 provided to support bidirectional transmission of encoded video that may occur during a video conference. In the case of bidirectional data transmission, each terminal 130, 140 can encode video data captured at a local location in order to transmit it to the other terminal via the network 150. Each terminal 130, 140 can also receive the encoded video data transmitted by the other terminal, decode the encoded data, and display the restored video data on a local display device.
[0026] In FIG. 1, terminals 110 to 140 may be, for example, servers, personal computers, and smartphones, and / or any other type of terminal. For example, the terminals (110 to 140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing devices. Network 150 represents any number of networks that transmit coded video data among terminals 110 to 140, including, for example, wired and / or wireless communication networks. The communication network 150 may exchange data over circuit-switched and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network 150 may not be important for the operation of the present disclosure, unless otherwise described herein below.
[0027] FIG. 2 shows the arrangement of a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be used in other video-related applications, including, for example, video conferencing, digital TV, storage of compressed video on digital media such as CDs, DVDs, memory sticks, etc.
[0028] As shown in FIG. 2, the streaming system 200 may include a capture subsystem 213 that includes a video source 201 and an encoder 203. The streaming system 200 may further include at least one streaming server 205 and / or at least one streaming client 206.
[0029] Video source 201 can create a stream 202 that includes, for example, a 3D mesh and metadata related to the 3D mesh. The 3D mesh can be composed of several polygons that describe the surface of a volumetric object. For example, the 3D mesh can include a plurality of vertices in 3D space where each vertex is associated with 3D coordinates (e.g., x, y, z). The video source 201 can include, for example, a 3D sensor (e.g., a depth sensor) or 3D imaging technology (e.g., a digital camera) and a computing device configured to generate a 3D mesh using data received from the 3D sensor or the 3D imaging technology. The sample stream 202 may have a high data volume compared to an encoded video bitstream and can be processed by an encoder 203 coupled to the video source 201. The encoder 203 may include hardware, software, or a combination thereof that enables or implements aspects of the disclosed subject matter, as described in more detail below. The encoder 203 may also generate an encoded video bitstream 204. The encoded video bitstream 204 may have a lower data volume compared to the uncompressed stream 202 and can be stored on a streaming server 205 for later use. One or more streaming clients 206 can access the streaming server 205 and retrieve a video bitstream 209 that can be a copy of the encoded video bitstream 204.
[0030] The streaming client 206 may include a video decoder 210 and a display 212. The video decoder 210 may decode, for example, a video bitstream 209 that is a copy of the input encoded video bitstream 204, and create an output video sample stream 211 that can be rendered on the display 212 or another rendering device (not shown). In some streaming systems, the video bitstreams 204, 209 may be encoded according to a particular video coding / compression standard.
[0031] FIG. 3 is an exemplary diagram of a framework 300 for dynamic mesh compression and mesh reconstruction using an encoder and a decoder.
[0032] As seen in FIG. 3, the framework 300 may include an encoder 301 and a decoder 351. The encoder 301 can include one or more input meshes 305, one or more meshes 310 having a UV atlas, an occupancy map 315, a geometry map 320, an attribute map 325, and metadata 330. The decoder 351 can include a decoded occupancy map 335, a decoded geometry map 340, a decoded attribute map 345, a decoded metadata 350, and a reconstructed mesh 360.
[0033] According to one or more embodiments of the present disclosure, the input mesh 305 may include one or more frames, and each of the one or more frames may be preprocessed by a series of operations and used to generate a mesh 310 having a UV atlas. As an example, the preprocessing operations may include, but are not limited to, tracking, parameterization, remeshing, voxelization, etc. In some embodiments, the preprocessing operations may be performed only on the encoder side and not on the decoder side.
[0034] The mesh 310 with a UV atlas can be a 2D mesh. The 2D mesh can be a chart of vertices each associated with coordinates (e.g., 2D coordinates) in a 2D space. Each vertex in the 2D mesh may be associated with a corresponding vertex in the 3D mesh, and the vertices in the 3D mesh are associated with coordinates in 3D space. The compressed 2D mesh can be a version of the 2D mesh with reduced information compared to the uncompressed 2D mesh. For example, the 2D mesh may be sampled at a sampling rate at which the compressed 2D mesh includes sampling points. The 2D mesh with a UV atlas can be a mesh in which each vertex of the mesh can be associated with UV coordinates on the 2D atlas. For example, the 2D atlas can be a two-dimensional plane in which each 3D coordinate in 3D space can be assigned 2D coordinates in the 2D plane. The connected 2D coordinates can be called a 2D chart or patch. The mesh 310 with a UV atlas can be processed based on sampling and converted into a plurality of maps. As an example, the UV atlas 310 can be processed based on sampling of the 2D mesh with a UV atlas and converted into an occupancy map, a geometry map, and an attribute map. The generated occupancy map 335, geometry map 340, and attribute map 345 can be encoded using an appropriate codec (e.g., HVEC, VVC, AV1, AVS3, etc.) and sent to a decoder. In some embodiments, metadata (e.g., connectivity information, etc.) can also be sent to the decoder.
[0035] In some embodiments, on the decoder side, a mesh can be reconstructed from the decoded 2D map. Post-processing and filtering can also be applied to the reconstructed mesh. In some examples, the metadata may be signaled to the decoder side for the purpose of 3D mesh reconstruction. The occupancy map can be inferred from the decoder side when the boundary vertices of each patch are signaled.
[0036] According to one aspect, decoder 351 may receive the encoded occupancy map, geometry map, and attribute map from the encoder. The decoder 351 may, in addition to the embodiments described herein, use suitable techniques and methods to decode the occupancy map, geometry map, and attribute map. In some embodiments, decoder 351 may generate a decoded occupancy map 335, a decoded geometry map 340, a decoded attribute map 345, and decoded metadata 350. Input mesh 305 may be reconstructed into a reconstructed mesh 360 based on the decoded occupancy map 335, the decoded geometry map 340, the decoded attribute map 345, and the decoded metadata 350 using one or more reconstruction filters and techniques. In some embodiments, metadata 330 may be sent directly to decoder 351, and decoder 351 may use the metadata to generate a reconstructed mesh 360 based on the decoded occupancy map 335, the decoded geometry map 340, and the decoded attribute map 345. Post-filtering techniques including, but not limited to, remeshing, parameterization, tracking, voxelization, etc. may also be applied to the reconstructed mesh 360.
[0037] According to some embodiments, the 3D mesh can be divided into several segments (or patches / charts). Each segment can be composed of a set of connected vertices related to their geometry, attributes, and connectivity information. As shown in FIG. 4, the UV parameterization process maps the mesh segment 400 onto 2D charts (402, 404) within the 2D UV atlas. 2D UV coordinates within the 2D UV atlas may be assigned to each vertex within the mesh segment. Vertices within the 2D chart (e.g., 2D mesh) can form connected components as their 3D counterparts. The geometry, attributes, and connectivity information of each vertex can also be inherited from their 3D counterparts in a similar manner.
[0038] According to some embodiments, the 3D mesh segment can also be mapped onto a plurality of separate 2D charts. When the 3D mesh segment is mapped onto separate 2D charts, the vertices within the 3D mesh segment may correspond to a plurality of vertices within the 2D UV atlas. As shown in FIG. 5, a 3D mesh segment 500 that can correspond to the 3D mesh segment 400 may be mapped onto two 2D charts (502A, 502B) in the 2D UV atlas instead of a single chart. As shown in FIG. 5, the 3D vertices v1 and v4 each have two 2D corresponding vertices v1' and v4'.
[0039] FIG. 6 shows an example of a general 2D UV atlas 600 of a 3D mesh including a plurality of charts, where each chart can include a plurality of (e.g., three or more) vertices related to their 3D geometry, attributes, and connectivity information.
[0040] Boundary vertices can be defined within a 2D UV space. As shown in FIG. 7, the filled vertices are boundary vertices because they are on the boundary edges of the connected components (patches / charts). The boundary edges can be determined by checking whether the edge appears in only one triangle. Geometry information (e.g., 3D xyz coordinates) and 2D UV coordinates can be signaled in a bitstream.
[0041] In one or more examples, when boundary vertices in a 3D mesh correspond to multiple vertices in a 2D UV atlas, as shown in FIG. 5, the mapping from 3D XYZ coordinates to 2D UV coordinates can be one-to-many. Thus, a UV-XYZ (e.g., called UV2XYZ) index can be signaled to indicate the mapping function. UV2XYZ can be a 1D array of indices corresponding to mapping each 2D UV vertex to a 3D XYZ vertex.
[0042] A dynamic mesh sequence can require a large amount of data because it can consist of a significant amount of information that changes over time. In particular, the boundary information represents a significant portion of the entire mesh. Therefore, an efficient compression technique is needed to efficiently compress the boundary information.
[0043] Embodiments of the present disclosure may be used separately or combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0044] According to one or more embodiments, a number of methods are proposed for the coding of the UV2XYZ (UV-XYZ) index array of the boundaries in mesh compression. The methods can be applied individually or in any combination. The methods may be applied to a mesh with only one frame or to a static mesh where the mesh content does not change over time.
[0045] In one or more examples, the total number of boundary vertices is N, and the number of boundary vertices with unique xyz coordinates is M (M ≤ N). UV2XYZ may be a 1D array having a length of N, and each element in the array may indicate an index for a unique XYZ coordinate. For example, for i = 0, 1, ..., N - 1, UV2XYZ[i] ∈ {0, 1, ..., M - 1}.
[0046] In one or more examples, UV2XYZ may be separated into segments where each segment contains the UV2XYZ index of the chart. According to some embodiments, there are a number of ways to code the UV2XYZ array reversibly.
[0047] According to one or more embodiments, due to the nature of vertex storage, when assigning the indices of neighboring vertices to the UV2XYZ array, the neighboring UV2XYZ values can have an index difference of only 1. For example, within a certain range, there may be a monotonic increasing or decreasing trend. Thus, the UV2XYZ array can be generally represented in the following form: (a, a + 1, a + 2, ..., a + n a )(b, b - 1, b - 2,...., b - n b )(c, c + 1, c + 2,...., c + n c ) ,... This can be separated into several subsequences. Each subsequence can include consecutive indices with steps of +1 or -1.
[0048] Thus, in one or more examples, the UV2XYZ array may be represented and coded as a run-length direction tuple: (a,n a ,1), (b,n b ,-1), (c,n c ,1),...
[0049] In one or more examples, the UV2XYZ array may be specified as follows: {100,101,102,99,98,97,96,103,104}
[0050] In the above example, this array may equivalently be written in tuple format such as (100,2,1), (99,3,-1), (103,1,1). For example, the array in tuple format can be specified as follows: {100,2,1,99,3,-1,103,1,1}
[0051] According to some embodiments, the tuple of (RUN, LEN, DIR) can include the following parameters: (1) RUN can be the starting index in the run, (2) LEN can be the length of the run minus 1, and (3) DIR is the direction of the run (1 for increasing, -1 for decreasing). A run can indicate a set of one or more consecutive integers that are increasing or decreasing.
[0052] According to some embodiments, the tuple may be coded in different ways. In another example, the RUN parameter may be coded by fixed-length coding, and the bit length of the codeword is determined by the number of unique XYZ boundary vertices M (e.g., [log2M]). In one or more examples, the LEN parameter may be coded by fixed-length coding, and the bit length of the codeword may be determined by the number of UV boundary vertices in the current chart. For example, the bit length may be determined as [log2N j and N jis the number of UV boundary vertices in the j-th chart. N j The value of may also be coded in the bitstream.
[0053] In one or more examples, the parameter LEN may be coded by fixed-length coding, and the bit length of the codeword may be determined by the remaining number of UV boundary vertices in the current chart. For example, the bit length may be [Number] determined as, where n k is the LEN of the k-th run-length direction tuple in the j-th chart, and the index i is the index of the currently coded tuple. In one or more examples, the parameter DIR may be coded by 1-bit bypass coding. In one or more examples, the parameter DIR may be coded by arithmetic coding using context.
[0054] According to some embodiments, the parameter LEN may be coded by variable-length coding such as truncated binary coding. In one or more examples, the bit length of the codeword in variable-length coding may be determined by the number of UV boundary vertices in the current chart. For example, the bit length may be [log2N j determined as, where N j is the number of UV boundary vertices in the j-th chart. In one or more examples, N j may also be coded in the bitstream.
[0055] [Number] may be determined as, n k is the LEN of the k-th run-length direction tuple in the j-th chart, and the index i is the index of the currently coded tuple.
[0056] In one or more embodiments, the parameter LEN may be coded using the length of the previous run as a predictor, and the residual may be entropy coded. In one or more embodiments, the sign of the residual may be coded using a 1-bit flag. The absolute value of the residual may be coded using exponential Golomb coding.
[0057] According to one or more embodiments, for the tuple (RUN, LEN, DIR), one or more binarized codewords may be context coded, while one or more other binarized codewords may be bypass coded without using any context.
[0058] In one or more examples, to save bits further, the encoder may rearrange the boundary vertices in the chart (e.g., by rotating the UV2XYZ indices) such that the first run-length direction tuple in the chart is the longest (e.g., for i = 1, 2,...., n0 ≧ n i ). In this case, the boundary UV and the boundary XYZ may be rearranged according to the new UV2XYZ indices.
[0059] FIG. 8 shows a process 800 for coding a 3D mesh and generating a video bitstream according to one or more embodiments. Process 800 may be performed by an encoder 301. The process may start with an operation S802 where the 3D mesh is converted into one or more 2D meshes. For example, as shown in FIG. 5, the 3D mesh 500 is converted into 2D meshes 502A and 502B via UV parameterization.
[0060] The process proceeds to operation S804, where a 2D-3D index array is generated. The 2D-3D index array may be a UV2XYZ array that maps each vertex in one or more 2D meshes to a vertex in a 3D mesh. For example, the 2D-3D index array can map each 2D coordinate in one or more 2D meshes to a unique XYZ coordinate in the 3D mesh. The 2D-3D index can be formatted according to the 3D tuple format described above.
[0061] The process proceeds to operation S806, where a coded video bitstream is generated. The coded video bitstream can include one or more 2D meshes generated in operation S802 and the 2D-3D index.
[0062] FIG. 9 shows a process 900 for decoding a coded video bitstream according to one or more embodiments. The process 900 may be performed by a decoder 351. The process can start at operation S900 where a coded video bitstream is received. The coded video bitstream can include one or more 2D meshes corresponding to a 3D mesh. The coded video bitstream can further include a 2D-3D index array such as a UV2XYZ array. The process proceeds to operation S904, where the 3D mesh is reconstructed using the 2D-3D index array and one or more 2D meshes. For example, referring to FIG. 5, a 3D mesh segment can be reconstructed using 2D mesh segments 502A and 502B and the 2D-3D index array.
[0063] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, FIG. 10 shows a computer system 1000 suitable for implementing particular embodiments of the present disclosure.
[0064] Computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, and linking to create code containing instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or via interpretation, microcode execution, etc.
[0065] Instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, gaming machines, Internet of Things devices, etc.
[0066] The components shown in FIG. 10 for computer system 1000 are examples and are not intended to imply any limitation regarding the use or functionality scope of the computer software implementing the embodiments of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one or combination of the components exemplified in the non-limiting embodiments of computer system 1000.
[0067] Computer system 1000 may include specific human interface input devices. Such human interface input devices may respond to input by one or more human users via, for example, tactile input (keystrokes, swipes, movements of a data glove, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), olfactory input (not shown). The human interface device may also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (voice, music, ambient sound, etc.), images (scanned images, photographic images, obtained from a still image camera, etc.), video (2D video, 3D video including stereoscopic video, etc.).
[0068] The input human interface device may include one or more (only one of each is illustrated) of a keyboard 1001, a mouse 1002, a trackpad 1003, a touch screen 1010, a data glove, a joystick 1005, a microphone 1006, a scanner 1007, and a camera 1008.
[0069] The computer system 1000 may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, via tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen 1010, a data glove, or a joystick 1005, although there may be a tactile feedback device that does not function as an input device). For example, such devices may include audio output devices (such as a speaker 1009, headphones (not shown)), visual output devices (each with or without touch screen input capability, each with or without tactile feedback capability, some of which may be capable of outputting more than three dimensions, such as two-dimensional visual output or three-dimensional output by means such as stereoscopic output, including a screen 1010 such as a CRT screen, an LCD screen, a plasma screen, an OLED screen, virtual reality glasses (not shown), a holographic display, and a smoke tank (not shown)), and a printer (not shown).
[0070] The computer system 1000 may also include human-accessible storage devices and their associated media, such as an optical medium including a CD / DVD ROM / RW 1020 having a CD / DVD or similar medium 1021, a thumb drive 1022, a removable hard drive or solid state drive 1023, legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).
[0071] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include a transmission medium, a carrier wave, or other transient signals.
[0072] Computer system 1000 may also include an interface to one or more communication networks. The network may be wireless, wired, or optical. The network may further be local, wide area, metropolitan, vehicular and industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicular and industrial including CANBus. A particular network generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus 1049 (such as a USB port of computer system 1000), and other networks are generally integrated into the core of computer system 1000 by attachment to the system bus as described below (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, computer system 1000 can communicate with other entities. Such communication may be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANBus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. Such communication may also include communication to cloud computing environment 1055. Specific protocols and protocol stacks can be used with each of those networks and network interfaces as described above.
[0073] The foregoing human interface device, the human-accessible memory device, and the network interface 1054 may be attached to the core 1040 of the computer system 1000.
[0074] The core 1040 may include one or more central processing units (CPUs) 1041, a graphics processing unit (GPU) 1042, a dedicated programmable processing device in the form of a field programmable gate array (FPGA) 1043, a hardware accelerator 1044 for specific tasks, etc. These devices may be connected via a system bus 1048 together with a read-only memory (ROM) 1045, a random access memory 1046, an internal mass storage 1047 such as an internal hard drive or SSD that is not accessible to internal users. In some computer systems, the system bus 1048 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus 1048 or via a peripheral bus 1049. Architectures for peripheral buses include PCI, USB, etc. A graphics adapter 1050 may be included in the core 1040.
[0075] The CPU 1041, GPU 1042, FPGA 1043, and accelerator 1044 may execute specific instructions that can together constitute the aforementioned computer code. The computer code may be stored in the ROM 1045 or the RAM 1046. Temporary data may also be stored in the RAM 1046, while persistent data may be stored, for example, in the internal mass storage 1047. Fast storage and retrieval to / from any of the memory devices may be enabled by the use of a cache memory that may be closely associated with one or more CPUs 1041, GPUs 1042, the mass storage 1047, the ROM 1045, the RAM 1046, etc.
[0076] The computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of the present disclosure or may be of the kind well-known and available to those having skill in the computer software arts.
[0077] By way of example and not limitation, a computer system 1000 having an architecture, specifically a core 1040, may provide functionality as a result of a processor (including, e.g., a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be associated with user-accessible mass storage as introduced above, as well as media associated with specific storage of the core 1040 that is non-transitory in nature, such as core internal mass storage 1047 or ROM 1045. The software implementing various embodiments of the present disclosure may be stored on such devices and executed by the core 1040. The computer-readable media may include one or more memory devices or chips, depending on specific requirements. The software may cause the core 1040, specifically a processor therein (including, e.g., a CPU, GPU, FPGA, etc.), to define data structures stored in RAM 1046 and modify such data structures according to processes defined by the software, thereby executing specific processes or specific portions of specific processes described herein. Additionally or alternatively, the computer system may provide functionality as a result of logic hard-wired into a circuit (e.g., accelerator 1044) or otherwise embodied, which may operate instead of or in conjunction with the software to execute specific processes or specific portions of specific processes described herein. References to software may, as necessary, include logic, and vice versa. References to computer-readable media may, as necessary, include circuits (such as integrated circuits (ICs)) that store software for execution, circuits that embody logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.
[0078] Although several non-limiting embodiments have been described, there are changes, rearrangements, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure, even though not explicitly shown or described herein.
[0079] The above disclosure also encompasses the embodiments listed below.
[0080] (1) A method implemented by at least one processor in a decoder, the method comprising: receiving a coded video bitstream including (i) one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh, and (ii) a 2D-3D index array that maps each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh; and reconstructing the 3D mesh using the 2D-3D index array and mapping each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh, wherein the 2D-3D index array is encoded in a tuple format including a first parameter specifying a start index of a run of consecutive integers in the 2D-3D index array, a second parameter specifying a length of the run, and a third parameter specifying a direction of the run.
[0081] (2) The method according to feature (1), wherein the first parameter specifying the start index of the run is coded by fixed-length coding, and the bit length of the codeword in the fixed-length coding is based on the number of unique 3D boundary vertices in the 3D mesh.
[0082] (3) The method according to feature (1) or (2), wherein the second parameter specifying the length of the run is coded by fixed-length coding.
[0083] (4) The bit length of the coded word in fixed-length coding is the method according to feature (3), based on the number of 2D boundary vertices within one or more 2D meshes.
[0084] (5) The number of 2D boundary vertices is the method according to feature (4), signaled in the coded video bitstream.
[0085] (6) The bit length of the coded word in fixed-length coding is the method according to feature (3), based on the remaining number of 2D boundary vertices within the current 2D mesh of one or more 2D meshes that have not yet been added to the 2D-3D index array.
[0086] (7) The third parameter specifying the run direction is coded by one of 1-bit bypass coding or arithmetic coding using context, according to the method described in any one of features (1) to (6).
[0087] (8) The second parameter specifying the run length is coded by variable-length coding, according to the method described in any one of features (1) to (7).
[0088] (9) The bit length of the coded word in variable-length coding is the method according to feature (8), based on the number of 2D boundary vertices within the current 2D mesh of one or more 2D meshes.
[0089] (10) The bit length of the coded word in variable-length coding is the method according to feature (8), based on the remaining number of 2D boundary vertices within the current 2D mesh of one or more 2D meshes that have not yet been added to the 2D-3D index array.
[0090] (11) The second parameter specifying the run length is predicted based on the length of the previous run, according to the method described in any one of features (1) to (10).
[0091] (12) The 2D-3D index array is ordered such that the tuple among the plurality of tuples having the longest run length is in front of the 2D-3D index array, the method according to any one of features (1) to (11).
[0092] (13) At least one memory configured to store program code, and at least one processor configured to read the program code and operate as commanded by the program code, wherein the program code causes the at least one processor to: (i) one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh, and (ii) receive a coded video bitstream including a 2D-3D index array that maps each vertex in one or more 2D meshes to respective vertices in the 3D mesh; and a reconstruction code configured to cause the at least one processor to reconstruct the 3D mesh using the 2D-3D index array and map each vertex in one or more 2D meshes to respective vertices in the 3D mesh, wherein the 2D-3D index array is encoded in a tuple format including a first parameter specifying a start index of a run of consecutive integers in the 2D-3D index array, a second parameter specifying the length of the run, and a third parameter specifying the direction of the run, a decoder.
[0093] (14) The first parameter specifying the start index of the run is coded by fixed-length coding, and the bit length of the codeword in the fixed-length coding is based on the number of unique 3D boundary vertices in the 3D mesh, the decoder according to feature (13).
[0094] (15) The second parameter specifying the length of the run is coded by fixed-length coding, the decoder according to feature (13) or (14).
[0095] (16) The bit length of the codeword in fixed-length coding is the decoder according to any one of features (13) to (15), based on the number of 2D boundary vertices in one or more 2D meshes.
[0096] (17) The number of 2D boundary vertices is the decoder according to feature (16), signaled in the coded video bitstream.
[0097] (18) The bit length of the codeword in fixed-length coding is the decoder according to feature (15), based on the remaining number of 2D boundary vertices in the current 2D mesh of one or more 2D meshes that have not yet been added to the 2D-3D index array.
[0098] (19) The third parameter specifying the run direction is coded by one of 1-bit bypass coding or arithmetic coding using context, for the decoder according to any one of features (13) to (18).
[0099] (20) When executed by a processor in the decoder, the processor is caused to (i) receive a coded video bitstream including one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh and (ii) a 2D-3D index array that maps each vertex in the one or more 2D meshes to respective vertices in the 3D mesh, and to use the 2D-3D index array to reconstruct the 3D mesh and map each vertex in the one or more 2D meshes to the 3D mesh, where the 2D-3D index array is encoded in a tuple format including a first parameter specifying a start index of a run of consecutive integers for each tuple in the 2D-3D index array, a second parameter specifying the length of the run, and a third parameter specifying the direction of the run, a non-transitory computer-readable medium storing the instructions.
Description of Signs
[0100] 100 Communication system, 110 Terminal, First terminal, 120 Terminal, Second terminal, 130 Terminal, 140 Terminal, 150 Communication network, 200 Streaming system, 201 Video source, 202 Sample stream, 203 Encoder, 204 Encoded video bitstream, 205 Streaming server, 206 Streaming client, 209 Video bitstream, 210 Video decoder, 211 Video sample stream, 212 Display, 213 Capture subsystem, 300 Framework, 301 Encoder, 305 Input mesh, 310 Mesh with UV atlas, 315 Occupancy map, 320 Geometry map, 325 Attribute map, 330 Metadata, 335 Decoded occupancy map, 340 Decoded geometry map, 345 Decoded attribute map, 350 Decoded metadata, 351 Decoder, 360 Reconstructed mesh, 400 3D mesh segment, 402 2D chart, 404 2D chart, 500 3D mesh segment, 502A 2D mesh segment, 502B 2D mesh segment, 600 UV atlas, 800 Process, S802 Operation, S804 Operation, S806 Operation, 900 Process, S902 Operation, S904 Operation, 1000 Computer system, 1001 Keyboard, 1002 Mouse, 1003 Trackpad, 1005 Joystick, 1006 Microphone, 1007 Scanner, 1008 Camera, 1009 Speaker, 1010 Touch screen, 1020 CD / DVD ROM / RW, 1021 CD / DVD / Similar media, 1022 Thumb drive, 1023 Removable hard drive or solid state drive, 1040 Core, 1041 Central processing unit (CPU), 1042 Graphics processing unit (GPU), 1043 Field programmable gate array (FPGA), 1044 Hardware accelerator, 1045 Read-only memory (ROM), 1046 Random access memory, 1047 Internal mass storage, 1048 System bus, 1049 Peripheral bus, 1050 Graphics adapter, 1054Network interface, 1055 cloud computing environment
Claims
Claim 1 A method performed by at least one processor in a decoder, the method comprising: receiving a coded video bitstream comprising (i) one or more two-dimensional (2D) meshes corresponding to a three-dimensional (3D) mesh and (ii) a 2D-3D index array mapping each vertex in the one or more 2D meshes to a respective vertex in the 3D mesh; reconstructing the 3D mesh using the 2D-3D index array to map each vertex in the one or more 2D meshes to the respective vertex in the 3D mesh; wherein the 2D-3D index array is encoded in a tuple format comprising a first parameter specifying a start index of a run of consecutive integers in the 2D-3D index array, a second parameter specifying a length of the run, and a third parameter specifying a direction of the run; method. Claim 2 The method of claim 1, wherein the first parameter specifying the start index of the run is coded by fixed-length coding, and a bit length of a codeword in the fixed-length coding is based on a number of unique 3D boundary vertices in the 3D mesh. Claim 3 The method of claim 1, wherein the second parameter specifying the length of the run is coded by fixed-length coding. Claim 4 The method of claim 3, wherein the bit length of the codeword in the fixed-length coding is based on a number of 2D boundary vertices in the one or more 2D meshes. Claim 5 The method of claim 4, wherein the number of 2D boundary vertices is signaled in the coded video bitstream. Claim 6 The method of claim 3, wherein the bit length of the codeword in the fixed-length coding is based on a remaining number of 2D boundary vertices in a current 2D mesh of the one or more 2D meshes not yet added to the 2D-3D index array. Claim 7 The method of claim 1, wherein the third parameter specifying the direction of the run is coded by one of 1-bit bypass coding or arithmetic coding using context. Claim 8 The method according to claim 1, wherein the second parameter specifying the length of the run is coded by variable length coding.
9. The method according to claim 8, wherein the bit length of the code word in the variable length coding is based on the number of 2D boundary vertices in the current 2D mesh of the one or more 2D meshes.
10. The method according to claim 8, wherein the bit length of the code word in the variable length coding is based on the remaining number of 2D boundary vertices in the current 2D mesh of the one or more 2D meshes that have not yet been added to the 2D-3D index array.
11. The method according to claim 1, wherein the second parameter specifying the length of the run is predicted based on the length of the previous run.
12. The method according to claim 1, wherein the 2D-3D index array is ordered such that a tuple among a plurality of tuples having the longest run length is in front of the 2D-3D index array.
13. An apparatus configured to perform the method according to any one of claims 1 to 12.
14. A computer program for causing a processor to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and device for coding / Decoding three dimensional mesh information with novelty
JP2000224582A
3D mesh information encoding and decoding device and method
JP2008538435A
Information generation device, information processing device, control method, program, and data structure
JP2019086918A