Distortion information for each iteration of vertex reconstruction

By using the vertex count and distortion information of the original mesh to simplify and reconstruct the sub-mesh, the resource waste and rendering problems in the existing technology are solved, and more efficient 3D multimedia data transmission and display are achieved.

CN120642338APending Publication Date: 2025-09-12SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480013633.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-04
Filing Date
2024-04-17
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When generating and reconstructing three-dimensional multimedia data, existing technologies cannot effectively utilize the number of vertices and distortion information of the original mesh, resulting in resource waste and rendering problems, and cannot select the appropriate number of subdivision iterations according to application context and user needs.

Method used

By using the vertex count of the original mesh and the distortion information of each subdivision iteration to simplify and rebuild the sub-mesh, we avoid repeatedly generating the base mesh, improve compression efficiency, and select the appropriate number of iterations based on the application context.

Benefits of technology

It achieves more efficient resource utilization and rendering quality, reduces resource waste, adapts to the needs of different application scenarios, and improves the transmission and display efficiency of 3D multimedia data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120642338A_ABST
    Figure CN120642338A_ABST
Patent Text Reader

Abstract

In one embodiment, an apparatus includes a communication interface configured to receive a compressed bitstream including a base grid bitstream; and a processor operably coupled to the communication interface. The processor is configured to decode the plurality of sub-grids from the base grid bitstream. The processor is further configured to subdivide a sub-grid of the plurality of sub-grids according to a subdivision iteration count to generate at least one subdivided sub-grid, the process includes determining a plurality of vertex positions for the at least one submesh by using a number of vertices associated with the original submesh and by using distortion information between the original submesh and the base mesh at each subdivision iteration associated with the subdivision iteration count. The processor is further configured to reconstruct at least a portion of the grid frame using the vertex positions corresponding to the at least one fine sub-grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to multimedia devices and processes. More particularly, the present disclosure relates to using distortion information for each iteration of vertex reconstruction. Background Art

[0002] Thanks to the availability of high-performance handheld devices such as smartphones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are becoming new ways to experience immersive content. 360° video enables consumers to experience an immersive, "real-life," "be-there" experience by capturing a 360° view of the world, while 3D volumetric video provides a full six degrees of freedom (DoF) experience of immersion and movement within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movements in real time to determine the area of ​​360° video or volumetric video content that the user wants to view or interact with. Multimedia data with 3D properties, such as point clouds or 3D polygon meshes, can be used in immersive environments. This data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream. Summary of the Invention

[0003] In an embodiment, a device includes: a communication interface configured to receive a compressed bitstream including a base mesh sub-bitstream; and a processor operably coupled to the communication interface. The processor is configured to decode a plurality of sub-meshes from the base mesh sub-bitstream. The processor is configured to subdivide a sub-mesh of the plurality of sub-meshes according to a subdivision iteration count to generate at least one subdivided sub-mesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided sub-mesh by using a number of vertices associated with an original sub-mesh and using distortion information between the original sub-mesh and the base mesh at each subdivision iteration associated with the subdivision iteration count. The processor is configured to reconstruct at least a portion of a mesh frame using the vertex positions corresponding to the at least one subdivided sub-mesh.

[0004] In an embodiment, a method includes receiving a compressed bitstream comprising a base mesh sub-bitstream. The method includes decoding a plurality of sub-meshes from the base mesh sub-bitstream. The method includes subdividing a sub-mesh of the plurality of sub-meshes according to a subdivision iteration count to generate at least one subdivided sub-mesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided sub-mesh by using a number of vertices associated with an original sub-mesh and using distortion information between the original sub-mesh and the base mesh at each subdivision iteration associated with the subdivision iteration count. The method includes reconstructing at least a portion of a mesh frame using the vertex positions corresponding to the at least one subdivided sub-mesh.

[0005] In an embodiment, a device includes: a communication interface; and a processor operably coupled to the communication interface. The processor is configured to subdivide a submesh according to a subdivision iteration count to generate at least one subdivided submesh having vertex positions. The processor is configured to determine a number of vertices associated with the submesh to be used to simplify the vertex positions. The processor is configured to determine distortion information between the submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count. The processor is configured to reconstruct at least a portion of a mesh frame using the vertex positions of the at least one subdivided submesh. The processor is configured to create a compressed bitstream, the compressed bitstream including information about the submesh, the subdivision iteration count, the number of vertices associated with the submesh, and the distortion information.

[0006] In an embodiment, a method includes subdividing a submesh according to a subdivision iteration count to generate at least one subdivided submesh having vertex positions. The method includes determining a number of vertices associated with the submesh to be used to simplify the vertex positions. The method includes determining distortion information between the submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count. The method includes reconstructing at least a portion of a mesh frame using the vertex positions of the at least one subdivided submesh. The method includes creating a compressed bitstream including information about the submesh, the subdivision iteration count, the number of vertices associated with the submesh, and the distortion information.

[0007] In one embodiment, a device includes: a communication interface configured to receive a compressed bitstream including a signaling element identifying a first sub-mesh whose simplified mesh is to be copied when reconstructing a second sub-mesh and how to manipulate the copied first simplified mesh; and a processor operably coupled to the communication interface. The processor is configured to decode at least a portion of the bitstream. The processor is configured to identify that the current sub-mesh is a second sub-mesh for which a copy of the simplified mesh of the first sub-mesh should be used. The processor is configured to copy and manipulate the simplified mesh based on the signaling element. The processor is configured to output decoded and reconstructed content.

[0008] In one embodiment, a method includes receiving a compressed bitstream including a signaling element that identifies a first sub-mesh whose simplified mesh is to be replicated when reconstructing a second sub-mesh and how to manipulate the replicated first simplified mesh, and a processor operably coupled to a communication interface. The method includes decoding at least a portion of the bitstream. The method also includes identifying that the current sub-mesh is a second sub-mesh that should use a replicated simplified mesh of the first sub-mesh. The method includes replicating the simplified mesh and manipulating the simplified mesh based on the signaling element. The method also includes outputting the decoded and reconstructed content.

[0009] In an embodiment, a device includes: a communication interface; and a processor operably coupled to the communication interface. The processor is configured to identify that a first simplified mesh for a first original sub-mesh is identical to a second simplified mesh for a second original sub-mesh. The processor is configured to select the first simplified mesh for use in reconstructing the second sub-mesh. The processor is configured to construct a signaling element to instruct a decoder to copy the first simplified mesh when reconstructing the second sub-mesh, and to instruct the decoder how to manipulate the copied first simplified mesh. The processor is configured to encode and output a bitstream including the signaling element.

[0010] In an embodiment, a method includes identifying that a first simplified mesh for a first original sub-mesh is identical to a second simplified mesh for a second original sub-mesh. The method includes selecting the first simplified mesh for use in reconstructing the second sub-mesh. The method includes constructing a signaling element to instruct a decoder to copy the first simplified mesh when reconstructing the second sub-mesh, and instructing the decoder how to manipulate the copied first simplified mesh. The method includes encoding and outputting a bitstream including the signaling element.

[0011] Other technical features may be easily understood by those skilled in the art based on the following drawings, descriptions and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like parts: Figure 1 An example communication system according to the present disclosure is shown; Figure 2 and Figure 3 An example electronic device according to the present disclosure is shown; Figure 4 An example intra-frame encoding process according to the present disclosure is shown; Figure 5 An example trellis frame decoding process according to the present disclosure is shown; Figure 6 An example process for reconstructing a sub-grid according to the present disclosure is shown; Figure 7 An example process for reconstructing a simplified sub-mesh according to the present disclosure is shown; Figure 8 An example encoding method according to the present disclosure that allows reconstruction of a simplified sub-grid is shown; Figure 9 An example decoding method for reconstructing a simplified sub-grid according to the present disclosure is shown; 10A and 10B illustrate an example process for creating a simplified mesh corresponding to an original mesh according to the present disclosure; Figure 11An example encoding method for creating and signaling repeated base grid information according to the present disclosure is shown; and Figure 12 An example decoding method for using a replicated base mesh during mesh reconstruction according to the present disclosure is shown. DETAILED DESCRIPTION

[0013] Before proceeding with the detailed description below, it may be helpful to set forth the definitions of specific words and phrases used throughout this patent document. The term "in conjunction with" and its derivatives refer to any direct or indirect communication between two or more elements, regardless of whether those elements are in physical contact with one another. The terms "send," "receive," and "communicate," and their derivatives, encompass both direct and indirect communication. The terms "include," "comprise," and their derivatives, mean, but are not limited to, these. The term "or" is inclusive, meaning and / or. The phrase "associated with" and its derivatives mean including, included within, interconnected with, containing, contained within, connected to or connected with, coupled to or coupled with, communicable with, collaborate with, interlaced, juxtaposed, proximate to, bound to or bound with, having, having the property of, having a relationship to, or having a relationship with, etc. The term "controller" means any device, system, or component thereof that controls at least one operation. Such a controller may be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller may be centralized or distributed, whether local or remote. When used with a list of items, the phrase "at least one of" means that different combinations of one or more of the listed items can be used, and only one item in the list may be required. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A, B, and C.

[0014] Furthermore, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed from computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof, suitable for implementation in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer (such as read-only memory (ROM), random-access memory (RAM), hard drives, compact disks (CDs), digital video disks (DVDs), or any other type of memory). "Non-transitory" computer-readable media does not include wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable media includes both media that can permanently store data and media that can store data and later rewrite it (such as rewritable optical disks or erasable memory devices).

[0015] Definitions for other specific words and phrases are provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to prior, as well as future uses of such defined words and phrases.

[0016] Described below Figures 1 to 12 The various embodiments used to describe the principles of the present disclosure are merely exemplary and should not be interpreted in any way as limiting the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any type of suitably arranged device or system.

[0017] As mentioned above, thanks to the availability of high-performance handheld devices such as smartphones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are emerging as new ways to experience immersive content. 360° video enables consumers to experience an immersive, "real-life," "be-there" experience by capturing a 360° view of the world, while 3D volumetric video provides a full six degrees of freedom (DoF) experience of immersion and movement within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Display and navigation sensors can track the user's head movements in real time to determine the area of ​​the 360° video or volumetric video content that the user wants to view or interact with. Multimedia data with 3D properties, such as point clouds or 3D polygon meshes, can be used in immersive environments. This data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream.

[0018] A point cloud is a set of 3D points with properties such as color, normal direction, reflectivity, point size, etc. that represent the surface or volume of an object. Point clouds are common in various applications such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view playback, and six-degree-of-freedom (DoF) immersive media. If point clouds are not compressed, they typically require a large amount of bandwidth to transmit. Due to the large bitrate requirements, point clouds are typically compressed before transmission. Compressing 3D objects such as point clouds typically requires specialized hardware. To avoid the use of specialized hardware to compress 3D point clouds, 3D point clouds can be transformed into traditional two-dimensional (2D) frames, which can be compressed and later reconstructed so that they can be viewed by the user.

[0019] Polygonal 3D meshes, particularly triangle meshes, are another common format for representing 3D objects. A mesh typically consists of a collection of vertices, edges, and faces that represent the surface of a 3D object. A triangle mesh is a simple polygonal mesh where the faces are simple triangles that cover the surface of the 3D object. Typically, there may be one or more attributes associated with a mesh. In one scenario, one or more attributes may be associated with each vertex in the mesh. For example, a texture attribute (RGB) may be associated with each vertex. In another scenario, each vertex may be associated with a pair of coordinates (u, v). The (u, v) coordinates may refer to a location in a texture map associated with the mesh. For example, the (u, v) coordinates may refer to a row index and a column index, respectively, in a texture map. A mesh can be thought of as a point cloud with additional connectivity information.

[0020] Point clouds or meshes can be dynamic, meaning they change over time. In these cases, a point cloud or mesh at a specific moment in time is referred to as a point cloud frame or mesh frame, respectively. Because point clouds and meshes contain large amounts of data, they require compression for efficient storage and transmission. This is especially true for dynamic point clouds and meshes, which can contain 60 or more frames per second.

[0021] As part of the encoding process, a base mesh can be encoded using an existing mesh codec, and a reconstructed base mesh can be constructed from the encoded original mesh. The reconstructed base mesh can then be subdivided into one or more subdivided meshes, and a displacement field created for each subdivided mesh. For example, if the reconstructed base mesh includes triangles covering the surface of a 3D object, the triangles are subdivided according to the number of subdivision levels applied, such as creating a first subdivided mesh of four triangles for each triangle of the reconstructed base mesh, a second subdivided mesh of sixteen triangles for each triangle of the reconstructed base mesh, and so on. Each displacement field represents the difference between the vertex positions of the original mesh and the subdivided mesh associated with the displacement field. Each displacement field is wavelet transformed to create a level-of-detail (LOD) signal, which is encoded as part of the compressed bitstream. During decoding, the displacement of each displacement field is added to its associated subdivided mesh to reconstruct a version of the original mesh.

[0022] To create the base mesh and displacement field, the original mesh is first downsampled to generate base curves / polylines, referred to as "simplified" curves. A subdivision scheme is then applied to the simplified polylines to generate "tessellated" curves. A subdivision scheme using an iterative interpolation scheme may be applied that involves inserting a new point in the middle of each edge of the polyline at each iteration. However, the number of vertices of the reconstructed mesh may not be the same as the number of vertices of the original mesh because the number of vertices of the reconstructed mesh increases as the number of subdivision iterations increases. The encoder selects the number of iterations so that the similarity between the original mesh and the reconstructed base mesh reaches a certain similarity level. However, the similarity level selected by the encoder may not be appropriate for various reasons.

[0023] For example, when a 3D scene consists of a certain number of objects, objects in the background of the 3D scene may not need to be reconstructed to achieve high similarity, but objects in the front of the 3D scene may need to be reconstructed to achieve high similarity. For example, in a six-DoF application, each individual user can freely choose the position and orientation of the view in 3D space. This means that the encoder cannot determine which objects may require high similarity and which may not. This determination must be made by the decoder. Furthermore, as the number of interpolation rounds increases, the number of points used for the base mesh increases, and the distortion between the original and base mesh decreases. In some cases, the number of vertices in the reconstructed mesh may differ from the original mesh, and often the base mesh may have more vertices than the original mesh. This can cause rendering issues when the renderer has limited capabilities, so objects may need to be simplified to meet the rendering system's limitations. Furthermore, applying a large number of iterations can waste resources because the reconstructed mesh may not be identical to the original. Therefore, applying iterations may be more efficient when the distortion between the original and reconstructed meshes reaches a certain level.

[0024] Therefore, the present disclosure provides a method for using the number of vertices of the original mesh as a reference number for simplification while avoiding unnecessarily sacrificing the quality of the mesh. The present disclosure also provides the following technology: allowing the decoder to use the distortion information at each iteration to understand the similarity between the original mesh and the reconstructed base mesh at each iteration, so that the decoder can select the number of iterations to be applied to generate the base mesh according to the context of the application and the user.

[0025] As described above, a base mesh is created that is a simplified version of the original mesh to minimize the amount of compressed data. Because the original mesh is simplified to a simplified version, in some cases, the base meshes created for two different original meshes may be identical, even if the two original meshes are different. This may result in the generation, storage, and use of more than one identical base mesh, resulting in a waste of resources. Therefore, the present disclosure provides a method for eliminating duplicate base meshes to further improve compression efficiency.

[0026] In some examples of the present disclosure, the term "submesh" may refer to a partition of a base mesh. In some examples, in the present disclosure, a "submesh" may refer to geometric data reconstructed after the submesh is subdivided and displacements are added.

[0027] Figure 1 An example communication system 100 according to the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown in FIGURE 1 is for illustration only. Other embodiments of the communication system 100 may be used without departing from the scope of this disclosure.

[0028] like Figure 1 As shown, the communication system 100 includes a network 102 that supports communication between various components in the communication system 100. For example, the network 102 can transmit IP packets, frame relay frames, asynchronous transfer mode (ATM) cells, or other information between network addresses. The network 102 includes all or part of one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), a global network such as the Internet, or any other communication system at one or more locations.

[0029] In this example, network 102 supports communication between server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, TVs, interactive displays, wearable devices, head-mounted displays (HMDs), and the like. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device capable of providing computing services to one or more client devices, such as client devices 106-116. Each server 104 may, for example, include one or more processing devices, one or more memories for storing instructions and data, and one or more network interfaces for supporting communication over network 102. As described in more detail below, server 104 may transmit a compressed bitstream representing a point cloud or mesh to one or more display devices, such as client devices 106-116. In embodiments, each server 104 may include an encoder.

[0030] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other computing device via network 102. Client devices 106-116 include desktop computers 106, mobile phones or mobile devices 108 (such as smartphones), PDAs 110, laptop computers 112, tablet computers 114, and head-mounted display (HMD) 116. However, any other or additional client devices may be used in communication system 100. Smartphones represent one type of mobile device 108, which is a handheld device with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and internet data communications. HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments, any of client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 may record 3D volumetric video and then encode the video so that it can be transmitted to one of client devices 106-116. In another example, the laptop computer 112 may be used to generate a 3D point cloud or mesh, and then encode and send the 3D point cloud or mesh to one of the client devices 106 - 116 .

[0031] In this example, some client devices 108-116 communicate indirectly with network 102. For example, mobile device 108 and PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). Furthermore, laptop computer 112, tablet computer 114, and HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. It should be noted that this is for illustration only, and each client device 106-116 may communicate directly with network 102 or indirectly via any suitable intermediary device or network. In embodiments, server 104 or any client device 106-116 may be configured to compress a point cloud or mesh, generate a bitstream representing the point cloud or mesh, and transmit the bitstream to another client device, such as any client device 106-116.

[0032] In embodiments, any of the client devices 106-114 securely and efficiently transmits information to another device, such as, for example, the server 104. Furthermore, any of the client devices 106-116 can trigger information transfer between itself and the server 104. Any of the client devices 106-114 can function as a VR display when attached to a head-mounted device via a mount, and can function similarly to the HMD 116. For example, when the mobile device 108 is attached to the mount system and worn over the user's eyes, the mobile device 108 can function similarly to the HMD 116. The mobile device 108 (or any other client device 106-116) can trigger information transfer between itself and the server 104.

[0033] In an embodiment, any of the client devices 106-116 or the server 104 may perform the following operations: create a 3D point cloud or mesh, compress a 3D point cloud or mesh, transmit a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination thereof. For example, the server 104 may compress the 3D point cloud or mesh to generate a bitstream, and then transmit the bitstream to one or more of the client devices 106-116. As another example, one of the client devices 106-116 may compress the 3D point cloud or mesh to generate a bitstream, and then transmit the bitstream to another of the client devices 106-116 or the server 104. In other words, the server 104 and / or the client devices 106-116 may be devices as described herein. According to the present disclosure, the server 104 and / or the client devices 106-116 may use the number of vertices of the original mesh and / or distortion information for each reconstruction iteration to simplify the sub-mesh. Additionally or alternatively, according to the present disclosure, the server 104 and / or the client devices 106-116 may use a copy of the simplified mesh to reconstruct one or more sub-meshes. In an embodiment, the server 104 and / or the client devices 106-116 may construct and transmit signaling information that instructs another device to use the number of vertices of the original mesh and / or distortion information for each reconstruction iteration to simplify the sub-mesh and / or to create and use a copy of the simplified mesh to reconstruct one or more sub-meshes.

[0034] although Figure 1 One example of a communication system 100 is shown, but may be used for Figure 1 Various changes may be made. For example, the communication system 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of this disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.

[0035] Figure 2 and Figure 3 An example electronic device according to the present disclosure is shown. In particular, Figure 2 An example server 200 is shown and may represent Figure 1 The server 104 in FIG. The server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single seamless resource pool, cloud-based servers, etc. Figure 1 The server 200 may be accessed by one or more of the client devices 106 - 116 or another server.

[0036] like Figure 2 As shown, server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers (such as encoders). In an embodiment, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communication between at least one processing device (such as a processor 210 ), at least one storage device 215 , at least one communication interface 220 , and at least one input / output (I / O) unit 225 .

[0037] The processor 210 executes instructions that may be stored in the memory 230. The processor 210 may include any suitable number and type of processors or other devices arranged in any suitable manner. Example types of processors 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.

[0038] In an embodiment, the processor 210 may encode a 3D point cloud or mesh stored in the storage device 215. In an embodiment, encoding the 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding. In an embodiment, the processor 210 may use the number of vertices of the original mesh and / or distortion information for each reconstruction iteration to simplify a sub-mesh. Additionally or alternatively, as described in the present disclosure, the processor 210 may create and use a copy of the simplified mesh to reconstruct one or more sub-meshes. In an embodiment, the processor 210 may construct and transmit signaling information that instructs another device to use the number of vertices of the original mesh and / or distortion information for each reconstruction iteration to simplify a sub-mesh and / or create and use a copy of the simplified mesh to reconstruct one or more sub-meshes.

[0039] Memory 230 and persistent storage 235 are examples of storage 215, which represent any structure capable of storing and facilitating retrieval of information, such as data of a temporary or persistent nature, program code, or other suitable information. Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage. For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing the patches onto 2D frames, instructions for compressing the 2D frames, and instructions for encoding the 2D frames in a particular order to generate a bitstream. The instructions stored in memory 230 may also include instructions for viewing an omnidirectional 360° scene (e.g., through a VR headset such as Figure 1Persistent storage 235 may include one or more components or devices that support long-term storage of data (such as read-only memory, a hard drive, flash memory, or an optical disk).

[0040] The communication interface 220 supports communication with other systems or devices. For example, the communication interface 220 may include support for Figure 1 The communication interface 220 may include a network interface card or wireless transceiver for communicating with the network 102. The communication interface 220 may support communication via any suitable physical or wireless communication link. For example, the communication interface 220 may send a bitstream containing a 3D point cloud to another device (such as one of the client devices 106-116).

[0041] The I / O unit 225 allows for the input and output of data. For example, the I / O unit 225 can provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. The I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it should be noted that the I / O unit 225 can be omitted, such as when I / O interaction with the server 200 occurs via a network connection.

[0042] It should be noted that although Figure 2 Described as indicating Figure 1 The server 104 of FIG. 1 may be configured as a server 104, but the same or similar architecture may be used in one or more of the various client devices 106-116. For example, a desktop computer 106 or a laptop computer 112 may have a server 104 configured as a server 104. Figure 2 The same or similar structure as shown in FIG.

[0043] Figure 3 An example electronic device 300 is shown and may represent Figure 1 The electronic device 300 may be a mobile communication device (such as, for example, a mobile station, a user station, a wireless terminal, a desktop computer (similar to Figure 1 Desktop computer 106), portable electronic device (similar to Figure 1 108, PDA 110, laptop computer 112, tablet computer 114, or HMD 116) etc. In an embodiment, Figure 1 One or more of the client devices 106-116 may include a configuration that is the same as or similar to the configuration of the electronic device 300. In an embodiment, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 may be used with data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.

[0044] like Figure 3 As shown, electronic device 300 includes antenna 305, radio frequency (RF) transceiver 310, transmit (TX) processing circuitry 315, microphone 320, and receive (RX) processing circuitry 325. RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a WiFi transceiver, a ZigBee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes speaker 330, a processor, an input / output (I / O) interface (IF) 345, an input 350, a display 355, memory 360, and sensors 365. Memory 360 includes an operating system (OS) 361 and one or more applications 362.

[0045] The RF transceiver 310 receives incoming RF signals from the antenna 305, transmitted from an access point (such as a base station, a Wi-Fi router, or a Bluetooth device) or other device on the network 102 (such as Wi-Fi, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). The RF transceiver 310 downconverts the incoming RF signal to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to the RX processing circuitry 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. The RX processing circuitry 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or to a processor for further processing (such as for web browsing data).

[0046] TX processing circuitry 315 receives analog or digital voice data from microphone 320, or other outgoing baseband data from the processor. The outgoing baseband data may include web data, email, or interactive video game data. TX processing circuitry 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency (IF) signal. RF transceiver 310 receives the outgoing processed baseband or IF signal from TX processing circuitry 315 and up-converts the baseband or IF signal into an RF signal that is transmitted via antenna 305.

[0047] Processor 340 may include one or more processors or other processing devices. Processor 340 may execute instructions stored in memory 360 (such as OS 361) to control the overall operation of electronic device 300. For example, processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals via RF transceiver 310, RX processing circuitry 325, and TX processing circuitry 315 according to well-known principles. Processor 340 may include any suitable number and type of processors or other devices arranged in any suitable manner. For example, in an embodiment, processor 340 includes at least one microprocessor or microcontroller. Example types of processor 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application-specific integrated circuits, and discrete circuits.

[0048] The processor is capable of executing other processes and programs stored in memory 360, such as operations to receive and store data. Processor 340 can move data into or out of memory 360 as needed for the executed process. In an embodiment, processor 340 is configured to run one or more applications 362 based on OS 361 or in response to signals received from an external source or operator. For example, applications 362 may include encoders, decoders, VR or AR applications, camera applications (for still images and video), video phone calling applications, email clients, social media clients, SMS messaging clients, virtual assistants, etc. In an embodiment, processor 340 is configured to receive and send media content.

[0049] In an embodiment, the processor 340 may use the number of vertices of the original mesh and / or the distortion information for each reconstruction iteration to simplify the sub-mesh. Additionally or alternatively, as described in the present disclosure, the processor 340 may create and use a copy of the simplified mesh to reconstruct one or more sub-meshes. In an embodiment, the processor 340 may construct and transmit signaling information instructing another device to use the number of vertices of the original mesh and / or the distortion information for each reconstruction iteration to simplify the sub-mesh and / or create and use a copy of the simplified mesh to reconstruct one or more sub-meshes.

[0050] The processor 340 is also coupled to an I / O interface 345 that provides the electronic device 300 with the ability to connect to other devices, such as the client devices 106 - 114 . The I / O interface 345 is the communication path between these accessories and the processor 340 .

[0051] Processor 340 is also coupled to input 350 and display 355. An operator of electronic device 300 can use input 350 to enter data or input into electronic device 300. Input 350 can be a keyboard, touch screen, mouse, trackball, voice input, or other device capable of serving as a user interface to allow the user to interact with electronic device 300. For example, input 350 may include voice recognition processing, allowing the user to enter voice commands. In another example, input 350 may include a touch panel, a (digital) pen sensor, keys, or an ultrasonic input device. A touch panel can recognize touch input, for example, using at least one scheme (such as capacitive, pressure-sensitive, infrared, or ultrasonic). Input 350 can be associated with sensors 365 and / or cameras by providing additional input to processor 340. In embodiments, sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, and the like. Input 350 may also include control circuitry. In a capacitive solution, the input 350 may recognize touch or proximity.

[0052] Display 355 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active-matrix OLED (AMOLED), or other display capable of rendering text and / or graphics (such as from a website, video, game, image, etc.). Display 355 can be sized to fit within the HMD. Display 355 can be a single display screen or multiple displays capable of creating a stereoscopic display. In some embodiments, display 355 is a heads-up display (HUD). Display 355 can display 3D objects (such as a 3D point cloud or mesh).

[0053] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion of memory 360 may include flash memory or other ROM. Memory 360 may include a persistent storage device (not shown), which represents any structure capable of storing and facilitating retrieval of information (such as data, program code, and / or other suitable information). Memory 360 may include one or more components or devices that support long-term storage of data (such as read-only memory, a hard drive, flash memory, or an optical disk). Memory 360 may also include media content. The media content may include various types of media (such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, etc.).

[0054] The electronic device 300 also includes one or more sensors 365, which can measure physical quantities or detect the activation state of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensors 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyro sensor and an accelerometer), an eye tracking sensor, an air pressure sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalography (EEG) sensor, an electrocardiography (ECG) sensor, an IR sensor, an ultrasonic sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red, green, and blue (RGB) sensor), and the like. The sensors 365 may also include control circuitry for controlling any of the sensors included therein.

[0055] As discussed in more detail below, one or more of these sensors 365 can be used to control a user interface (UI), detect UI input, determine the orientation and facing direction of a user's three-dimensional content display recognition, etc. Any of these sensors 365 can be located within the electronic device 300, within an auxiliary device operably connected to the electronic device 300, within a head-mounted device configured to support the electronic device 300, or within a single device that includes the electronic device 300 and the head-mounted device.

[0056] The electronic device 300 may create media content (such as generating a virtual object or capturing (or recording) content through a camera). The electronic device 300 may encode the media content to generate a bitstream so that the bitstream can be sent directly to another electronic device or indirectly (such as through a Figure 1 The electronic device 300 may receive the bit stream directly from another electronic device, or indirectly (such as through a Figure 1 The network 102) receives the bit stream.

[0057] Although Figure 2 and Figure 3 An example of an electronic device is shown, but the Figure 2 and Figure 3 Make various changes. For example, you can combine, further subdivide or omit Figure 2 and Figure 3 As a specific example, processor 340 may be divided into multiple processors (such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs)). In addition, as with computing and communications, electronic devices and servers may have a variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.

[0058] Figure 4 An example intra-coding process 400 according to the present disclosure is shown. Figure 4 The intra-frame encoding process 400 shown in FIG. 4 is for illustration only. Figure 4 The scope of this disclosure is not limited to any particular implementation of the intra-frame coding process. For ease of explanation, Figure 4 The process 400 can be described as using Figure 3 However, process 400 may be used with any other suitable system and any other suitable electronic device.

[0059] like Figure 4 As shown, the intra-frame encoding process 400 encodes the grid frame using an intra-frame encoder 402. Figure 2 The server 200 shown in Figure 3 The electronic device 300 shown in FIG. 4 is used to represent or run an intra encoder 402. A base mesh 404, which typically has a smaller number of vertices than the original mesh, is created and quantized and compressed in a lossy or lossless manner, and then encoded into a compressed base mesh bitstream. Figure 4 As shown, the static mesh decoder decodes and reconstructs the base mesh, thereby providing a reconstructed base mesh 406. The reconstructed base mesh 406 is then subjected to one or more levels of subdivision, and a displacement field is created for each subdivision, representing the difference between the original mesh and the subdivided, reconstructed base mesh. In inter-frame coding of mesh frames, the base mesh 404 is encoded by transmitting vertex motion rather than directly compressing the base mesh. In either case, a displacement field 408 is created. Each displacement in the displacement field 408 has three components represented by x, y, and z. These can be relative to a standard coordinate system or a local coordinate system, where x, y, and z represent displacements in the local normal, tangent, and bitangent directions. It will be understood that multiple levels of subdivision can be applied, such that multiple subdivided mesh frames are created, and a displacement field is also created for each subdivided mesh frame.

[0060] Let the number of 3-D displacement vectors in the grid frame displacement 408 be N. Let the displacement field be given by The displacement field 408 is subjected to one or more levels of wavelet transformation 410 to create a level of detail (LOD) signal , where k represents the index of the detail level, represents the number of samples in the level of detail signal at level k, and numLOD represents the number of LODs. scalar quantized.

[0061] like Figure 4 As shown, the quantized LOD signal corresponding to the displacement field 408 is encoded into a compressed bitstream. In an embodiment, the quantized LOD signal is packed into a 2D image / video using an image packing operation and compressed losslessly or in a lossy manner using an image or video encoder. However, another entropy encoder (such as an asymmetric digital system (ANS) encoder or a binary arithmetic entropy encoder) can be used to losslessly encode the quantized LOD signal. There may be other dependencies that can be exploited based on previous samples, across components, and across LODs. The displacement component provides a displacement vector, which can be encoded into a geometric video component using any video codec indicated by the profile or using an SEI message. Optionally, the profile may indicate that the displacement component is encoded using arithmetic coding.

[0062] Also like Figure 4 As shown, image unpacking of the LOD signal is performed, and an inverse quantization operation and an inverse wavelet transform operation are performed to reconstruct the LOD signal. Another inverse quantization operation is performed on the reconstructed base mesh 406 and combined with the reconstructed LOD signal to reconstruct the deformed mesh. Attribute transfer operations are performed using the deformed mesh, static / dynamic mesh, and attribute map. A point cloud is a set of 3D points with attributes representing the surface or volume of an object (such as color, normal, reflectivity, point size, etc.). These attributes are encoded as a compressed attribute bitstream. Figure 4 As shown, encoding the compressed attribute bitstream may also include padding operations, color space conversion operations, and video encoding operations. In an embodiment, the atlas 405 may also be encoded as a compressed atlas bitstream. The atlas component provides information to the decoding and / or rendering system on how to perform inverse reconstruction. For example, the atlas may provide information on how to perform subdivision of the base mesh, how to apply displacement vectors to the subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.

[0063] Can be controlled by control process 412 Figure 4 The intra-frame encoding process 400 outputs a compressed bit stream that can be sent to and decoded by an electronic device (such as the server 104 or the client devices 106-116). Figure 4 As shown, the output compressed bitstream may include a compressed atlas bitstream, a compressed base grid bitstream, a compressed displacement bitstream, and a compressed attribute bitstream as sub-bitstreams of the compressed bitstream.

[0064] although Figure 4 An example intra-frame encoding process 400 is shown, but may be used for Figure 4Various changes may be made. For example, the number and arrangement of the various components of the intra-coding process 400 may be varied as needed or desired. Furthermore, the intra-coding process 400 may be used in any other suitable process and is not limited to the specific process described above. In an embodiment, only the first (x) component of the displacement may be created and encoded, and the other two components (y and z) may be assumed to be zero. In this case, a flag may be signaled in the bitstream to indicate that the bitstream only contains data corresponding to the first (x) component, and that the other two components (y and z) should be assumed to be zero when decompressing and reconstructing the displacement field 408. As another example, as described in the present disclosure, Figure 4 The intra-coding process 400 may include encoding a bitstream and / or implementing appropriate signaling to allow a decoder to use the number of vertices of the original mesh and / or distortion information for each reconstruction iteration to simplify sub-meshes and / or create and use a copy of the simplified mesh to reconstruct one or more sub-meshes.

[0065] Figure 5 An example trellis frame decoding process 500 according to the present disclosure is shown. Figure 5 The frame decoding process 500 shown in FIG. 5 is for illustration only. Figure 5 The scope of this disclosure is not limited to any particular implementation of the trellis frame decoding process. For ease of explanation, Figure 5 The process 500 can be described as using Figure 3 However, process 500 may be used with any other suitable system and any other suitable electronic device.

[0066] The decoding process 500 involves a demultiplexer 502 that receives an input bitstream. The demultiplexer separates various component bitstreams from the input bitstream, including a compressed base grid bitstream, a compressed displacement bitstream, and a compressed attribute bitstream (such as a bitstream about the Figure 4 The compressed attribute bitstream is decoded using a video decoder 504, the decoded attributes are processed using a color space conversion operation 506, and the original attributes for the mesh are restored.

[0067] The decoding process 500 also includes decoding the displacement bitstream using a video decoder 508, which, in embodiments, can be the same video decoder as the video decoder 504. The decoded displacement data is subjected to an image unpacking operation 510, an inverse quantization operation 512, and an inverse wavelet transform operation 514 as part of recovering position displacement data 516. Recovering the position displacement data 516 can also include performing one or more subdivision operations 518 on the grid frame recovered using a base grid decoder 520 and extracting x, y, and z components 522 (normal, tangent, bitangent) from the subdivided grid frame. The base grid decoder 520 can perform an inverse quantization operation 521 before performing the subdivision operation 518.

[0068] The base grid decoder 520 takes the base grid bitstream provided by the demultiplexer 502 and reconstructs an intra base grid frame from the base grid bitstream using the static grid decoder 524. The grid buffer 530 provides the decoded intra frame to the motion decoder 532. At step 534, the motion decoder 532 also receives the inter data and uses the intra data, the inter data, and the associated tables to reconstruct the base grid.

[0069] although Figure 5 An example frame decoding process 500 is shown, but may be used for Figure 5 Various changes may be made. For example, the number and arrangement of the various components of the frame decoding process 500 may be varied as needed or desired. Furthermore, the frame decoding process 500 may be used in any other suitable process and is not limited to the specific process described above. Furthermore, although shown as a series of steps, Figure 5 The steps in may overlap, occur in parallel, or occur any number of times. Figure 4 As described, the atlas bitstream may also be decoded to obtain an atlas that provides information on how to perform inverse reconstruction. For example, the atlas may provide information on how to perform subdivision of the base mesh, how to apply displacement vectors to subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.

[0070] Figure 6 An example process 600 for reconstructing a sub-mesh according to the present disclosure is shown. Figure 6 Process 600 is shown for illustration only. Figure 6 The scope of this disclosure is not limited to any particular implementation of the process for reconstructing a subgrid. For ease of explanation, Figure 6 The process 600 can be described as using Figure 3 However, process 600 may be used with any other suitable system and any other suitable electronic device.

[0071] like Figure 6As shown, to create the base mesh and displacement field, the original mesh 602 is first downsampled to generate a base curve / polyline, referred to as a simplified curve or mesh 604. A subdivision scheme is then applied to the simplified polyline to generate a subdivided curve or mesh 606. For example, Figure 6 As shown, a subdivision scheme using an iterative interpolation scheme is applied that includes inserting a new point in the middle of each edge of the polyline at each iteration. The displacement data can then be used to shift vertices to more closely match the original mesh 602 and create a reconstructed curve or mesh 608.

[0072] However, if Figure 6 As shown, the number of vertices of the reconstructed mesh 608 may be different from the number of vertices of the original mesh 602 because the number of vertices of the reconstructed mesh 608 increases as the number of subdivision iterations increases. The encoder typically selects the number of iterations to achieve a specific similarity level between the original mesh 602 and the reconstructed base mesh 608. However, the similarity level selected by the encoder may not be appropriate for various reasons.

[0073] For example, when a 3D scene consists of a certain number of objects, objects in the background of the 3D scene may not need to be reconstructed to achieve high similarity, but objects in the front of the 3D scene may need to be reconstructed to achieve high similarity. For example, in a six-DoF application, each individual user can freely choose the position and orientation of their view in 3D space. This means that the encoder cannot determine which objects may require high similarity and which may not. This determination must be made by the decoder. Furthermore, as the number of interpolation rounds increases, the number of points used for the base mesh increases, and the distortion between the original and base meshes decreases. In some cases, the number of vertices in the reconstructed mesh 608 may differ from the number of vertices in the original mesh 602, and often the number of vertices in the base mesh 604 may be greater than that of the original mesh 602. This can cause rendering issues when the renderer has limited capabilities, necessitating object simplification to meet the rendering system's limitations. Furthermore, applying a large number of iterations can waste resources, as the reconstructed mesh 608 may not be identical to the original mesh 602. Therefore, applying iterations may be more efficient when the distortion between the original mesh 602 and the reconstructed mesh 608 reaches a certain level.

[0074] Therefore, the present disclosure provides a method for using the number of vertices of the original mesh as a reference number for simplification while avoiding unnecessarily sacrificing the quality of the mesh. The present disclosure also provides the following technology: allowing the decoder to use the distortion information at each iteration to understand the similarity between the original mesh and the reconstructed base mesh at each iteration, so that the decoder can select the number of iterations to be applied to generate the base mesh according to the context of the application and the user.

[0075] Furthermore, since the original mesh 602 is simplified into a simplified version, namely the base mesh 604, in some cases, the base meshes 604 created for two different original meshes may be identical, even though the two original meshes are different. This may result in the generation, storage, and use of more than one identical base mesh, leading to a waste of resources. Therefore, the present disclosure provides a method for eliminating duplicate base meshes to further improve compression efficiency.

[0076] although Figure 6 An example process 600 for reconstructing a subgrid is shown, but may be used for Figure 6 Various changes may be made. Process 600 may be used in any other suitable process and is not limited to the specific process described above. In addition, although shown as a series of steps, Figure 6 The steps in can overlap, occur in parallel, or occur any number of times.

[0077] Figure 7 An example process 700 for reconstructing a simplified sub-mesh according to the present disclosure is shown. Figure 7 Process 700 is shown for illustration only. Figure 7 The scope of this disclosure is not limited to any particular implementation of the process for reconstructing the bitstream consistency of a simplified sub-grid. Figure 7 The process 700 can be described as using Figure 3 However, process 700 may be used with any other suitable system and any other suitable electronic device.

[0078] like Figure 7 As shown, similar to Figure 6 As described, process 700 includes downsampling the original mesh 702 to a simplified mesh or base mesh 704. However, to avoid Figure 6 The reconstructed mesh described has a larger number of vertices than the original mesh 702 and thus causes wasted resources or incompatibility issues with certain renderers, and a simplified subdivided mesh 706 may be created. The present disclosure provides methods to indicate the number of vertices of the original mesh and the similarity between the base mesh and the original mesh at each subdivision iteration.

[0079] To simplify the subdivided mesh, in embodiments, the total number of original vertices for each sub-mesh or patch may be signaled to the decoder. Additionally or alternatively, in embodiments, distortion information between the original mesh 702 and the reconstructed mesh after each subdivision iteration is signaled per sub-mesh or patch. In embodiments, the signaled distortion may be one or more of the following distortion metrics: a point cloud-based D1 metric, a point cloud-based D2 metric, a point cloud-based luminance / chrominance peak signal-to-noise ratio (PSNR) metric, a rendered image geometry- and luminance / chrominance PSNR metric, a perceptual-based distortion metric, or the like.

[0080] In an embodiment, the subdivision information is signaled by the encoder using the following syntax, and the decoder processes the syntax elements and recovers information about the number of vertices of the original mesh and the distortion between the original mesh and the reconstructed base mesh at each subdivision iteration.

[0081] number_of_vertices_of_original_submesh distortion_type subdivision_iteration_count for(i==0;i <subdivision_iteration_count;i++) distortion[i] Here, “number_of_vertices_of_original_submesh” indicates the number of vertices of the original submesh, “distortion_type” indicates the type of distortion metric used, “subdivision_iteration_count” indicates the number of subdivision iterations to be applied to generate the reconstructed mesh 708, and “distortion” indicates the difference between the original mesh 702 and the reconstructed base mesh 708 after performing the i-th subdivision iteration.

[0082] The available distortion types are listed in Table 1 below.

[0083] [Table 1] Distortion metric types

[0084] In an embodiment, the decoding process using this information is as follows: 1. Set layer_number to 0; 2. Decode the layer with layer_number; 3. Subdivide the layer with layer_number; 4. Get the number of vertices in the layer with layer_number+1 (based on the number of vertices in the subdivision layer with layer_number)? The difference in the number of vertices in the layer with layer_number; 5. Decode the displacement map values ​​in the video, where the number of displacement maps is equal to the difference value from step 4; 6. Increase layer_number by 1; 7. Check if layer_number is equal to the value of subdivision_iteration_count. If yes, go to step 8. If no, go to step 2; and 8. End.

[0085] Various standards for vertex mesh and dynamic mesh encoding have been proposed. The following documents are incorporated herein by reference in their entirety as if fully set forth herein: “V-Mesh Test Model v1,” ISO / IEC SC29 WG07 N00404, July 2022; “V-DMC Test Model v2 (TMM v2)”, ISO / IEC SC29 WG07 N00456, October 2022; “V-DMC Test Model v3 (TMM v3)”, ISO / IEC SC29 WG07 N00530, February 2023; “WD 1.0 of V-DMC,” ISO / IEC SC29 WG07, N00486, December 2023; “WD 2.0 of V-DMC,” ISO / IEC SC29 WG07 N00546, February 2023; “WD 3.0 of V-DMC,” ISO / IEC SC29 WG07 N00611, May 2023; “WD 4.0 of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 07 N00611, August 2023; and “WD 5.0 ​​of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 7 N00744, August 2023.

[0086] “WD of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 7 N00822, August 2023.

[0087] In an embodiment, a Supplemental Enhancement Information (SEI) message may be used. For example, the SEI message may be used to indicate the number of vertices of the original mesh 702 and define the degree of similarity between the base mesh 704 and the original mesh 702 at each subdivision iteration. For example, the syntax of the SEI message may indicate the total number of original vertices for each sub-mesh or patch. Distortion information between the original mesh and the reconstructed base mesh after each subdivision iteration may be signaled on a per-sub-mesh or patch basis. The following semantics may be defined: [Table 2] Subgrid distortion indication SEI payload syntax

[0088] Various embodiments may use the semantics described above. The SEI message indicates the number of vertices of the original sub-mesh and the similarity between the base mesh and the original mesh at each subdivision iteration, allowing the decoder to estimate the quality loss of the decoded sub-mesh. In some cases, when the estimated quality of the reconstructed sub-mesh is sufficient for the intended use, the reconstruction process can be stopped at a specific number of iterations by using the similarity information provided by the SEI message.

[0089] In the above semantics, “sdi_number_of_submesh_indicated” indicates the number of submesh similarity information signaled by this SEI message, “sdi_submesh_id_length_minus1 plus 1” specifies the number of bits used to represent the syntax element sdi_submesh_id[i], where “sdi_submesh_id[i]” indicates the identifier of the i-th submesh. The number of bits used to represent “sdi_submesh_id[i]” is “sdi_submesh_id_length_minus1+1”. In addition, “sdi_number_of_vertices_of_original_submesh[i]” indicates the number of vertices of the original submesh, and “sdi_subdivision_iteration_count[i]” indicates the number of subdivision iterations to be applied to generate the reconstructed base mesh. Optionally, “sdi_number_of_distortion_indicated_minus1[i] plus 1” indicates the number of distortions associated with the i-th submesh signaled by the SEI message. In addition, "sdi_distortion_metrics_type[i][j]" indicates the type of distortion metric for the j-th distortion of the i-th sub-mesh. Available example distortion metric types are listed in Table 1. In addition, "sdi_distortion[i][j][k]" indicates the j-th distortion metric between the original mesh and the reconstructed j-th sub-mesh after performing the k-th subdivision iteration.

[0090] like Figure 7 As shown, after creating the simplified subdivided mesh 706 , the displacement data may still be used to shift the simplified number of vertices and output a reconstructed mesh 708 having a number of vertices that matches or is similar to the original mesh 702 .

[0091] although Figure 7 An example process 700 for reconstructing a simplified sub-mesh is shown, but may be used for Figure 7 Various changes may be made. The example process 700 may be used in any other suitable process and is not limited to the specific process described above. In addition, although shown as a series of steps, Figure 7 The steps in can overlap, occur in parallel, or occur any number of times.

[0092] Figure 8 An example encoding method 800 that allows reconstruction of a simplified sub-grid according to the present disclosure is shown. For ease of explanation, Figure 8 The method 800 is described as using Figure 3However, the method 800 may be used with any other suitable system and any other suitable electronic device.

[0093] like Figure 8 As shown, in step 802, the electronic device 300 subdivides the submesh according to the subdivision iteration count to generate at least one subdivided submesh having vertex positions. In step 804, the electronic device 300 determines the number of vertices associated with the submesh to be used to simplify the vertex positions. In step 806, the electronic device 300 determines distortion information between the submesh and the base mesh at each subdivision iteration associated with the subdivision iteration count. As described in the present disclosure, the distortion information may include a distortion metric type, and the distortion metric type includes at least one of a point-to-point (D1) metric based on a point cloud, a point-to-plane (D2) metric based on a point cloud, a PSNR metric based on a point cloud, a PSNR metric based on a rendered image, and a distortion metric based on perception.

[0094] In step 808, the electronic device 300 uses the vertex positions of at least one subdivided sub-mesh to reconstruct at least a portion of the mesh frame. In step 810, the electronic device 300 creates a compressed bitstream that includes information about the sub-mesh, a subdivision iteration count, a number of vertices associated with the sub-mesh, and distortion information. As described in the present disclosure, in an embodiment, the electronic device 300 may signal the number of vertices associated with the sub-mesh and the distortion information as signaling elements. In an embodiment, the signaling elements are signaled as part of an SEI message. The output bitstream may be sent to an external device or to a memory on the electronic device 300.

[0095] although Figure 8 An example of an encoding method 800 that allows reconstruction of a simplified sub-grid is shown, but may be used for Figure 8 For example, although shown as a series of steps, Figure 8 The various steps in can overlap, occur in parallel, or occur any number of times. It will be understood that method 800 can be used with any number of sub-grids, and the sub-grids described with respect to method 800 are for illustration purposes only.

[0096] Figure 9 An example decoding method 900 for reconstructing a simplified sub-grid according to the present disclosure is shown. For ease of explanation, Figure 9 The method 900 is described as using Figure 3 However, the method 900 may be used with any other suitable system and any other suitable electronic device.

[0097] like Figure 9As shown, in step 902, the electronic device 300 receives a compressed bitstream including a base mesh sub-bitstream. In step 904, the electronic device 300 decodes a plurality of sub-meshes from the base mesh sub-bitstream. In step 906, the electronic device 300 subdivides a sub-mesh of the plurality of sub-meshes according to the subdivision iteration count to generate at least one subdivided sub-mesh. This includes, as shown in step 908, the electronic device determining a plurality of vertex positions for the at least one subdivided sub-mesh by using the number of vertices associated with the original sub-mesh and by using distortion information between the original sub-mesh and the base mesh at each subdivision iteration associated with the subdivision iteration count. In an embodiment, the number of vertices is determined based on a difference between the number of vertices associated with a next iteration of the subdivision iteration count and a current iteration of the subdivision iteration count.

[0098] In an embodiment, the number of vertices and distortion information associated with the original submesh are included in the compressed bitstream in association as a signaling element. In an embodiment, the signaling element further comprises one or more of a subdivision iteration count, a submesh identifier, and an amount of distortion associated with the submesh. In an embodiment, the signaling element is signaled as part of an SEI message. In an embodiment, the distortion information comprises a distortion metric type. In an embodiment, the distortion metric type comprises at least one of a point cloud-based point-to-point (D1) metric, a point cloud-based point-to-plane (D2) metric, a point cloud-based PSNR metric, a rendered image-based PSNR metric, and a perceptual-based distortion metric.

[0099] At step 910, the electronic device 300 reconstructs at least a portion of a mesh frame using vertex positions corresponding to at least one subdivided sub-mesh. In one embodiment, the electronic device 300 stops reconstructing the mesh frame at a specific iteration of the subdivision iteration count based on distortion information indicating that the reconstruction quality exceeds a threshold. At step 912, the electronic device 300 outputs decoded content (such as a 3D video including the reconstructed mesh frame). For example, the output decoded content may be transmitted to an external device or a storage device on the electronic device 300.

[0100] although Figure 9 An example of a decoding method 900 for reconstructing a simplified sub-grid is shown, but may be used for Figure 9 For example, although shown as a series of steps, Figure 9 The various steps in can overlap, occur in parallel, or occur any number of times. It will be understood that method 900 can be used with any number of sub-grids, and the sub-grids described with respect to method 900 are for illustration purposes only.

[0101] 10A and 10B illustrate an example process 1000 for creating a simplified mesh corresponding to an original mesh according to the present disclosure. The process 1000 shown in FIG. 10 is for illustration only. FIG. 10 does not limit the scope of the present disclosure to any particular embodiment of the process for creating a simplified mesh corresponding to an original mesh. For ease of explanation, the process 1000 of FIG. 10 may be described as using Figure 3 However, process 1000 may be used with any other suitable system and any other suitable electronic device.

[0102] As this article about Figure 6 As described above, since the original mesh is simplified into a simplified version, namely the base mesh, in some cases, the base meshes created for two different original meshes may be the same, even if the two original meshes are different. This may result in the generation, storage, and use of more than one identical base mesh, resulting in a waste of resources. Therefore, the present disclosure provides a method for eliminating duplicate base meshes to further improve compression efficiency.

[0103] For example, FIG10A shows two different original grids, a first original grid 1002 and a second original grid 1004. As shown in FIG10A , the shapes of the curves of the first original grid 1002 and the second original grid 1004 are different. However, as shown in FIG10B , even though the first original grid 1002 and the second original grid 1004 are different, the simplified grid 1006 created using the first original grid 1002 and the second original grid 1004 are identical. Therefore, the base grid for the two grids (simplified grid 1006) will be the same during encoding / decoding of the compressed bitstream, resulting in the transmission and processing of redundant data.

[0104] However, the present disclosure provides a method whereby, to avoid such redundancy and for efficiency, only one simplified mesh 1006 may be used to calculate the displacement fields for two different original meshes 1002, 1004, and when generating the reconstructed mesh, the decoder may use only one simplified mesh 1006. It should be understood that there may even be more than two original meshes having the same base mesh, and thus a single base mesh may even be used for three or more original sub-meshes.

[0105] There may be situations where the simplified mesh is a transformed version of the simplified mesh of another mesh. Therefore, embodiments of the present disclosure include transforming one base mesh to use it as a base mesh for computing a displacement field for another mesh.

[0106] Although FIG10 illustrates one example process 1000 for creating a simplified mesh corresponding to an original mesh, various modifications may be made to FIG10. Process 1000 may be used in any other suitable process and is not limited to the specific process described above. Furthermore, it will be understood that the simplified mesh and original mesh shown are for illustrative purposes only and may differ from those shown in FIG10.

[0107] Figure 11 An example encoding method 1100 for creating and signaling repeated base grid information according to the present disclosure is shown. For ease of explanation, Figure 11 The method 1100 is described as using Figure 3 However, the method 1100 may be used with any other suitable system and any other suitable electronic device.

[0108] like Figure 11 As shown, at step 1102, the electronic device 300 identifies that the first simplified mesh for the first original sub-mesh is the same as the second simplified mesh for the second original sub-mesh, such as described with respect to FIG. 10A and FIG. 10B . At step 1104, the electronic device 300 selects the first simplified mesh to be used for reconstructing the second sub-mesh. At step 1106, the electronic device 300 constructs a signaling element to instruct the decoder to copy the first simplified mesh when reconstructing the second sub-mesh, and to instruct the decoder how to manipulate the copied first simplified mesh. For example, in an embodiment, the identification of the base mesh to be copied and the operation to be applied to the mesh (which will be used as the base mesh for reconstructing another mesh) can be signaled by the encoder electronic device 300 using the following syntax: submesh_identifier translation_flag rotation_flag flipping_flag if(translation_flag== true){ translation_offset_u translation_offset_v translation_offset_w } if (rotation_flag == true) { anchor_point_u anchor_point_v anchor_point_w rotation_yaw rotation_pitch rotation_roll } if (flipping_flag == true) { flipping_surface_center_u flipping_surface_center_v flipping_surface_center_w flipping_surface_normal_u flipping_surface_normal_v flipping_surface_normal_w } As shown above, the syntax may include a "submesh_identifier," which indicates the submesh whose base mesh will be copied to create the base mesh for the current submesh. The syntax may also include a "translation_flag," a "rotation_flag," and / or a "flipping_flag," which indicate whether translation, rotation, and flipping, respectively, will be applied to the copied base mesh to apply the number of subdivision iterations to generate the base mesh for the current submesh. The syntax may also include a "translation_offset_u," a "translation_offset_v," and / or a "translation_offset_w," which indicate the amount of translation in the u, v, and w axes, respectively, to be applied to the copied base mesh to generate the base mesh for the current submesh. The syntax may also include an "anchor_point_u," an "anchor_point_v," and / or an "anchor_point_w," which indicate the location of a point in 3D space that will be used as an anchor point for the rotation. The syntax may also include "rotation_yaw", "rotation_pitch", and / or "rotation_roll", which indicate the amount of rotation of the yaw, pitch, and roll axes, respectively, to be applied to the copied base mesh to generate the base mesh for the current sub-mesh.

[0109] The syntax may also include "flipping_surface_center_u", "flipping_surface_center_v" and / or "flipping_surface_center_w", which indicate the center point of the surface to be applied to the copied base mesh to generate the flipping operation for the base mesh of the current sub-mesh. The syntax may also include "flipping_surface_normal_u", "flipping_surface_normal_v" and / or "flipping_surface_normal_w", which indicate the normal vector of the surface to be applied to the copied base mesh to generate the flipping operation for the base mesh of the current sub-mesh.

[0110] It should be understood that the syntax shown is for illustrative purposes and that other syntaxes may be used without departing from the scope of the present disclosure. At step 1108, the electronic device 300 encodes and outputs the bitstream including the signaling element. The output bitstream may be sent to an external device or a storage device on the electronic device 300.

[0111] although Figure 11 An example of an encoding method 1100 for creating and signaling repeated base grid information is shown, but may be used for Figure 11 For example, although shown as a series of steps, Figure 11 The steps in can overlap, occur in parallel, or occur any number of times.

[0112] Figure 12 An example decoding method 1200 for using a replicated base grid during grid reconstruction according to the present disclosure is shown. For ease of explanation, Figure 12 The method 1200 is described as using Figure 3 However, the method 1200 may be used with any other suitable system and any other suitable electronic device.

[0113] like Figure 12As shown, in step 1202, the electronic device 300 receives a compressed bitstream including a signaling element, which identifies a first sub-mesh, wherein the simplified mesh of the first sub-mesh will be copied when reconstructing the second sub-mesh. The signaling element also identifies how to manipulate the copied first simplified mesh. In step 1204, the electronic device 300 decodes at least a portion of the bitstream. In step 1206, the electronic device 300 recognizes that the current sub-mesh is a second sub-mesh for which a copy of the simplified mesh of the first sub-mesh should be used. For example, based on the signaling element, in an embodiment, a specific base mesh can be identified and its copy moved to another position in 3D space to serve as a base mesh for reconstructing another mesh. In an embodiment, a specific base mesh is identified and its copy is rotated relative to a single point in 3D space or its copy is flipped relative to a specific surface defined in 3D space.

[0114] The electronic device 300 may process the syntax elements of the signaling elements to create a replicated base grid, such as processing information about the subgrid identifier and information about the Figure 11 The electronic device 300 may further comprise a syntax element for describing translation, rotation and / or flip information. In step 1208, the electronic device 300 may copy the simplified mesh and manipulate the simplified mesh based on the signaling element. In step 1210, the electronic device 300 may output the decoded and reconstructed content.

[0115] although Figure 12 One example of a decoding method 1200 for using a replicated base mesh during mesh reconstruction is shown, but may be used for Figure 12 For example, although shown as a series of steps, Figure 12 The steps in can overlap, occur in parallel, or occur any number of times.

[0116] In an embodiment, an apparatus comprises: a communication interface configured to receive a compressed bitstream comprising a base grid sub-bitstream; and a processor operably coupled to the communication interface.

[0117] The processor is configured to decode a plurality of sub-grids from a base-grid sub-bitstream.

[0118] The processor is configured to: subdivide a submesh from the plurality of submeshes according to a subdivision iteration count to generate at least one subdivided submesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided submesh by using a number of vertices associated with an original submesh and by using distortion information between the original submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count.

[0119] The processor is configured to reconstruct at least a portion of the mesh frame using vertex positions corresponding to at least one subdivided sub-mesh.

[0120] The number of vertices and distortion information associated with the original sub-mesh are included as signaling elements in the compressed bitstream.

[0121] The signaling element also includes one or more of a subdivision iteration count, a sub-mesh identifier, and a number of distortions associated with the sub-mesh.

[0122] The signaling elements are signaled as part of a Supplemental Enhancement Information (SEI) message.

[0123] The distortion information includes the distortion metric type.

[0124] The distortion metric type includes at least one of a point cloud based point-to-point (D1) metric, a point cloud based point-to-plane (D2) metric, a point cloud based peak signal-to-noise ratio (PSNR) metric, a rendered image based PSNR metric, and a perception-based distortion metric.

[0125] The processor is further configured to stop reconstruction of the mesh frame at a particular iteration of the subdivision iteration count based on the distortion information indicating that the reconstruction quality is above a threshold.

[0126] The number of vertices is determined based on a difference between the number of vertices associated with the next iteration of the subdivision iteration count and the current iteration of the subdivision iteration count.

[0127] In an embodiment, a method comprises receiving a compressed bitstream comprising a base grid sub-bitstream.

[0128] The method includes decoding a plurality of sub-grids from a base-grid sub-bitstream.

[0129] The method includes subdividing a submesh from among a plurality of submeshes according to a subdivision iteration count to generate at least one subdivided submesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided submesh by using a number of vertices associated with an original submesh and by using distortion information between the original submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count.

[0130] The method includes reconstructing at least a portion of a mesh frame using vertex positions corresponding to at least one subdivided sub-mesh.

[0131] The number of vertices and distortion information associated with the original sub-mesh are included as signaling elements in the compressed bitstream.

[0132] The signaling element also includes one or more of a subdivision iteration count, a sub-mesh identifier, and a number of distortions associated with the sub-mesh.

[0133] The signaling elements are signaled as part of a Supplemental Enhancement Information (SEI) message.

[0134] The distortion information includes the distortion metric type.

[0135] The distortion metric type includes at least one of a point cloud based point-to-point (D1) metric, a point cloud based point-to-plane (D2) metric, a point cloud based peak signal-to-noise ratio (PSNR) metric, a rendered image based PSNR metric, and a perception-based distortion metric.

[0136] The method includes stopping reconstruction of the mesh frame at a specific iteration of a subdivision iteration count based on distortion information indicating that the reconstruction quality is above a threshold.

[0137] The number of vertices is determined based on a difference between the number of vertices associated with the next iteration of the subdivision iteration count and the current iteration of the subdivision iteration count.

[0138] In an embodiment, a device includes: a communication interface; and a processor operably coupled to the communication interface.

[0139] The processor is configured to subdivide the sub-mesh according to the subdivision iteration count to generate at least one subdivided sub-mesh having vertex positions.

[0140] The processor is configured to determine a number of vertices associated with the sub-mesh to be used to simplify vertex positions.

[0141] The processor is configured to determine distortion information between the sub-mesh and the base mesh at each subdivision iteration associated with the subdivision iteration count.

[0142] The processor is configured to reconstruct at least a portion of the mesh frame using vertex positions of at least one subdivided sub-mesh.

[0143] The processor is configured to create a compressed bitstream comprising information about the sub-meshes, a subdivision iteration count, a number of vertices associated with the sub-meshes, and distortion information.

[0144] The processor is configured to signal the number of vertices and distortion information associated with the sub-mesh as signaling elements.

[0145] The signaling elements are signaled as part of a Supplemental Enhancement Information (SEI) message.

[0146] The distortion information includes a distortion metric type, the distortion metric type including at least one of a point cloud-based point-to-point (D1) metric, a point cloud-based point-to-plane (D2) metric, a point cloud-based peak signal-to-noise ratio (PSNR) metric, a rendered image-based PSNR metric, and a perception-based distortion metric.

[0147] In an embodiment, a method includes subdividing a submesh according to a subdivision iteration count to generate at least one subdivided submesh having vertex positions.

[0148] The method includes determining a number of vertices associated with a sub-mesh to be used to simplify vertex positions.

[0149] The method includes determining distortion information between the sub-mesh and the base mesh at each subdivision iteration associated with a subdivision iteration count.

[0150] The method includes reconstructing at least a portion of a mesh frame using vertex positions of at least one subdivided sub-mesh.

[0151] The method includes creating a compressed bitstream including information about a sub-mesh, a subdivision iteration count, a number of vertices associated with the sub-mesh, and distortion information.

[0152] The method includes signaling a number of vertices and distortion information associated with a sub-mesh as signaling elements.

[0153] The signaling elements are signaled as part of a Supplemental Enhancement Information (SEI) message.

[0154] The distortion information includes a distortion metric type, the distortion metric type including at least one of a point cloud-based point-to-point (D1) metric, a point cloud-based point-to-plane (D2) metric, a point cloud-based peak signal-to-noise ratio (PSNR) metric, a rendered image-based PSNR metric, and a perception-based distortion metric.

[0155] Although the present disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. The present disclosure is intended to include such changes and modifications as fall within the scope of the appended claims. Nothing in this application should be interpreted as implying that any particular element, step, or function is essential to be included within the scope of the claims. The scope of a patented subject matter is defined by the claims.

Claims

1. A device comprising: a communication interface configured to receive a compressed bitstream comprising a base grid sub-bitstream; as well as a processor operatively coupled to the communication interface, the processor configured to: decoding a plurality of sub-grids from the base-grid sub-bitstream; subdividing a submesh of the plurality of submeshes according to a subdivision iteration count to generate at least one subdivided submesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided submesh by using a number of vertices associated with an original submesh and by using distortion information between the original submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count; as well as At least a portion of a mesh frame is reconstructed using the vertex positions corresponding to the at least one subdivided sub-mesh.

2. The device according to claim 1, wherein The number of vertices associated with the original sub-mesh and the distortion information are associatedly included in the compressed bitstream as signaling elements.

3. The apparatus of claim 2, wherein: The signaling element also includes one or more of the following: said subdivision iteration count; subgrid identifier; and The amount of distortion associated with this sub-mesh.

4. The apparatus of claim 3, wherein: The signaling element is signaled as part of a Supplemental Enhancement Information, SEI, message.

5. The apparatus of claim 1, wherein: The distortion information includes a distortion metric type.

6. The apparatus of claim 5, wherein: The distortion metric type includes at least one of the following: Point-to-point D1 measurement based on point cloud; Point-to-plane D2 metric based on point cloud; Peak signal-to-noise ratio (PSNR) metric based on point cloud; PSNR metric based on rendered images; as well as Perception-based distortion metrics.

7. The apparatus of claim 1, wherein: The processor is further configured to stop the reconstruction of the mesh frame at a specific iteration of the subdivision iteration count based on the distortion information indicating that the reconstruction quality is above a threshold.

8. The apparatus of claim 1, wherein: The number of vertices is determined based on a difference between a number of vertices associated with a next iteration of the subdivision iteration count and a current iteration of the subdivision iteration count.

9. A method comprising: receiving a compressed bitstream comprising a base grid sub-bitstream; decoding a plurality of sub-grids from the base-grid sub-bitstream; subdividing a submesh of the plurality of submeshes according to a subdivision iteration count to generate at least one subdivided submesh, wherein the process includes determining a plurality of vertex positions for the at least one subdivided submesh by using a number of vertices associated with an original submesh and by using distortion information between the original submesh and a base mesh at each subdivision iteration associated with the subdivision iteration count; as well as At least a portion of a mesh frame is reconstructed using the vertex positions corresponding to the at least one subdivided sub-mesh.

10. The method of claim 9, wherein: The number of vertices associated with the original sub-mesh and the distortion information are associatedly included in the compressed bitstream as signaling elements.

11. The method according to claim 10, wherein: The signaling element also includes one or more of the following: said subdivision iteration count; subgrid identifier; and The amount of distortion associated with this sub-mesh.

12. The method of claim 11, wherein: The signaling element is signaled as part of a Supplemental Enhancement Information, SEI, message.

13. The method of claim 9, wherein: The distortion information includes a distortion metric type.

14. The method of claim 13, wherein: The distortion metric type includes at least one of the following: Point-to-point D1 measurement based on point cloud; Point-to-plane D2 metric based on point cloud; Peak signal-to-noise ratio (PSNR) metric based on point cloud; PSNR metric based on rendered images; as well as Perception-based distortion metrics.

15. The method of claim 9, further comprising: Based on the distortion information indicating that the reconstruction quality is above a threshold, the reconstruction of the mesh frame is stopped at a specific iteration of the subdivision iteration count.