Partial decoding and reconstruction of subgrids
By partially decoding and reconstruction of the subgrid of 3D multimedia data and processing independent of other subgrids, the problem of inefficient decoding in the prior art is solved, efficient 3D object decoding and reconstruction is realized, and real-time interaction and dynamic viewing is supported.
Patent Information
- Application Number
- CN202480007048.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-01-08
- Publication Date
- 2025-08-05
AI Technical Summary
When encoding and decoding 3D multimedia data, especially in the partial decoding and reconstruction of sub-grids, there is a problem that vertex and attribute data cannot be decoded and reconstruction independently of other sub-grids, resulting in real-time and inefficient.
The partial decoding and reconstruction method of the subgrid is adopted. By encoding the geometric data and attribute data into the displacement subbit stream and attribute subbit stream respectively, and decoding and reconstruction is performed independently of the data of other subgrids during the decoding process, the bitstream consistency limitation is used to ensure partial decoding and reconstruction of vertices and attribute data.
It realizes efficient partial decoding and reconstruction of 3D objects in multimedia devices, supports real-time interaction and dynamic viewing, and improves the processing capabilities and user experience of the device.
Smart Images

Figure CN120435869A_ABST
Abstract
Description
Technical Field
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 438,198 filed on January 10, 2023, U.S. Provisional Patent Application No. 63 / 438,417 filed on January 11, 2023, U.S. Provisional Patent Application No. 63 / 439,456 filed on January 17, 2023, and U.S. Patent Application No. 18 / 399,529 filed on December 28, 2023, the entire contents of the foregoing applications are incorporated herein by reference.
[0002] The present disclosure relates generally to multimedia devices and processes. More particularly, the present disclosure relates to partial decoding and reconstruction of sub-grids. Background Art
[0003] Thanks to the ready availability of powerful handheld devices such as smartphones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are emerging as new ways to experience immersive content. While 360° video enables consumers to experience an immersive, "real-life," "there" experience by capturing a 360° outside-in view of the world, 3D volumetric video provides a full six degrees of freedom (DoF) experience, allowing users to immerse themselves in and move around the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Displays and navigation sensors can track the user's head movements in real time to determine the area of the 360° video or volumetric content that the user wants to view or interact with. Intrinsically 3D multimedia data, such as point clouds or 3D polygon meshes, can be used in immersive environments. This data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream. Summary of the Invention
[0004] Solution to the problem
[0005] The present disclosure provides partial decoding and reconstruction of sub-grids.
[0006] In an embodiment, an apparatus may include a communication interface configured to receive a compressed bitstream having sub-bitstreams, the sub-bitstreams including a base mesh sub-bitstream, a displacement sub-bitstream, and an attribute sub-bitstream. The apparatus may include a processor operably coupled to the communication interface. The processor may be configured to decode at least a portion of the compressed bitstream. The processor may be configured to decode a plurality of sub-meshes from the base mesh sub-bitstream. The processor may decode geometry data from the displacement sub-bitstream. The processor may decode attribute data from the attribute sub-bitstream. The processor may also be configured to subdivide a sub-mesh from the plurality of sub-meshes to generate a subdivided sub-mesh. The processor may also be configured to reconstruct vertex positions of the subdivided sub-mesh using the decoded geometry data and attributes of the subdivided sub-mesh using the decoded attribute data, independently of decoded data corresponding to one or more other sub-meshes. The processor may also be configured to reconstruct at least a portion of a mesh frame using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-mesh.
[0007] In an embodiment, a method may include receiving a compressed bitstream having sub-bitstreams, the sub-bitstreams including a base mesh sub-bitstream, a displacement sub-bitstream, and an attribute sub-bitstream. The method may include decoding at least a portion of the compressed bitstream. The method may include decoding a plurality of sub-meshes from the base mesh sub-bitstream. The method may include decoding geometry data from the displacement sub-bitstream. The method may include decoding attribute data from the attribute sub-bitstream. The method may include subdividing a sub-mesh of the plurality of sub-meshes to generate a subdivided sub-mesh. The method may include reconstructing vertex positions of the subdivided sub-mesh using the decoded geometry data, and reconstructing attributes of the subdivided sub-mesh using the decoded attribute data, independently of decoded data corresponding to one or more other sub-meshes. The method may include reconstructing at least a portion of a mesh frame using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-mesh.
[0008] In an embodiment, an apparatus may include a communication interface and a processor operably coupled to the communication interface. The processor may be configured to encode geometry data and attribute data associated with a separate sub-mesh into a displacement sub-bitstream and an attribute sub-bitstream, respectively. The geometry data and attribute data associated with the separate sub-mesh may be separable from data corresponding to one or more other sub-meshes in the displacement sub-bitstream and the attribute sub-bitstream during decoding. The processor may also be configured to combine the displacement sub-bitstream and the attribute sub-bitstream into a compressed bitstream.
[0009] In an embodiment, a method may include encoding geometry data and attribute data associated with a separate sub-mesh into a displacement sub-bitstream and an attribute sub-bitstream, respectively. The geometry data and attribute data associated with the separate sub-mesh may be separable from data corresponding to one or more other sub-meshes in the displacement sub-bitstream and the attribute sub-bitstream during decoding. The method may include combining the displacement sub-bitstream and the attribute sub-bitstream into a compressed bitstream.
[0010] Other technical features will be clear to those skilled in the art from the following drawings, description and claims.
[0011] Before proceeding to the following detailed description, it may be beneficial to set forth the definitions of certain words and phrases used throughout this patent document. The term "coupling" and its derivatives refer to any direct or indirect communication between two or more elements, regardless of whether these elements are in physical contact with each other. The terms "send," "receive," and "communicate" and their derivatives encompass both direct and indirect communication. The terms "include" and "comprises" and their derivatives refer to, including but not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with..." and its derivatives refer to including, being included within, interconnected with, containing, being contained within, being connected to or connected with, being coupled to or coupled with, being able to communicate with, collaborating with, interweaving, juxtaposing, being close to, being bound to or bound with, having, having the property of, having a relationship to or with, etc. The term "controller" refers to any device, system, or part thereof that controls at least one operation. Such a controller can be implemented in hardware or a combination of hardware and software and / or firmware. The functions associated with any particular controller can be centralized or distributed, whether locally or remotely. The phrase "at least one of" when used with a list of items means that different combinations of one or more of the listed items can be used, and that only one of the items in the list may be required. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A, B, and C.
[0012] Furthermore, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed of computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, related data, or portions thereof, adapted to be implemented in suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as read-only memory (ROM), random-access memory (RAM), hard drive, compact disk (CD), digital video disk (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit transitory electrical or other signals. Non-transitory computer-readable media includes media in which data can be permanently stored as well as media in which data can be stored and later overwritten, such as rewritable optical disks or erasable memory devices.
[0013] Definitions for certain other words and phrases are provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, in which like reference numerals represent like parts:
[0015] Figure 1 An example communication system according to the present disclosure is shown;
[0016] Figure 2 An example electronic device according to the present disclosure is shown;
[0017] Figure 3 An example electronic device according to the present disclosure is shown;
[0018] Figure 4 An example intra-frame encoding process according to the present disclosure is shown;
[0019] Figure 5 An example trellis frame decoding process according to the present disclosure is shown;
[0020] Figure 6 An example process for bitstream conformance of partial reconstruction and / or decoding of a sub-mesh according to the present disclosure is shown;
[0021] Figure 7 An example process for defining coordinates associated with a sub-bitstream according to the present disclosure is shown;
[0022] Figure 8 An example process for defining coordinates associated with sub-bitstreams in a non-overlapping manner according to the present disclosure is shown;
[0023] Figure 9 An example encoding method for partial reconstruction and / or decoding of a sub-mesh according to the present disclosure is shown; and
[0024] Figure 10 An example decoding method for partial reconstruction and / or decoding of a sub-mesh according to the present disclosure is shown. DETAILED DESCRIPTION
[0025] Described below Figures 1 to 10 The various embodiments used to describe the principles of the present disclosure are illustrative only and should not be construed in any way to limit the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure can be implemented in any type of appropriately arranged device or system.
[0026] As mentioned above, thanks to the readily available availability of powerful handheld devices such as smartphones, three-hundred-sixty-degree (360°) video and three-dimensional (3D) volumetric video are emerging as new ways to experience immersive content. While 360° video enables consumers to experience an immersive, "real-life," "being there" experience by capturing a 360° outside-in view of the world, 3D volumetric video provides a full six-degree-of-freedom (DoF) experience, allowing users to immerse themselves in and move around within the content. Users can interactively change their viewpoint and dynamically view any part of the captured scene or object they desire. Displays and navigation sensors can track the user's head movements in real time to determine the area of the 360° video or volumetric content that the user wants to view or interact with. Intrinsically 3D multimedia data, such as point clouds or 3D polygon meshes, can be used in immersive environments. This data can be stored in a video format and encoded and compressed for transmission to other devices as a bitstream.
[0027] A point cloud is a collection of 3D points and properties representing the surface or volume of an object, such as color, normal direction, reflectivity, point size, etc. Point clouds are common in various applications, such as gaming, 3D mapping, visualization, medical applications, augmented reality, virtual reality, autonomous driving, multi-view playback, and six degrees of freedom (DoF) immersive media, to name a few. If uncompressed, point clouds typically require a large amount of bandwidth for transmission. Due to the large bitrate requirements, point clouds are typically compressed before transmission. Compressing 3D objects, such as point clouds, typically requires specialized hardware. To avoid using specialized hardware to compress 3D point clouds, 3D point clouds can be transformed into traditional two-dimensional (2D) frames, which can be compressed and later reconstructed and viewed by the user.
[0028] Polygonal 3D meshes, particularly triangle meshes, are another popular format for representing 3D objects. A mesh typically consists of a set of vertices, edges, and faces that represent the surface of a 3D object. A triangle mesh is a simple polygonal mesh where the faces are simple triangles that cover the surface of the 3D object. Typically, there can be one or more attributes associated with a mesh. In a scene, one or more attributes can be associated with each vertex in the mesh. For example, a texture attribute (RGB) can be associated with each vertex. In a scene, each vertex can be associated with a pair of coordinates (u, v). The (u, v) coordinates can refer to a location in a texture map associated with the mesh. For example, the (u, v) coordinates can refer to the row and column indices, respectively, in a texture map. A mesh can be thought of as a point cloud with additional connectivity information.
[0029] Point clouds or meshes can be dynamic, meaning they change over time. In these cases, a point cloud or mesh at a specific moment in time can be referred to as a point cloud frame or mesh frame, respectively. Because point clouds and meshes contain large amounts of data, they require compression for efficient storage and transmission. This is especially true for dynamic point clouds and meshes, which can contain 60 or more frames per second.
[0030] As part of the encoding process, a base mesh can be encoded using an existing mesh codec, and a reconstructed base mesh can be constructed from the encoded original mesh. The reconstructed base mesh can then be subdivided into one or more subdivided meshes, and a displacement field is created for each subdivided mesh. For example, if the reconstructed base mesh includes triangles covering the surface of a 3D object, the triangles are subdivided according to multiple subdivision levels to create a first subdivided mesh in which each triangle of the reconstructed base mesh is subdivided into four triangles, a second subdivided mesh in which each triangle of the reconstructed base mesh is subdivided into sixteen triangles, and so on, depending on how many subdivision levels are applied. Each displacement field represents the difference between the vertex positions of the original mesh and the subdivided mesh associated with the displacement field. Each displacement field is wavelet transformed to create a level of detail (LOD) signal that is encoded as part of the compressed bitstream. During decoding, the displacement of each displacement field is added to its associated subdivided mesh to reconstruct a version of the original mesh.
[0031] Typically, mesh encoding and decoding operations are strongly sequential. For base meshes with a large number of vertices and high frame rates, mesh codecs may have difficulty achieving real-time encoding and decoding. To alleviate this problem, submeshes are used. The base mesh can be divided into multiple submeshes. Submeshes may not be mutually exclusive, i.e., some vertices and triangles may be common to different submeshes. Submeshes can be encoded and decoded without using any information from other submeshes. This allows multiple instances of the mesh codec to operate in parallel on different submeshes. This also supports the ability to perform partial decoding of a mesh by decoding only some of the submeshes present in the bitstream. Each decoded submesh can undergo subdivision, and the decoded displacement field is then used to refine the positions of the subdivided points belonging to that submesh.
[0032] However, the creation of sub-meshes is not sufficient to ensure the functionality of partial decoding and reconstruction of the full-resolution mesh. For example, if the displacement video and the attribute video are not tiled, partial decoding is not possible for those videos. Furthermore, a wavelet transform is typically applied to the displacement field on the encoder side, and a corresponding inverse wavelet transform is applied on the decoder side. If the forward wavelet transform uses vertices corresponding to different sub-meshes, a partial reconstruction of the vertices corresponding to a sub-mesh cannot be performed on the decoder side unless the relevant wavelet coefficients corresponding to other sub-meshes required for the inverse wavelet transform have been decoded. In some cases, in addition, an inverse wavelet transform can be performed on the relevant wavelet coefficients corresponding to other sub-meshes.
[0033] The present disclosure provides for partial reconstruction and / or partial decoding of a bitstream when multiple sub-meshes are used. Various embodiments of the present disclosure include bitstream consistency constraints on the bitstream and syntax elements to assist in the partial decoding and / or reconstruction of vertices and corresponding attributes of a sub-mesh. As further described in the present disclosure, in embodiments, bitstream consistency can be imposed, which requires that the vertices, vertex connectivity, and attribute data corresponding to a sub-mesh can be reconstructed independently of the decoded data corresponding to other sub-meshes. As further described in the present disclosure, in embodiments, bitstream consistency can be imposed, which requires that the vertices, connectivity, and attribute data corresponding to a sub-mesh can be decoded and reconstructed independently of the other sub-meshes.
[0034] In some examples of the present disclosure, the term "submesh" may refer to a partition of a base mesh. In some examples, in the present disclosure, a "submesh" may refer to geometric data reconstructed after the submesh is subdivided and displacements are added.
[0035] According to embodiments of the present disclosure, geometric data may include structural information of a 3D object. The structural information may include at least one of vertex positions, texture vertex positions (UV coordinates), or connectivity. Vertex positions may include three-dimensional coordinate information for a vertex that constitutes each triangle of a mesh. Texture vertex positions may include two-dimensional coordinates corresponding to the vertex on a texture map. Connectivity may include indices of the three vertices of a triangle that constitutes the mesh.
[0036] According to an embodiment of the present disclosure, the attribute data may include color-related information about the surface of the 3D object. In one example, the attribute data may include an image that is a two-dimensional texture map having an RGB color space.
[0037] Figure 1 An example communication system 100 according to the present disclosure is shown. Figure 1 The embodiment of the communication system 100 shown is for illustration only. The communication system 100 may be used in accordance with embodiments of the present disclosure without departing from the scope of the present disclosure.
[0038] like Figure 1 As shown, the communication system 100 includes a network 102 that facilitates communication between various components in the communication system 100. For example, the network 102 can transmit IP packets, frame relay frames, asynchronous transfer mode (ATM) cells, or other information between network addresses. The network 102 includes all or part of one or more local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), a global network such as the Internet, or any other communication system or systems at one or more locations.
[0039] In the example, network 102 facilitates communication between server 104 and various client devices 106-116. Client devices 106-116 may be, for example, smartphones, tablets, laptops, personal computers, TVs, interactive displays, wearable devices, head-mounted displays (HMDs), and the like. Server 104 may represent one or more servers. Each server 104 includes any suitable computing or processing device that can provide computing services to one or more client devices, such as client devices 106-116. Each server 104 may, for example, include one or more processing devices, one or more memories for storing instructions and data, and one or more network interfaces for facilitating communication over network 102. As described in more detail below, server 104 may transmit a compressed bitstream representing a point cloud or mesh to one or more display devices, such as client devices 106-116. In embodiments, each server 104 may include an encoder. In embodiments, server 104 may perform partial decoding and / or partial reconstruction of sub-meshes as described in the present disclosure.
[0040] Each client device 106-116 represents any suitable computing or processing device that interacts with at least one server (such as server 104) or other computing device via network 102. Client devices 106-116 include desktop computers 106, mobile phones or mobile devices 108 (such as smartphones), PDAs 110, laptop computers 112, tablet computers 114, and head-mounted display (HMD) 116. However, any other or additional client devices may be used in communication system 100. Smartphones represent one type of mobile device 108: handheld devices with a mobile operating system and integrated mobile broadband cellular network connectivity for voice, short message service (SMS), and internet data communications. HMD 116 can display a 360° scene including one or more dynamic or static 3D point clouds. In embodiments, any of client devices 106-116 may include an encoder, a decoder, or both. For example, mobile device 108 may record 3D volumetric video and then encode the video so that it can be transmitted to one of client devices 106-116. In an example, laptop computer 112 may be used to generate a 3D point cloud or mesh, which is then encoded and sent to one of client devices 106 - 116 .
[0041] In the example, some client devices 108-116 communicate indirectly with the network 102. For example, the mobile device 108 and the PDA 110 communicate via one or more base stations 118, such as cellular base stations or eNodeBs (eNBs). Additionally, the laptop 112, the tablet 114, and the HMD 116 communicate via one or more wireless access points 120, such as IEEE 802.11 wireless access points. Note that these are for illustration purposes only, and each client device 106-116 can communicate directly with the network 102 or indirectly with the network 102 via any suitable intermediary device or network. In an embodiment, the server 104 or any client device 106-116 can be configured to compress a point cloud or mesh, generate a bitstream representing the point cloud or mesh, and send the bitstream to another client device, such as any client device 106-116.
[0042] In an embodiment, any of the client devices 106-114 securely and efficiently transmits information to another device, such as, for example, the server 104. Furthermore, any of the client devices 106-116 can trigger information transmission between itself and the server 104. Any of the client devices 106-114, when attached to a head-mounted device via a mount, can function as a VR display and act similarly to the HMD 116. For example, the mobile device 108, when attached to a mount system and worn over the user's eyes, can act similarly to the HMD 116. The mobile device 108 (or any of the other client devices 106-116) can trigger information transmission between itself and the server 104.
[0043] In an embodiment, any of the client devices 106 to 116 or the server 104 may create a 3D point cloud or mesh, compress a 3D point cloud or mesh, transmit a 3D point cloud or mesh, receive a 3D point cloud or mesh, decode a 3D point cloud or mesh, render a 3D point cloud or mesh, or a combination thereof. For example, the server 104 may compress the 3D point cloud or mesh to generate a bitstream and then transmit the bitstream to one or more of the client devices 106-116. As an example, one of the client devices 106-116 may compress the 3D point cloud or mesh to generate a bitstream and then transmit the bitstream to another of the client devices 106-116 or the server 104. In accordance with the present disclosure, the server 104 and / or the client devices 106-116 may perform partial decoding and / or partial reconstruction of a sub-mesh as described in the present disclosure.
[0044] although Figure 1 One example of a communication system 100 is shown, but may be Figure 1Various changes may be made. For example, the communication system 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of this disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.
[0045] Figure 2 and Figure 3 An example electronic device according to the present disclosure is shown. In particular, Figure 2 An example server 200 is shown and may represent Figure 1 The server 200 may represent one or more encoders, decoders, local servers, remote servers, cluster computers, and components that act as a single pool of seamless resources, cloud-based servers, etc. The server 200 may be composed of Figure 1 One or more of the client devices 106-116 or another server accesses.
[0046] like Figure 2 As shown, server 200 may represent one or more local servers, one or more compression servers, or one or more encoding servers, such as encoders. In an embodiment, the encoder may perform decoding. Figure 2 As shown, server 200 includes a bus system 205 that supports communications between at least one processing device (such as processor 210 ), at least one storage device 215 , at least one communication interface 220 , and at least one input / output (I / O) unit 225 .
[0047] The processor 210 executes instructions that may be stored in the memory 230. The processor 210 may include any suitable number and type of processors or other devices in any suitable arrangement. Example types of processors 210 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application specific integrated circuits, and discrete circuits.
[0048] In an embodiment, the processor 210 may encode a 3D point cloud or mesh stored in the storage device 215. In an embodiment, encoding the 3D point cloud also decodes the 3D point cloud or mesh to ensure that when the point cloud or mesh is reconstructed, the reconstructed 3D point cloud or mesh matches the 3D point cloud or mesh before encoding. In an embodiment, the processor 210 may perform partial decoding and / or partial reconstruction of a sub-mesh as described in the present disclosure.
[0049] Memory 230 and persistent storage 235 are examples of storage devices 215, which represent any structure capable of storing and facilitating retrieval of information, such as temporary or persistent data, program code, or other suitable information. Memory 230 may represent random access memory or any other suitable volatile or non-volatile storage device. For example, the instructions stored in memory 230 may include instructions for decomposing a point cloud into patches, instructions for packing patches onto 2D frames, instructions for compressing 2D frames, and instructions for encoding 2D frames in a particular order to generate a bitstream. The instructions stored in memory 230 may also include instructions for viewing the image in a manner similar to that described in the preceding text, such as through a VR headset (such as a Figure 1 The persistent storage 235 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk.
[0050] The communication interface 220 supports communication with other systems or devices. For example, the communication interface 220 may include a communication interface that facilitates communication with other systems or devices. Figure 1 The communication interface 220 may include a network interface card or wireless transceiver for communication with the network 102. The communication interface 220 may support communication via any suitable physical or wireless communication link. For example, the communication interface 220 may send a bitstream containing a 3D point cloud to another device, such as one of the client devices 106-116.
[0051] The I / O unit 225 allows for the input and output of data. For example, the I / O unit 225 can provide a connection for user input via a keyboard, mouse, keypad, touch screen, or other suitable input device. The I / O unit 225 can also send output to a display, printer, or other suitable output device. However, it is noted that the I / O unit 225 can be omitted, such as when I / O interaction with the server 200 occurs via a network connection.
[0052] Note that although Figure 2 Described as indicating Figure 1 The server 104 of FIG. 106 may be configured as a server, but the same or similar architecture may be used in one or more of the various client devices 106-116. For example, a desktop computer 106 or a laptop computer 112 may have a server that is configured as a server. Figure 2 The structures shown are the same or similar structures.
[0053] Figure 3 An example electronic device 300 is shown and may represent Figure 1 The electronic device 300 may be a mobile communication device such as, for example, a mobile station, a subscriber station, a wireless terminal, a desktop computer (similar to a Figure 1 Desktop computer 106), portable electronic device (similar to Figure 1 In an embodiment, Figure 1 One or more of the client devices 106-116 may include the same or similar configuration as the electronic device 300. In an embodiment, the electronic device 300 is an encoder, a decoder, or both. For example, the electronic device 300 may be used with data transmission, image or video compression, image or video decompression, encoding, decoding, and media rendering applications.
[0054] like Figure 3 As shown, electronic device 300 includes antenna 305, radio frequency (RF) transceiver 310, transmit (TX) processing circuit 315, microphone 320, and receive (RX) processing circuit 325. RF transceiver 310 may include, for example, an RF transceiver, a Bluetooth transceiver, a Wi-Fi transceiver, a Zigbee transceiver, an infrared transceiver, and various other wireless communication signals. Electronic device 300 also includes speaker 330, processor 340, input / output (I / O) interface (IF) 345, input 350, display 355, memory 360, and sensor 365. Memory 360 includes operating system (OS) 361 and one or more applications 362.
[0055] The RF transceiver 310 receives incoming RF signals from an antenna 305 transmitted from an access point (such as a base station, a WI-FI router, or a Bluetooth device) or other device of a network 102 (such as WI-FI, Bluetooth, cellular, 5G, LTE, LTE-A, WiMAX, or any other type of wireless network). The RF transceiver 310 downconverts the incoming RF signal to generate an intermediate frequency (IF) or baseband signal. The IF or baseband signal is sent to the RX processing circuitry 325, which generates a processed baseband signal by filtering, decoding, and / or digitizing the baseband or IF signal. The RX processing circuitry 325 sends the processed baseband signal to the speaker 330 (such as for voice data) or to the processor 340 for further processing (such as for web browsing data).
[0056] The TX processing circuitry 315 receives analog or digital voice data from the microphone 320 or other outgoing baseband data from the processor 340. The outgoing baseband data may include web data, email, or interactive video game data. The TX processing circuitry 315 encodes, multiplexes, and / or digitizes the outgoing baseband data to generate a processed baseband or intermediate frequency signal. The RF transceiver 310 receives the outgoing processed baseband or intermediate frequency signal from the TX processing circuitry 315 and up-converts the baseband or intermediate frequency signal into an RF signal that is transmitted via the antenna 305.
[0057] Processor 340 may include one or more processors or other processing devices. Processor 340 may execute instructions stored in memory 360 (such as OS 361) to control the overall operation of electronic device 300. For example, processor 340 may control the reception of forward channel signals and the transmission of reverse channel signals by RF transceiver 310, RX processing circuit 325, and TX processing circuit 315 according to well-known principles. Processor 340 may include any suitable number and type of processors or any other device in any suitable arrangement. For example, in an embodiment, processor 340 includes at least one microprocessor or microcontroller. Example types of processor 340 include microprocessors, microcontrollers, digital signal processors, field programmable gate arrays, application-specific integrated circuits, and discrete circuits.
[0058] Processor 340 is also capable of executing other processes and programs residing in memory 360, such as operations for receiving and storing data. Processor 340 can move data into or out of memory 360 as needed by the executed processes. In embodiments, processor 340 is configured to execute one or more applications 362 based on OS 361 or in response to signals received from an external source or operator. For example, applications 362 may include encoders, decoders, VR or AR applications, camera applications (for still images and video), video phone calling applications, email clients, social media clients, SMS messaging clients, virtual assistants, etc. In embodiments, processor 340 is configured to receive and send media content. In embodiments, processor 340 may perform partial decoding and / or partial reconstruction of sub-grids as described herein.
[0059] Processor 340 is also coupled to an I / O interface 345 that provides electronic device 300 with the ability to connect to other devices, such as client devices 106 - 114 . I / O interface 345 is the communication path between these accessories and processor 340 .
[0060] Processor 340 is also coupled to input 350 and display 355. An operator of electronic device 300 can use input 350 to enter data or input into electronic device 300. Input 350 can be a keyboard, touch screen, mouse, trackball, voice input, or other device capable of serving as a user interface for user interaction with electronic device 300. For example, input 350 can include voice recognition processing, allowing the user to enter voice commands. In some examples, input 350 can include a touch panel, a (digital) pen sensor, keys, or an ultrasonic input device. A touch panel can recognize touch input using at least one scheme, such as capacitive, pressure-sensitive, infrared, or ultrasonic. Input 350 can be associated with sensors 365 and / or cameras by providing additional inputs to processor 340. In embodiments, sensors 365 include one or more inertial measurement units (IMUs) (such as accelerometers, gyroscopes, and magnetometers), motion sensors, optical sensors, cameras, pressure sensors, heart rate sensors, altimeters, and the like. Input 350 may also include control circuitry. In a capacitive solution, the input 350 may recognize touch or proximity.
[0061] Display 355 can be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED), an active-matrix OLED (AMOLED), or other display capable of rendering text and / or graphics, such as from a website, video, game, image, etc. Display 355 can be resized to fit within the HMD. Display 355 can be a single display screen or multiple display screens capable of creating a stereoscopic display. In an embodiment, display 355 is a heads-up display (HUD). Display 355 can display 3D objects, such as a 3D point cloud or mesh.
[0062] Memory 360 is coupled to processor 340. A portion of memory 360 may include RAM, and another portion of memory 360 may include flash memory or other ROM. Memory 360 may include a persistent storage device (not shown), which represents any structure capable of storing and facilitating retrieval of information (such as data, program code, and / or other suitable information). Memory 360 may include one or more components or devices that support long-term storage of data, such as read-only memory, a hard drive, flash memory, or an optical disk. Memory 360 may also include media content. Media content may include various types of media, such as images, videos, three-dimensional content, VR content, AR content, 3D point clouds, meshes, and the like.
[0063] The electronic device 300 also includes one or more sensors 365, which can measure physical quantities or detect the activation state of the electronic device 300 and convert the measured or detected information into electrical signals. For example, the sensors 365 may include one or more buttons for touch input, a camera, a gesture sensor, an IMU sensor (such as a gyroscope or gyro sensor and an accelerometer), an eye tracking sensor, an air pressure sensor, a magnetic sensor or magnetometer, a grip sensor, a proximity sensor, a color sensor, a biophysical sensor, a temperature / humidity sensor, an illumination sensor, an ultraviolet (UV) sensor, an electromyography (EMG) sensor, an electroencephalography (EEG) sensor, an electrocardiography (ECG) sensor, an IR sensor, an ultrasound sensor, an iris sensor, a fingerprint sensor, a color sensor (such as a red, green, and blue (RGB) sensor), and the like. The sensors 365 may also include control circuitry for controlling any of the sensors included therein.
[0064] As discussed in more detail below, one or more of these sensors 365 can be used to control a user interface (UI), detect UI input, determine the user's orientation and the direction the user is facing for three-dimensional content display recognition, etc. Any of these sensors 365 can be located within the electronic device 300, within an auxiliary device operably connected to the electronic device 300, within a head-mounted device configured to hold the electronic device 300, or in a single device in which the electronic device 300 includes a head-mounted device.
[0065] The electronic device 300 can create media content, such as generating a virtual object or capturing (or recording) content through a camera. The electronic device 300 can encode the media content to generate a bit stream so that the bit stream can be sent directly to another electronic device, or such as through a Figure 1 The electronic device 300 may receive the bit stream directly from another electronic device, or such as through Figure 1 The network 102 receives the bit stream indirectly.
[0066] although Figure 2 and Figure 3 An example of an electronic device is shown, but the Figure 2 and Figure 3 Make various changes. For example, Figure 2 and Figure 3 The various components in the can be combined, further subdivided, or omitted, and additional components can be added according to specific needs. As a specific example, processor 340 can be divided into multiple processors, such as one or more central processing units (CPUs) and one or more graphics processing units (GPUs). In addition, as with computing and communications, electronic devices and servers can have a wide variety of configurations, and Figure 2 and Figure 3 The present disclosure is not limited to any particular electronic device or server.
[0067] Figure 4 An example intra-coding process 400 according to the present disclosure is shown. Figure 4 The intra-coding process 400 is shown for illustration only. Figure 4 The scope of this disclosure is not limited to any particular implementation of the intra-frame coding process. For ease of explanation, Figure 4 The process 400 can be described as using Figure 3 However, process 400 may be used with any other suitable system and any other suitable electronic device.
[0068] like Figure 4 As shown, the intra-frame encoding process 400 encodes the grid frame using an intra-frame encoder 402. The intra-frame encoder 402 may be composed of Figure 2 The server 200 or Figure 3 The electronic device 300 shown represents or performs a process of creating a base mesh 404, which typically has a smaller number of vertices than the original mesh, and quantizing and compressing the base mesh in a lossy or lossless manner, and then encoding the base mesh into a compressed base mesh bitstream. Figure 4 As shown, the static mesh decoder decodes and reconstructs the base mesh, thereby providing a reconstructed base mesh 406. The reconstructed base mesh 406 then undergoes one or more levels of subdivision, and a displacement field is created for each subdivision that represents the difference between the original mesh and the subdivided reconstructed base mesh. In inter-frame coding of mesh frames, the base mesh 404 is encoded by sending vertex motions rather than directly compressing the base mesh. In either case, a displacement field 408 is created. Each displacement in the displacement field 408 has three components represented by x, y, and z. These can be relative to a standard coordinate system or a local coordinate system, where x, y, and z represent displacements in the local normal, tangent, and bitangent directions. It should be understood that multiple levels of subdivision can be applied, so that multiple subdivided mesh frames are created, and a displacement field is also created for each subdivided mesh frame.
[0069] Let the number of 3D displacement vectors in the grid frame displacement 408 be N. Let the displacement field be given by The displacement field 408 undergoes one or more levels of wavelet transform 410 to create a level of detail (LOD) signal , where k represents the index of the level of detail, N k represents the number of samples in the level of detail signal at level k, and numLOD represents the number of LODs. k (i) Scalar quantization.
[0070] like Figure 4 As shown, the quantized LOD signal corresponding to the displacement field 408 is encoded as a compressed bitstream. In various embodiments, the quantized LOD signal is encapsulated as a 2D image / video using an image packing operation and compressed in a lossless or lossy manner using an image or video encoder. However, another entropy encoder such as an asymmetric digital system (ANS) encoder or a binary arithmetic entropy encoder can be used to losslessly encode the quantized LOD signal. There may be other dependencies based on previous samples, across components, and across LODs that can be exploited. The displacement component provides a displacement vector that can be encoded as a geometric video component using any video codec indicated by the profile or using SEI messages. Alternatively, the profile can indicate that the displacement component is encoded using arithmetic coding.
[0071] Also like Figure 4 As shown, image decapsulation of the LOD signal is performed, and an inverse quantization operation and an inverse wavelet transform operation are performed to reconstruct the LOD signal. An inverse quantization operation is performed on the reconstructed base mesh 406, which is combined with the reconstructed LOD signal to reconstruct the deformed mesh. An attribute transfer operation is performed using the deformed mesh, the static / dynamic mesh, and the attribute map. A point cloud is a collection of 3D points and attributes representing the surface or volume of an object (such as color, normal direction, reflectivity, point size, etc.). These attributes are encoded as a compressed attribute bitstream. Figure 4 As shown, the encoding of the compressed attribute bitstream may also include padding operations, color space conversion operations, and video encoding operations. In various embodiments, atlas 405 may also be encoded as a compressed atlas bitstream. The atlas component provides information to the decoding and / or rendering system on how to perform inverse reconstruction. For example, the atlas may provide information on how to perform subdivision of the base mesh, how to apply displacement vectors to the subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.
[0072] Figure 4 The various functions or operations shown in FIG4 may be controlled by the control process 412. The intra-frame encoding process 400 outputs a compressed bitstream, which may be sent to and decoded by an electronic device such as the server 104 or the client devices 106-116, for example. Figure 4 As shown, the output compressed bitstream may include a compressed atlas bitstream, a compressed base grid bitstream, a compressed displacement bitstream and a compressed attribute bitstream as sub-bitstreams of the compressed bitstream.
[0073] although Figure 4 An example intra-frame encoding process 400 is shown, but may be used for Figure 4Various changes may be made. For example, the number and placement of the various components of the intra-coding process 400 may be varied as needed or desired. Furthermore, the intra-coding process 400 may be used in any other suitable process and is not limited to the specific process described above. In an embodiment, only the first (x) component of the displacement may be created and encoded, and the other two components (y and z) may be assumed to be zero. In this case, a flag may be signaled in the bitstream to indicate that the bitstream contains only data corresponding to the first (x) component, and that the other two components (y and z) should be assumed to be zero when decompressing and reconstructing the displacement field 408. As an example, Figure 4 The intra-coding process 400 may include encoding a bitstream and / or implementing appropriate signaling to allow a decoder to perform partial reconstruction and / or partial decoding of a sub-grid, as described in this disclosure.
[0074] Figure 5 An example trellis frame decoding process 500 according to the present disclosure is shown. Figure 5 The frame decoding process 500 is shown for illustration only. Figure 5 The scope of this disclosure is not limited to any particular implementation of the trellis frame decoding process. For ease of explanation, Figure 5 The process 500 can be described as using Figure 3 However, process 500 may be used with any other suitable system and any other suitable electronic device.
[0075] The decoding process 500 involves a demultiplexer 502 that receives an incoming bitstream. The demultiplexer separates various component bitstreams from the incoming bitstream, including a compressed base trellis bitstream, a compressed displacement bitstream, and a compressed attribute bitstream, such as a reference bitstream. Figure 4 The compressed attribute bitstream is decoded using a video decoder 504, the decoded attributes are processed using a color space conversion operation 506, and the original attributes of the grid are restored.
[0076] The decoding process 500 also includes decoding the displacement bitstream using a video decoder 508, which, in embodiments, can be the same video decoder as the video decoder 504. The decoded displacement data undergoes an image depacketization operation 510, an inverse quantization operation 512, and an inverse wavelet transform operation 514 as part of recovering position displacement data 516. Recovering the position displacement data 516 can also include performing one or more subdivision operations 518 on the grid frame recovered using the base grid decoder 520 and extracting x, y, and z components 522 (normals, tangents, and bitangents) from the subdivided grid frame. The base grid decoder 520 can perform an inverse quantization operation 521 before performing the subdivision operation 518.
[0077] The base mesh decoder 520 takes the base mesh bitstream provided by the demultiplexer 502 and reconstructs the intra base mesh frame from the base mesh bitstream using the static mesh decoder 524. The mesh buffer 530 provides the decoded intra frame to the motion decoder 532. At step 534, the motion decoder 532 also receives the inter frame data and uses the intra frame data, the inter frame data, and the associated tables to reconstruct the base mesh.
[0078] although Figure 5 An example frame decoding process 500 is shown, but may be used for Figure 5 Various changes may be made. For example, the number and placement of the various components of the frame decoding process 500 may be varied as needed or desired. Furthermore, the frame decoding process 500 may be used in any other suitable process and is not limited to the specific process described above. Furthermore, although shown as a series of steps, Figure 5 The steps in can overlap, occur in parallel, or occur any number of times. Figure 4 As described, the atlas bitstream can also be decoded to obtain an atlas that provides information on how to perform inverse reconstruction. For example, the atlas can provide information on how to perform subdivision of the base mesh, how to apply displacement vectors to the subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.
[0079] As described herein, mesh encoding and decoding operations are typically strongly sequential. For base meshes with a large number of vertices and a high frame rate, it may be difficult for a mesh codec to achieve real-time encoding and decoding. To alleviate this problem, sub-meshes are used. The base mesh can be divided into multiple sub-meshes. The sub-meshes may not be mutually exclusive, i.e., some vertices and triangles may be common to different sub-meshes. A sub-mesh can be encoded and decoded without using any information from other sub-meshes. This allows multiple instances of the mesh codec to operate in parallel on different sub-meshes. This also supports the ability to perform partial decoding of a mesh by decoding only some of the sub-meshes present in the bitstream. Each decoded sub-mesh can undergo subdivision, and the decoded displacement field is then used to refine the positions of the subdivided points belonging to that sub-mesh.
[0080] However, the creation of sub-meshes is not sufficient to ensure the functionality of partial decoding and reconstruction of the full-resolution mesh. For example, if the displacement video and the attribute video are not tiled, partial decoding is not possible for those videos. Furthermore, a wavelet transform is typically applied to the displacement field on the encoder side, and a corresponding inverse wavelet transform is applied on the decoder side. If the forward wavelet transform uses vertices corresponding to different sub-meshes, a partial reconstruction of the vertices corresponding to a sub-mesh cannot be performed on the decoder side unless the relevant wavelet coefficients corresponding to other sub-meshes required for the inverse wavelet transform have been decoded. In some cases, in addition, an inverse wavelet transform can be performed on the relevant wavelet coefficients corresponding to other sub-meshes.
[0081] Figure 6 An example process 600 for bitstream consistency of partial reconstruction and / or decoding of a sub-mesh according to the present disclosure is shown. Figure 6 Process 600 shown in FIG. 6 is for illustration only. Figure 6 The scope of this disclosure is not limited to any particular implementation of a process for bitstream consistency of partial reconstruction and / or decoding of a subgrid. For ease of explanation, Figure 6 The process 600 can be described as using Figure 3 However, process 600 may be used with any other suitable system and any other suitable electronic device.
[0082] like Figure 6 As shown, at step 602, the electronic device 300 initiates encoding of a dynamic grid bit stream, such as referring to Figure 4 As described. In step 604, it is determined whether partial reconstruction and / or partial decoding of the bitstream are supported. If not, the process 600 proceeds to step 608 to complete the encoding and output the bitstream. However, if in step 604, it is determined that partial reconstruction and / or partial decoding of the bitstream are supported, the process 600 moves to step 606. In step 606, one or more bitstream consistency requirements that can be recognized by the decoder due to various syntax or signaling elements are introduced during the encoding of the bitstream to assist in the partial reconstruction and / or decoding of data including vertices, connectivity and corresponding attributes of the sub-meshes. In one embodiment, all compressed bitstreams that comply with a specific dynamic mesh codec standard (such as V-DMC or its profile) automatically meet the consistency conditions without the need for any additional syntax or signaling elements. In step 608, the electronic device 300 completes the encoding of the bitstream and outputs the bitstream.
[0083] For example, in embodiments, bitstream consistency may be imposed, requiring that the vertices, connectivity, and attributes corresponding to a submesh can be reconstructed independently of the decoded data corresponding to other submeshes. This condition requires that the decoded data corresponding to different submeshes can be separated so that each subdivided submesh can be independently reconstructed. This may impose certain conditions on how the bitstream is generated. For example, if the displacement field undergoes a wavelet transform, such as wavelet transform 410, then when applying the forward wavelet transform, geometric data including displacements from other submeshes is excluded. That is, the forward wavelet transform is applied independently to each submesh. It should be understood that this condition does not mandate that all compressed data corresponding to a submesh can be independently decoded. That is, this condition allows for the independent reconstruction of the subdivided submeshes after the data is decoded, but does not necessarily allow for the independent decoding of the data. Enforcing that all compressed data corresponding to a submesh can be independently decoded would be a stronger condition and may not be implemented by all video codecs. Therefore, in various embodiments, ensuring that decoded data corresponding to different submeshes can be separated to allow for independent reconstruction of the subdivided submeshes can be supported as described above.
[0084] However, in embodiments, bitstream consistency may be imposed, which requires that the geometry and attribute data can be decoded and the vertices, connectivity, and attributes corresponding to a sub-mesh can be reconstructed independently of other sub-meshes. In embodiments, the decoded geometry data corresponds to displacements associated with the sub-mesh. In embodiments, because some video codecs may not include functionality for independently decodable slices, tiles, or sub-pictures, it may not be mandatory to place displacements and attribute data corresponding to different sub-meshes in independently decodable blocks or sub-pictures. Instead, a flag may be used to indicate partial decoding functionality as follows.
[0085] Consider a dynamic mesh bitstream containing multiple sub-meshes and having an attribute sub-bitstream. In various embodiments, a flag may be used to indicate whether each independently decodable unit of the attribute video sub-bitstream contains data corresponding to one and only one sub-mesh. A value of 1 for the flag provides an indication to the decoder that partial decoding and reconstruction of the sub-mesh is possible. In an embodiment, a flag, e.g., "one_submesh_per_independent_unit_attribute_flag", is included in sub-mesh information, e.g., "bmesh_sub_mesh_information()", to indicate partial decoding functionality of the attribute sub-bitstream. A value of 1 for the flag indicates that each independently decodable unit in the attribute sub-bitstream includes coded data from at most one sub-mesh. A value of 0 for the flag indicates that each independently decodable unit in the attribute sub-bitstream may include coded data corresponding to multiple sub-meshes.
[0086] In an embodiment, the semantics of a flag value of 1 may indicate that each independently decodable unit in the attribute sub-bitstream includes coded data corresponding to exactly one sub-grid. This does not allow independently decodable units that do not include coded data from a sub-grid. In an embodiment, the semantics of a flag value of 0 may indicate that at least one independently decodable unit in the attribute sub-bitstream includes coded data corresponding to multiple sub-grids.
[0087] In various embodiments, a similar flag as described above can be signaled for a shifted sub-bitstream containing coded data including a displacement field. For example, a flag, such as "one_submesh_per_independent_unit_geometry_flag," can be included in submesh information, such as "bmesh_sub_mesh_information()," to indicate partial decoding functionality for the shifted sub-bitstream. A flag value of 1 indicates that each independently decodable unit in the shifted sub-bitstream includes coded data from at most one submesh. A flag value of 0 indicates that each independently decodable unit in the shifted sub-bitstream can include coded data corresponding to multiple submeshes.
[0088] In an embodiment, the semantics of a flag value of 1 for a shifted sub-bitstream may indicate that each independently decodable unit in the shifted sub-bitstream includes coded data corresponding to exactly one sub-grid. In an embodiment, the semantics of a flag value of 0 may indicate that at least one independently decodable unit in the shifted sub-bitstream includes coded data corresponding to multiple sub-grids.
[0089] Various standards for vertex mesh and dynamic mesh encoding and decoding have been proposed. The following documents are incorporated herein by reference in their entirety as if fully set forth herein:
[0090] “V-Mesh Test Model v1,” ISO / IEC SC29 WG07 N00404, July 2022;
[0091] “V-DMC Test Model v2 (TMM v2)”, ISO / IEC SC29 WG07 N00456, October 2022;
[0092] “WD 1.0 of V-DMC,” ISO / IEC SC29 WG07, N00486, December 2022;
[0093] “WD 2.0 of V-DMC,” ISO / IEC SC29 WG07 N00546, January 2023;
[0094] “WD 3.0 of V-DMC,” ISO / IEC SC29 WG07 N00611, April 2023;
[0095] “WD 4.0 of V-DMC”, ISO / IEC JTC 1 / SC 29 / WG 07 N00611, August 2023; and
[0096] “WD 5.0 of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 7 N00744, August 2023.
[0097] In order to provide partial reconstruction and / or partial decoding of a sub-mesh according to the present disclosure, the standard may be updated to specify the syntax of the flag in WD 1.0 of V-DMC as follows.
[0098] H.8.1.3.2.2 Basic grid subgrid information
[0099] [Table 1]
[0100]
[0101] Various embodiments may use the semantics described above. Although in embodiments, the signaling of the flag may be within "basemesh_submesh_information()," in embodiments, the flag may be signaled as part of volumetric visibility information, in embodiments as part of a Supplemental Enhancement Information (SEI) message, or other syntax structures. In embodiments, if the video bitstream includes a hierarchy of independently decodable units, the flag may be interpreted as applying to the lowest level (i.e., smallest size) independently decodable unit.
[0102] although Figure 6 One example process 600 for bitstream consistency of partial reconstruction and / or decoding of a subgrid is shown, but may be used for Figure 6 Various changes may be made. 600 may be used for any other suitable process and is not limited to the specific process described above. In addition, although shown as a series of steps, Figure 6 The steps in can overlap, occur in parallel, or occur any number of times.
[0103] Figure 7 An example process 700 for defining coordinates associated with a sub-bitstream according to this disclosure is shown. Figure 7 Process 700 shown in FIG. 7 is for illustration only. Figure 7The scope of this disclosure is not limited to any particular implementation of the process for defining coordinates associated with a sub-bitstream. Figure 7 The process 700 can be described as using Figure 3 However, process 700 may be used with any other suitable system and any other suitable electronic device.
[0104] It is also useful to have information about which independently decodable units from the sub-bitstream are associated with each sub-grid, regardless of whether the corresponding independently decodable unit flag is 0 or 1. Since you can choose from a variety of video codecs such as AVC, HEVC, and VVC, each codec may have a different way of signaling independently decodable unit identifiers. To illustrate this, in an embodiment, the bounding rectangle or box 704 may be signaled relative to the 2D coordinates associated with that particular video sub-bitstream 702.
[0105] For example, for the attribute sub-bitstream corresponding to attribute j and atlas i, four syntax elements may be signaled for each sub-grid. To provide this information, the standard may be updated to specify the following in V-DMC WD 1.0 (for illustration purposes, the attribute identifier and atlas index are not shown):
[0106] H.8.1.3.2.2 Basic grid subgrid information
[0107] [Table 2]
[0108]
[0109]
[0110] "submesh_attribute_2d_pos_x[submeshIdx]" shown above can specify the x-coordinate of the upper left corner of the submesh bounding box 704 relative to the attribute video of the current submesh with ID submeshIdx. "submesh_attribute_2d_pos_y[submeshIdx]" shown above can specify the y-coordinate of the upper left corner of the submesh bounding box 704 relative to the attribute video of the current submesh with ID submeshIdx. "submesh_attribute_2d_size_x_minus1[submeshIdx]" shown above plus 1 can specify the width value of the submesh bounding box 704 relative to the attribute video of the current submesh with ID submeshIdx. The "submesh_attribute_2d_size_y_minus1[submeshIdx]" shown above plus 1 can specify the height value of the submesh bounding box 704 relative to the attribute video of the current submesh with ID submeshIdx. In an embodiment, the position and size of the submesh bounding box 704 can be specified in terms of the number of pixels relative to the attribute sub-bitstream. In an embodiment, the position and size of the submesh bounding box 704 can be specified as a multiple of an integer. In an embodiment, the integer can be a "PatchPackingBlockSize" element. In an embodiment, instead of ue(v), a fixed length (such as 16 bits) or a fixed length signaled in the bitstream can be used to signal the syntax element.
[0111] In an embodiment, a requirement for bitstream conformance may be that all pixels within the bounding box specified by the syntax elements "submesh_attribute_2d_pos_x", "submesh_attribute_2d_pos_y", "submesh_attribute_2d_size_x_minus1", and "submesh_attribute_2d_size_x_minus1" are within the attribute sub-bitstream. However, it will be understood that, in addition or alternatively, similar positioning data for similar one or more bounding boxes may be signaled for the displacement data to identify independently decodable units corresponding to sub-meshes in the displacement sub-bitstream.
[0112] Assuming the bitstream consistency requirement imposed above, i.e., that it should be possible to reconstruct the vertex, connectivity, and attribute data corresponding to a sub-mesh independently of the decoded data corresponding to other sub-meshes, in this case, in an embodiment, another bitstream consistency requirement may be introduced to provide a tight bounding box comprising the encoded data of a single sub-mesh, such as Figure 7. For example, a tight bounding box corresponding to a sub-grid may be defined as the smallest rectangle (in terms of width and height) that includes all 2D positions in the attribute (or displacement) sub-bitstream that includes the coded data corresponding to the sub-grid. Thus, as Figure 7 As shown, while in an embodiment, the bounding box 704 may be used to specify the 2D coordinates in which the coded data corresponding to the sub-mesh is retained, the tight bounding box 706 may be used to define a smaller area in the sub-bitstream 702 that includes the coded data corresponding to the sub-mesh. Thus, in an embodiment, a requirement for bitstream conformance may be that the bounding box signaled by the syntax elements "submesh_attribute_2d_pos_x", "submesh_attribute_2d_pos_y", "submesh_attribute_2d_size_x_minus1", and "submesh_attribute_2d_size_x_minus1" is as tight as the tight bounding box 706.
[0113] although Figure 7 One example process 700 for defining coordinates associated with a sub-bitstream is shown, but may be used for Figure 7 Various changes may be made. For example, the number and placement of the various components of process 700 may be varied as needed or desired. Furthermore, process 700 may be used in any other suitable process and is not limited to the specific process described above.
[0114] Figure 8 An example process 800 for defining coordinates associated with sub-bitstreams in a non-overlapping manner according to this disclosure is shown. Figure 8 Process 800 is shown in FIGURE 8 for illustration only. Figure 8 The scope of this disclosure is not limited to any particular implementation of the process for defining coordinates associated with sub-bitstreams in a non-overlapping manner. Figure 8 The process 800 can be described as using Figure 3 However, process 800 may be used with any other suitable system and any other suitable electronic device.
[0115] In an embodiment, such as reference Figure 7 As described, a requirement for bitstream consistency may be that the bounding boxes signaled corresponding to different subgrids are non-overlapping. For illustration purposes, Figure 8A sub-bitstream 802 is shown having a first bounding box 804, a second bounding box 806, and a third bounding box 808. The first bounding box 804 includes data corresponding to a first sub-grid, the second bounding box 806 includes data corresponding to a second sub-grid, and the third bounding box 808 includes data corresponding to a third sub-grid. Figure 8 As shown, the first bounding box 804 , the second bounding box 806 , and the third bounding box 808 are defined such that the bounding boxes do not overlap to prevent data of other subgrids (ie, more than one subgrid) from being within a single bounding box.
[0116] In an embodiment, a requirement for bitstream consistency may be that the bounding boxes of the signaling corresponding to different sub-grids are tight and non-overlapping. For example, such as reference Figure 7 The tight bounding box described for the tight bounding box 706 may be used for the first bounding box 804 , the second bounding box 806 , and the third bounding box 808 , such that the bounding boxes are all tight and non-overlapping.
[0117] For example, if the above-mentioned condition regarding non-overlap of tight bounding boxes corresponding to different sub-grids is not applied, sometimes part of a sub-grid can be placed in one corner, and another part of the sub-grid can be placed in another corner. In such a case, a single bounding box will result in overlap with a large number of independently decodable units that do not have encoded data corresponding to the sub-grid. To avoid this, in an embodiment, multiple bounding boxes can be signaled for each sub-grid. This can be achieved by signaling the number of bounding boxes minus 1. This can then be followed by signaling the top left position and size of the bounding box for each bounding box, similar to what is described above in the present disclosure. In an embodiment, multiple bounding boxes can be signaled for each sub-grid separately with respect to the attribute sub-bitstream and the displacement sub-bitstream.
[0118] In an embodiment, a video decoder for an attribute sub-bitstream may use bounding box information to derive which independently decodable units need to be decoded for each sub-grid, and only decode those units to achieve partial decoding. In an embodiment, a similar bounding box corresponding to each sub-grid may be signaled for a shifted sub-bitstream containing a displacement field. In an embodiment, a video decoder for a shifted sub-bitstream may use bounding box information to derive which independently decodable units need to be decoded for each sub-grid, and only decode those units to achieve partial decoding. In an embodiment, instead of signaling bounding boxes, the number of independently decodable units containing data corresponding to that sub-grid may be signaled for each sub-grid, followed by identifiers (IDs) of those units. For example, this may be a block ID, slice ID, or sub-picture ID.
[0119] In an embodiment, a combination of the volume annotation SEI message family syntax can be used to achieve similar functionality. As an example, each sub-mesh can be assigned to a scene object, and a combination of a scene object information SEI message and a volume rectangle information SEI message can be used to send information about the bounding box. One advantage of this approach is that the position of the bounding box can be updated within a sequence. In such a case, the scene object information SEI message can be modified to be able to associate a scene object with a specific sub-mesh. Similarly, the scene object information SEI message or the volume rectangle information SEI message can be modified to include information about whether the rectangle refers to an attribute sub-bitstream or a displacement sub-bitstream.
[0120] although Figure 8 One example process 800 is shown for defining coordinates associated with sub-bitstreams in a non-overlapping manner, but may be used for Figure 8 Various changes may be made. For example, the number and placement of the various components of process 800 may be varied as needed or desired. For example, although Figure 8 Three bounding boxes are shown in , but any number of bounding boxes may be used, such as depending on how many subgrids there are for the bitstream. Furthermore, process 800 may be used in any other suitable process and is not limited to the specific process described above.
[0121] Figure 9 An example encoding method 900 for partial reconstruction and / or decoding of a sub-grid according to the present disclosure is shown. For ease of explanation, Figure 9 The method 900 is described as using Figure 3 However, the method 900 may be used with any other suitable system and any other suitable electronic device.
[0122] like Figure 9 As shown, at step 902, the electronic device 300 receives geometric data and attribute data associated with a separate sub-mesh. In an embodiment, the geometric data corresponds to a displacement created based on subdividing one or more sub-meshes. At step 904, the electronic device 300 encodes the geometric data and attribute data associated with the separate sub-mesh into a displacement sub-bitstream and an attribute sub-bitstream, respectively, such as a reference image. Figure 4 In various embodiments, bitstream conformance requirements may be imposed, such as those described in reference Figure 6 As described, geometry data and attribute data associated with an individual sub-mesh can be separated from data corresponding to one or more other sub-meshes in the displacement sub-bitstream and the attribute sub-bitstream during decoding.
[0123] In step 906, the electronic device combines the shifted sub-bitstream and the attribute sub-bitstream into a compressed bitstream, as also described in reference Figure 4In an embodiment, during encoding, the electronic device 300 may set one or more flags in the compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of the shift sub-bitstream and the attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid, as also described in reference Figure 6 In an embodiment, one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
[0124] In an embodiment, during encoding, the electronic device 300 may signal a bounding box associated with a separate sub-grid, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream, as shown in FIG. Figure 7 and Figure 8 In various embodiments, the electronic device 300 may form the bounding box to be at least one of: (i) within the smallest possible area while still including all 2D locations containing encoded data corresponding to the individual sub-grids, and (ii) non-overlapping with one or more other bounding boxes associated with one or more other sub-grids.
[0125] In an embodiment, during encoding, the electronic device 300 may include a signaling element indicating the number of independently decodable units of at least one of the shifted sub-bitstream and the attribute sub-bitstream corresponding to the subdivided sub-grid, and an identifier for each of the independently decodable units. That is, for each sub-grid, the number of independently decodable units containing data corresponding to the sub-grid may be signaled, followed by the IDs of those units, which may be, for example, block IDs, slice IDs, or sub-picture IDs.
[0126] During encoding, in order to provide that decoded data related to a particular sub-grid can be separated from other data during decoding, the electronic device 300 can apply a forward wavelet transform (e.g., wavelet transform 410) to the displacement field such that geometric data including displacements from other sub-grids is excluded when applying the forward wavelet transform. In other words, the forward wavelet transform is applied independently to each sub-grid.
[0127] At step 908, the electronic device 300 outputs a compressed bitstream. The output bitstream may also include a compressed base grid bitstream and an atlas sub-bitstream, for example, Figure 4 The output bitstream may be sent to an external device or a storage device on the electronic device 300.
[0128] although Figure 9One example of an encoding method 900 for partial reconstruction and / or decoding of a sub-grid is shown, but may be used for Figure 9 For example, although shown as a series of steps, Figure 9 The steps in can overlap, occur in parallel, or occur any number of times.
[0129] Figure 10 An example decoding method 1000 for partial reconstruction and / or decoding of a sub-grid according to the present disclosure is shown. For ease of explanation, Figure 10 The method 1000 is described as using Figure 3 However, the method 1000 may be used with any other suitable system and any other suitable electronic device.
[0130] like Figure 10 As shown, in step 1002, the electronic device 300 receives a compressed bitstream having sub-bitstreams, the sub-bitstreams including a base grid sub-bitstream, a displacement sub-bitstream, and an attribute sub-bitstream. In an embodiment, the compressed bitstream may further include an atlas sub-bitstream. In step 1004, the electronic device decodes at least a portion of the compressed bitstream, which may include the electronic device decoding a plurality of sub-grids from the base grid sub-bitstream, decoding geometric data from the displacement sub-bitstream, and decoding attribute data from the attribute sub-bitstream, as also described in reference to FIG. Figure 5 described.
[0131] In an embodiment, to decode at least a portion of the compressed bitstream, the electronic device 300 decodes the geometric data and attributes of the subdivided sub-mesh independently of the encoded data associated with one or more other sub-meshes included in the compressed bitstream, as also described with reference to Figure 6 As described. In an embodiment, this may include the electronic device 300 identifying one or more flags in the compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of the displacement sub-bitstream and the attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid. In an embodiment, the one or more flags are signaled as part of the volume visibility information or as part of a supplemental enhancement information (SEI) message.
[0132] In an embodiment, the electronic device 300 may decode a bounding box associated with the subdivided sub-grid signaled by the encoder, wherein the bounding box corresponds to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream, as described in reference to FIG. Figure 7 and Figure 8In an embodiment, the bounding box occupies the smallest possible area while still including all 2D locations containing encoded data corresponding to the sub-mesh.
[0133] In one embodiment, the electronic device 300 may identify, from the signaling element, the number of independently decodable units of at least one of the shifted sub-bitstream and the attribute sub-bitstream corresponding to the subdivided sub-grid, and an identifier for each of the independently decodable units. That is, for each sub-grid, the number of independently decodable units containing data corresponding to the sub-grid may be signaled, followed by the IDs of those units, which may be, for example, block IDs, slice IDs, or sub-picture IDs.
[0134] In step 1006, the electronic device subdivides a sub-mesh in the plurality of sub-meshes to generate a subdivided sub-mesh. In step 1008, the electronic device reconstructs vertex positions of the subdivided sub-mesh using the decoded geometric data and reconstructs attributes of the subdivided sub-mesh using the decoded attribute data, independently of the decoded data corresponding to one or more other sub-meshes, as also described in reference to FIG. Figure 6 The decoded geometric data may be from an independently decodable unit of a signaled ID corresponding to the sub-grid. This may include the electronic device 300 applying an inverse wavelet transform to the decoded geometric data corresponding to the sub-grid, independently of the decoded geometric data corresponding to one or more other sub-grids, to obtain a displacement associated with the subdivided sub-grid.
[0135] At step 1010, the electronic device reconstructs at least a portion of the mesh frame using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-meshes. At step 1012, the electronic device 300 outputs decoded content, such as a 3D video including the reconstructed mesh frame. For example, the output decoded content may be sent to an external device or a storage device on the electronic device 300.
[0136] although Figure 10 One example of a decoding method 1000 for partial reconstruction and / or decoding of a sub-grid is shown, but may be used for Figure 10 For example, although shown as a series of steps, Figure 10 The steps in can overlap, occur in parallel, or occur any number of times.
[0137] Although the present disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims. The description in this application should not be interpreted as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the claimed subject matter is defined by the claims.
[0138] According to an embodiment of the present disclosure, a device is provided. The device may include a communication interface configured to receive a compressed bitstream having sub-bitstreams, the sub-bitstreams including a base mesh sub-bitstream, a displacement sub-bitstream, and an attribute sub-bitstream. A processor may be operably coupled to the communication interface. The processor may be configured to decode at least a portion of the compressed bitstream. The processor may be configured to decode multiple sub-meshes from the base mesh sub-bitstream. The processor may be configured to decode geometry data from the displacement sub-bitstream. The processor may be configured to decode attribute data from the attribute sub-bitstream. The processor may be configured to subdivide a sub-mesh from the multiple sub-meshes to generate a subdivided sub-mesh. The processor may be configured to reconstruct vertex positions of the subdivided sub-mesh using the decoded geometry data and reconstruct attributes of the subdivided sub-mesh using the decoded attribute data, independently of decoded data corresponding to one or more other sub-meshes. The processor may be configured to reconstruct at least a portion of a mesh frame using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-mesh.
[0139] According to an embodiment of the present disclosure, the processor may be configured to use an inverse wavelet transform on the decoded geometric data corresponding to the sub-mesh independently of the decoded geometric data corresponding to one or more other sub-meshes to obtain the displacement associated with the subdivided sub-mesh.
[0140] According to an embodiment of the present disclosure, the processor may be configured to decode geometric data and attributes of a subdivided sub-mesh independently of encoded data associated with one or more other sub-meshes included in the compressed bitstream.
[0141] According to an embodiment of the present disclosure, the processor may be configured to identify one or more flags in a compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of a displacement sub-bitstream and an attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid.
[0142] According to an embodiment of the present disclosure, one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
[0143] According to an embodiment of the present disclosure, the processor may be further configured to decode a bounding box associated with the subdivided sub-grid signaled by the encoder, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream.
[0144] According to an embodiment of the present disclosure, the bounding box may occupy the smallest possible area while still including all 2D locations containing encoded data corresponding to the sub-mesh.
[0145] According to an embodiment of the present disclosure, the processor may further be configured to identify, from the signaling element, the number of independently decodable units of at least one of the displacement sub-bitstream and the attribute sub-bitstream corresponding to the subdivided sub-grid and an identifier of each of the independently decodable units.
[0146] According to an embodiment of the present disclosure, a method is provided. The method may include receiving a compressed bitstream having sub-bitstreams, the sub-bitstreams including a base grid sub-bitstream, a displacement sub-bitstream, and an attribute sub-bitstream. The method may include decoding at least a portion of the compressed bitstream, including decoding multiple sub-grids from the base grid sub-bitstream, decoding geometric data from the displacement sub-bitstream, and decoding attribute data from the attribute sub-bitstream. The method may include subdividing a sub-grid of the multiple sub-grids to generate a subdivided sub-grid. The method may include reconstructing vertex positions of the subdivided sub-grid using the decoded geometric data, and reconstructing attributes of the subdivided sub-grid using the decoded attribute data, independently of decoded data corresponding to one or more other sub-grids. The method may include reconstructing at least a portion of a grid frame using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-grid.
[0147] According to an embodiment of the present disclosure, a method may include using an inverse wavelet transform on the decoded geometric data corresponding to the sub-mesh, independently of the decoded geometric data corresponding to one or more other sub-meshes, to obtain displacements associated with the subdivided sub-mesh.
[0148] According to an embodiment of the present disclosure, decoding at least a portion of the compressed bitstream may include decoding geometric data and attributes of the subdivided sub-mesh independently of encoded data associated with one or more other sub-meshes included in the compressed bitstream.
[0149] According to an embodiment of the present disclosure, a method may include identifying one or more flags in a compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of a displacement sub-bitstream and an attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid.
[0150] According to an embodiment of the present disclosure, one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
[0151] According to an embodiment of the present disclosure, a method may include decoding a bounding box associated with a subdivided sub-mesh signaled by an encoder, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of a displacement sub-bitstream and a property sub-bitstream.
[0152] According to an embodiment of the present disclosure, the bounding box may occupy the smallest possible area while still including all 2D locations containing encoded data corresponding to the sub-mesh.
[0153] According to an embodiment of the present disclosure, the method may include identifying, from a signaling element, a number of independently decodable units of at least one of a displacement sub-bitstream and an attribute sub-bitstream corresponding to the subdivided sub-grid and an identifier of each of the independently decodable units.
[0154] According to an embodiment of the present disclosure, a device is provided. The device may include a communication interface and a processor operably coupled to the communication interface. The processor may be configured to encode geometric data and attribute data associated with a separate sub-mesh into a displacement sub-bitstream and an attribute sub-bitstream, respectively. The geometric data and attribute data associated with the separate sub-mesh may be separable from data corresponding to one or more other sub-meshes in the displacement sub-bitstream and the attribute sub-bitstream during decoding. The processor may be configured to combine the displacement sub-bitstream and the attribute sub-bitstream into a compressed bitstream.
[0155] According to an embodiment of the present disclosure, the processor may further be configured to set one or more flags in the compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of the displacement sub-bitstream and the attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid.
[0156] According to an embodiment of the present disclosure, one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
[0157] According to an embodiment of the present disclosure, the processor may be further configured to signal a bounding box associated with the individual sub-grid, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream.
[0158] According to an embodiment of the present disclosure, the processor may also be configured to form the bounding box to be at least one of: within the smallest possible area while still including all 2D locations containing encoded data corresponding to the individual sub-grids; and not overlapping with one or more other bounding boxes associated with one or more other sub-grids.
Claims
1. A device comprising: a communication interface configured to receive a compressed bitstream having sub-bitstreams, the sub-bitstreams comprising a base trellis sub-bitstream, a shifted sub-bitstream, and an attribute sub-bitstream; and a processor operatively coupled to the communication interface, the processor configured to: decoding at least a portion of a compressed bitstream, wherein the processor is configured to decode a plurality of subgrids from a base grid subbitstream, decode geometry data from a displacement subbitstream, and decode attribute data from an attribute subbitstream; subdividing a subgrid of the plurality of subgrids to generate a subdivided subgrid; reconstructing vertex positions of the subdivided submesh using the decoded geometry data and reconstructing attributes of the subdivided submesh using the decoded attribute data, independent of decoded data corresponding to one or more other submeshes; and At least a portion of the mesh frame is reconstructed using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-meshes.
2. The device according to claim 1, wherein The processor is further configured to use an inverse wavelet transform on the decoded geometric data corresponding to the sub-mesh, independently of the decoded geometric data corresponding to the one or more other sub-meshes, to obtain displacements associated with the subdivided sub-mesh.
3. The device according to any one of claims 1 to 2, wherein: To decode at least a portion of the compressed bitstream, the processor is further configured to decode geometric data and attributes of the subdivided sub-mesh independently of encoded data associated with the one or more other sub-meshes included in the compressed bitstream.
4. The device according to claim 3, wherein The processor is further configured to identify one or more flags in the compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of the displacement sub-bitstream and the attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid.
5. The device according to claim 4, wherein The one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
6. The device according to any one of claims 1 to 5, wherein: The processor is further configured to decode a bounding box associated with the subdivided sub-grid signaled by the encoder, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream.
7. The device according to claim 6, wherein The bounding box occupies the smallest possible area while still including all 2D locations containing encoded data corresponding to the sub-mesh.
8. The device according to any one of claims 1 to 7, wherein The processor is further configured to identify from the signaling element a number of independently decodable units of at least one of the displacement sub-bitstream and the attribute sub-bitstream corresponding to the subdivided sub-grid and an identifier of each of the independently decodable units.
9. A method comprising: receiving a compressed bitstream having sub-bitstreams, the sub-bitstreams comprising a base trellis sub-bitstream, a shifted sub-bitstream, and an attribute sub-bitstream; decoding at least a portion of the compressed bitstream, including decoding a plurality of subgrids from the base grid subbitstream, decoding geometry data from the displacement subbitstream, and decoding attribute data from the attribute subbitstream; subdividing a subgrid of the plurality of subgrids to generate a subdivided subgrid; reconstructing vertex positions of the subdivided submesh using the decoded geometry data and reconstructing attributes of the subdivided submesh using the decoded attribute data, independent of decoded data corresponding to one or more other submeshes; and At least a portion of the mesh frame is reconstructed using the reconstructed vertex positions and reconstructed attributes corresponding to the subdivided sub-meshes.
10. The method according to claim 9, further comprising: An inverse wavelet transform is applied to the decoded geometric data corresponding to the sub-mesh, independently of the decoded geometric data corresponding to the one or more other sub-meshes, to obtain displacements associated with the subdivided sub-mesh.
11. A device comprising: Communication interface; and a processor operatively coupled to the communication interface, the processor configured to: Encoding the geometry data and attribute data associated with individual sub-meshes into the displacement sub-bitstream and attribute sub-bitstream, respectively, wherein geometry data and attribute data associated with an individual sub-mesh can be separated during decoding from data corresponding to one or more other sub-meshes in the displacement sub-bitstream and the attribute sub-bitstream; and The shift sub-bitstream and the attribute sub-bitstream are combined into a compressed bitstream.
12. The device according to claim 11, wherein The processor is further configured to set one or more flags in the compressed bitstream, the one or more flags indicating whether each independently decodable unit of at least one of the displacement sub-bitstream and the attribute sub-bitstream includes data corresponding to (i) one and only one sub-grid or (ii) at most one sub-grid.
13. The device according to claim 12, wherein The one or more flags are signaled as part of the volume visibility information or as part of a Supplemental Enhancement Information (SEI) message.
14. The device according to any one of claims 11 to 13, wherein The processor is further configured to signal a bounding box associated with an individual sub-mesh, the bounding box corresponding to two-dimensional (2D) coordinates of at least one of the displacement sub-bitstream and the attribute sub-bitstream.
15. The device according to claim 14, wherein The processor is further configured to form the bounding box as at least one of: in the smallest possible area while still including all 2D locations containing encoded data corresponding to individual sub-grids; and One or more other bounding boxes associated with one or more other sub-meshes do not overlap.