A method, apparatus and storage medium for decoding a geometry patch of a three-dimensional mesh
Patent Information
- Application Number
- CN202280032180.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2022-10-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-10-31
Smart Images

Figure CN117242777B_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 316,876, filed March 4, 2022, which is incorporated herein by reference in its entirety. This application also claims priority to U.S. Patent Application No. 17 / 973,824, filed October 26, 2022, which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure generally relates to mesh encoding (or compression) and mesh decoding (or decompression) processes, and particularly to methods and systems for mesh compression with limited geometric dynamic range. Background Technology
[0004] The background description provided herein is intended to present the overall context of this application. The extent of the work of the currently named inventors described in the background section and various aspects of this specification does not imply that it was prior art at the time of filing of this application, nor is it expressly or implied that it was acknowledged as prior art to this application.
[0005] Various technologies have been developed to capture, represent, and simulate real-world objects, environments, and more in three-dimensional (3D) space. 3D representations of the world enable more immersive and interactive forms of communication. Exemplary 3D representations of objects and environments include, but are not limited to, point clouds and meshes. A series of 3D representations of objects and environments can form video sequences. Redundancy and correlation within the sequence of 3D representations of objects and environments can be leveraged to compress and encode these video sequences into more compact digital forms. Summary of the Invention
[0006] This disclosure generally relates to the encoding (compression) and decoding (decompression) of 3D meshes, and specifically to mesh compression with a limited geometric dynamic range.
[0007] This disclosure describes a method for decoding geometric patches of a three-dimensional mesh. The method includes receiving an encoded bitstream by a device, the encoded bitstream comprising geometric patches of a three-dimensional mesh. The device includes a memory storing instructions and a processor communicating with the memory. The method further includes extracting a first syntax element from the encoded bitstream by the device, the first syntax element indicating whether the geometric patch is partitioned, the geometric patch comprising one or more partitions; and for each partition in the geometric patch, obtaining by the device a dynamic range of pixel values of points in each partition, wherein the points correspond to vertices in the three-dimensional mesh, the dynamic range enabling encoding of the geometric patch within a predetermined bit depth.
[0008] According to another aspect, embodiments of this disclosure provide an apparatus for encoding or decoding 3D meshes. The apparatus includes a memory storing instructions and a processor communicating with the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the described method.
[0009] In another aspect, embodiments of this disclosure provide a non-volatile computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the methods described above.
[0010] The above and other aspects and their implementations are described in more detail in the accompanying drawings, description and claims. Attached Figure Description
[0011] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings, wherein:
[0012] Figure 1 This is a simplified block diagram of a communication system according to an embodiment;
[0013] Figure 2 This is a simplified block diagram of a flow system according to an embodiment;
[0014] Figure 3 A block diagram of an encoder for encoding grid frames according to some embodiments is shown;
[0015] Figure 4 A block diagram of a decoder for decoding a compressed bitstream corresponding to a grid frame, according to some embodiments, is shown.
[0016] Figure 5 This is a simplified block diagram of a video decoder according to an embodiment;
[0017] Figure 6 This is a simplified block diagram of a video encoder according to an embodiment;
[0018] Figure 7 A block diagram of an encoder for encoding grid frames according to some embodiments is shown;
[0019] Figure 8 A schematic diagram of a frame for mesh compression according to some embodiments of the present disclosure is shown;
[0020] Figure 9 A flowchart outlining examples of processes according to some embodiments is shown;
[0021] Figure 10 A schematic diagram of a frame for mesh compression according to some embodiments of the present disclosure is shown;
[0022] Figure 11 Another schematic diagram of a framework for mesh compression according to some embodiments of the present disclosure is shown;
[0023] Figure 12 This is a schematic diagram of a computer device according to an embodiment. Detailed Implementation
[0024] Throughout the specification and claims, terms may have nuanced meanings implied or implied by the context, and not merely their explicitly stated meanings. The phrases “in one embodiment” or “in some embodiments” as used herein do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” as used herein do not necessarily refer to different embodiments. Similarly, the phrases “in one implementation” or “in some implementations” as used herein do not necessarily refer to the same implementation, and the phrases “in another implementation” or “in other embodiments” as used herein do not necessarily refer to different implementations. For example, the claimed subject matter is intended to include all or part of a combination of exemplary embodiments / implementations.
[0025] Generally, terms can be understood, at least in part, based on their usage in the context. For example, terms such as “and,” “or,” or “and / or” as used herein can include a variety of meanings, which can depend at least in part on the context in which such terms are used. Typically, “or,” when used with an associative list such as A, B, or C, is intended to mean A, B, and C, used herein in an inclusive sense, and A, B, or C, used herein in an exclusive sense. Additionally, the terms “one or more” or “at least one,” as used herein, depend at least in part on the context and can be used to describe any feature, structure, or characteristic in the singular or in the plural form to describe a combination of features, structures, or characteristics. Similarly, terms such as “a,” “an,” or “the” can also be understood to convey either a singular or a plural usage, depending at least in part on the context. Furthermore, the terms “based on” or “determined by” can be understood to not necessarily convey an exclusive set of factors, but may allow for additional factors that are not necessarily explicitly described, again depending at least in part on the context.
[0026] Technological advancements in 3D media processing, such as progress in 3D acquisition, 3D modeling, and 3D rendering, have facilitated the creation of 3D content ubiquitous across multiple platforms and devices. This 3D content contains information that can be processed to generate various forms of media, providing, for example, immersive viewing / presentation and interactive experiences. The applications of 3D content are extensive, including but not limited to virtual reality, augmented reality, metaverse interaction, games, immersive video conferencing, robotics, and computer-aided design (CAD). According to one aspect of this disclosure, to improve immersive experiences, 3D models are becoming increasingly complex, and the creation and use of 3D models require substantial data resources, such as data storage, data transmission, and data processing resources.
[0027] Compared to traditional 2D content, which is typically represented by datasets in the form of 2D pixel arrays (such as images), 3D content with full-resolution pixelation can be dauntingly resource-intensive and unnecessary in many (if not most) practical applications. In most immersive 3D applications, according to some aspects of this disclosure, a lower data-intensive representation of 3D content can be employed. For example, in most applications, only topographic information of objects in a 3D scene (a real-world scene captured by sensors such as a LiDAR device or an animated 3D scene generated by software tools) is likely needed, rather than volumetric information. Therefore, more efficient forms of datasets can be used to represent 3D objects and 3D scenes. For example, a 3D mesh can be used as a 3D model to represent immersive 3D content, such as 3D objects in a 3D scene.
[0028] A mesh (or mesh model) of one or more objects can comprise a set of vertices. Vertices can be connected to each other to form edges. Edges can be further connected to form faces. Faces can be further connected to form polygons. The 3D surfaces of various objects can be decomposed into, for example, faces and polygons. Each of a vertex, edge, face, polygon, or surface can be associated with various attributes such as color, normal, texture, etc. The normal of a surface can be called a surface normal; and / or the normal of a vertex can be called a vertex normal. Information about how vertices connect to form edges, faces, or polygons can be called connectivity information. Connectivity information is important for uniquely defining the components of a mesh because the same set of vertices can form different faces, surfaces, and polygons. Typically, the position of a vertex in 3D space can be represented by its 3D coordinates. A face can be represented by a set of sequentially connected vertices, each vertex associated with a set of 3D coordinates. Similarly, an edge can be represented by two vertices, each vertex associated with its 3D coordinates. Vertices, edges, and faces can be indexed in a 3D mesh dataset.
[0029] Meshes can be defined and described by a set of one or more of these basic element types. However, not all of the above types of elements are required to fully describe a mesh. For example, a mesh can be fully described using only vertices and their connectivity. As another example, a mesh can be fully described using only faces and a list of common vertices of those faces. Therefore, meshes can be of various optional types, composed of optional datasets and described by different formats. Example mesh types include, but are not limited to, face-vertex meshes, winged-edge meshes, half-edge meshes, quad-edge meshes, corner-table meshes, vertex-vertex meshes, etc. Accordingly, the mesh dataset can be stored along with information conforming to optional file formats, which have file extensions including, but not limited to, .raw, .blend, .fbx, .3ds, .dae, .dng, 3dm, .dsf, .dwg, .obj, .ply, .pmd, .stl, amf, .wrl, .wrz, .x3d, .x3db, .x3dv, .x3dz, .x3dbz, .x3dvz, .c4d, .lwo, .smb, .msh, .mesh, .veg, .z3d, .vtk, .l4d, etc. The attributes of these elements, such as color, normals, and textures, can be included in the mesh dataset in various ways.
[0030] In some implementations, the vertices of the mesh can be mapped to a pixelated 2D space, called UV space. Therefore, each vertex of the mesh can be mapped to a pixel in UV space. In some implementations, a vertex can be mapped to more than one pixel in UV space; for example, vertices at boundaries can be mapped to two or three pixels in UV space. Similarly, faces or surfaces in the mesh can be sampled into multiple 3D points, which may or may not be recorded in the mesh's vertices, and these multiple 3D points can also be mapped to pixels in 2D UV space. Mapping the vertices of faces or surfaces in the mesh and the sampled 3D points to UV space, and subsequent data analysis and processing in UV space, can facilitate the storage, compression, and encoding / decoding of 3D datasets of the mesh or mesh sequences, as described in further detail below. The mapped UV space dataset can be referred to as a UV image, or a 2D map, or a 2D image of the mesh.
[0031] After mapping vertices and sampled surface points in a 3D mesh to a 2D UV space, some pixels can be mapped to vertices and sampled surface points in the 3D mesh, while others may remain unmapped. Each mapped pixel in the 2D image of the mesh can be associated with information about a corresponding mapped vertex or surface point in the 3D mesh. Depending on the type of information included in the pixels in the UV space, various 2D images or 2D graphs of the mesh can be constructed. A collection of multiple 2D graphs can be used as a substitute or joint representation of the mesh.
[0032] For example, the simplest 2D map of a mesh can be constructed as an occupancy map. An occupancy map indicates pixels in UV space that are mapped to the mesh's 3D vertices or sampled surface points. Occupancy can be indicated by a binary indicator at each 2D pixel; for example, a binary value "1" indicates mapping or occupancy, while a binary value "0" indicates non-mapping or non-occupancy. Therefore, an occupancy map can be constructed as a 2D image. While a normal 2D image contains an array of three channels (RGB, YUV, YCrCb, etc.) with, for example, 8-bit color depth, a 2D occupancy map of such a mesh only requires one binary channel.
[0033] For another example, a 2D geometry can be constructed for the mesh. Instead of a single binary channel, the 2D geometry is a complete three-channel image, where the three color channels at each occupied pixel correspond to the three 3D coordinates of the corresponding mapped vertex or sampled 3D point in the mesh.
[0034] In some implementations, additional 2D maps can be constructed for the mesh. For example, a set of attributes for each vertex and sampled 3D point of the mesh can be extracted from the mesh dataset, and this set of attributes can be encoded into the three color channels of a 2D map image. Such a 2D map can be referred to as a property map of the mesh. A particular property map may contain the three-channel color of each occupied pixel in UV space. For another example, the texture attributes associated with each mapped vertex or sampled 3D point of the mesh can be parameterized as three-channel values and encoded into a 2D property map. For yet another example, the normal attributes associated with each mapped vertex or sampled 3D point of the mesh can be parameterized as three-channel values and encoded into a 2D property map. In some example implementations, multiple 2D property maps can be constructed to store all the necessary attribute information for the vertices and sampled surface points of the mesh.
[0035] The 2D graph above is merely an example. Other types of 2D graphs can be constructed for the mesh. Additionally, other datasets can be extracted from the 3D mesh to complement the 2D graph described above. Figure 1Together, they represent the original 3D mesh. For example, in addition to the 2D graph, the connections or connectivity information between vertices can be grouped and organized separately in the form of lists, tables, etc. For instance, connectivity information can use vertex indices to refer to vertices. Vertex indices can be mapped to their corresponding pixel locations in the 2D graph. As another example, surface texture, color, normals, displacement, and other information can be extracted and organized separately outside of the 2D graph, rather than as part of the 2D graph. Further metadata can be extracted from the 3D mesh to combine the aforementioned 2D graph and other datasets to represent the 3D mesh.
[0036] While the exemplary embodiments described above focus on static meshes, according to one aspect of this disclosure, 3D meshes can be dynamic. For example, a dynamic mesh can refer to a mesh in which at least one component (geometric information, connectivity information, mapping information, vertex attributes, and attribute graph) changes over time. Thus, a dynamic mesh can be described by a sequence of meshes or meshes (also called mesh frames), similar to a timing sequence of 2D image frames forming a video.
[0037] In some exemplary implementations, dynamic meshes can have constant connectivity information, time-varying geometry, and time-varying vertex properties. In some other examples, dynamic meshes can have time-varying connectivity information. In some examples, digital 3D content creation tools can be used to generate dynamic meshes with time-varying property maps and time-varying connectivity information. In some other examples, volumetric acquisition / detection / sensing techniques are used to generate dynamic meshes. Volumetric acquisition techniques can generate dynamic meshes with time-varying connectivity information, especially under real-time constraints.
[0038] Because dynamic meshes can contain a large amount of information that changes over time, they may require a significant amount of data. However, compression can be performed to utilize redundancy within mesh frames (intra-frame compression) and between mesh frames (inter-frame compression). Various mesh compression processes can be implemented to allow for the efficient storage and transmission of media content in mesh representations, particularly for mesh sequences.
[0039] This disclosure provides example architectures and techniques for mesh compression. These techniques can be used for various mesh compression methods, including but not limited to static mesh compression, dynamic mesh compression, compression of dynamic meshes with constant connectivity information, compression of dynamic meshes with time-varying connectivity information, compression of dynamic meshes with time-varying property graphs, etc. These techniques can be used for lossy and lossless compression in various applications, such as real-time immersive communication, storage, free-viewpoint video, augmented reality (AR), virtual reality (VR), etc. These applications may include features such as random access and scalable / progressive encoding / decoding.
[0040] While this disclosure explicitly describes techniques and implementations applicable to 3D meshes, the principles underlying the various implementations described herein can also be applied to other types of 3D data structures, including but not limited to point cloud (PC) data structures. For simplicity, the following references to 3D meshes are intended to be general and include other types of 3D representations, such as point clouds and other 3D volumetric datasets.
[0041] First, let's turn to an exemplary architecture-level implementation. Figure 1 A simplified block diagram of a communication system (100) according to an example embodiment of the present disclosure is illustrated. The communication system (100) may include a plurality of terminal devices capable of communicating with each other via, for example, a communication network (150) (alternately referred to as a network). For example, the communication system (100) may include a pair of terminal devices (110) and (120) interconnected via the network (150). Figure 1 In the example, the first pair of terminal devices (110) and (120) can perform unidirectional transmission of 3D meshes. For example, terminal device (110) can compress 3D meshes or sequences of 3D meshes, which may be generated by terminal device (110), obtained from storage, or captured by a 3D sensor (105) connected to terminal device (110). The compressed 3D meshes or sequences of 3D meshes can be transmitted via network (150) to another terminal device (120), for example, in the form of a bitstream (also called an encoded bitstream). Terminal device (120) can receive the compressed 3D meshes or sequences of 3D meshes from network (150), decompress the bitstream to reconstruct the original 3D meshes or sequences of 3D meshes, and process the reconstructed 3D meshes or sequences of 3D meshes appropriately for display or for other purposes / uses. Unidirectional data transmission may be common in media service applications, etc.
[0042] exist Figure 1 In the examples, one or both of the terminal devices (110) and (120) may be implemented as a server, a fixed or mobile personal computer, a laptop computer, a tablet computer, a smartphone, a gaming terminal, a media player, and / or a dedicated three-dimensional (3D) device, etc., but the principles of this disclosure are not limited thereto. The network (150) may represent any type of network or combination of networks transmitting compressed 3D meshes between the terminal devices (110) and (120). The network (150) may include, for example, wired (wired) and / or wireless communication networks. The network (150) may exchange data in circuit-switched and / or packet-switched channels. Representative networks include long-distance telecommunications networks, local area networks, wide area networks, cellular networks, and / or the Internet. For the purposes of this disclosure, the architecture and topology of the network (150) may be irrelevant to the operation of this disclosure, unless explained below.
[0043] Figure 2 A simplified example block diagram of a streaming system (200) according to an embodiment of the present disclosure is shown. Figure 2 The illustrations depict example applications of the disclosed implementations related to 3D meshes and compressed 3D meshes. The disclosed topics are equally applicable to other applications that support 3D meshes or point clouds, such as 3D telepresence applications, virtual reality applications, etc.
[0044] The streaming system (200) may include an acquisition or storage subsystem (213). The acquisition or storage subsystem (213) may include a 3D mesh generator or storage medium (201) that generates or provides an uncompressed 3D mesh (202) or point cloud (202), such as 3D mesh or point cloud generation tools / software, graphics generation components, or point cloud sensors, such as light detection and ranging (LIDAR) systems, 3D cameras, 3D scanners, 3D mesh storage, etc. In some exemplary embodiments, the 3D mesh (202) includes the vertices of a 3D mesh or 3D points of a point cloud (both referred to as a 3D mesh). The 3D mesh (202) is depicted as thick lines to emphasize the correspondingly high data volume when compared to a compressed 3D mesh (204) (the bitstream of the compressed 3D mesh). The compressed 3D mesh (204) may be generated by an electronics device (220) including an encoder (203) coupled to the 3D mesh (202). The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter as described in more detail below. A compressed 3D mesh (204) (or a stream of compressed 3D mesh (204)) may be stored in a streaming server (205) for future use, the compressed 3D mesh (204) being depicted as thin lines to emphasize a lower data volume compared to a stream of uncompressed 3D mesh (202). Such as Figure 2 One or more streaming client subsystems (206) and (208) can access the streaming server (205) to retrieve copies (207) and (209) of the compressed 3D mesh (204). The client subsystem (206) may include a decoder (210) in, for example, an electronic device (230). The decoder (210) may be configured to decode the incoming copy (207) of the compressed 3D mesh and create an output stream of a reconstructed 3D mesh (211) that can be rendered on a rendering device (212) or used for other purposes.
[0045] Note that electronic devices (220) and (230) may include other components (not shown). For example, electronic device (220) may include a decoder (not shown), and electronic device (230) may also include an encoder (not shown).
[0046] In some streaming systems, compressed 3D meshes (204), (207), and (209) (e.g., a stream of compressed 3D meshes) can be compressed according to certain standards. In some examples, as described in further detail below, video codec standards are used to utilize redundancy and correlation in the compression of 3D meshes after first projecting the 3D meshes to map them into a 2D representation suitable for video compression. Non-limiting examples of these standards include High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), etc., as described in further detail below.
[0047] A compressed 3D mesh or 3D mesh sequence can be generated by an encoder, while a decoder can be configured to decompress the compressed or encoded 3D mesh. Figure 3 The illustration shows a high-level exemplary data flow of a 3D mesh in such an encoder (301) and decoder (303). Figure 3As shown, the original input 3D mesh or 3D mesh sequence (302) can be preprocessed by orbital remeshing, parameterization, and / or voxelization to generate input data to a mapping unit used to map the 3D mesh to a 2D UV space (304). In some embodiments, the 2D UV space (304) may include a mesh with a UV atlas. The 3D mesh can be sampled to include 3D surface points that may not be in the vertices, and these sampled 3D surface points are added to the mapping in the UV space. Various 2D graphs can be generated in encoder 301, including but not limited to occupancy graphs (310), geometry graphs (312), and attribute graphs (314). These image-type graphs can be compressed by encoder 301 using, for example, video encoding / decoding / compression techniques. For example, a video encoder can help compress 3D mesh frames by using intra-frame prediction techniques and inter-frame predictions performed by other 3D mesh reference frames. Other non-image or non-graph data or metadata (316) can also be encoded in various ways to remove redundancy, thereby generating compressed non-graph data via entropy encoding (as a non-limiting example). The encoder 301 can then combine or multiplex the compressed 2D graph and non-graph data, and further encode the combined data to generate an encoded bitstream (or coded bitstream). The encoded bitstream can then be stored or transmitted for use by the decoder 303. The decoder can be configured to decode the bitstream, demultiplex the decoded bitstream to obtain compressed 2D graph and non-graph data, and pre-decompress to generate a decoded occupancy graph (320), a decoded geometry graph (322), a decoded attribute graph (324), and decoded non-graph data and metadata (326). The decoder 303 can then be further configured to reconstruct a 3D mesh or a sequence of 3D meshes (330) from the decoded 2D graphs (320, 322, and 324) and the decoded non-graph data (326).
[0048] More in detail, Figure 4 A block diagram of an example 3D mesh encoder (400) for encoding 3D mesh frames according to some embodiments of the present disclosure is shown. In some exemplary embodiments, the mesh encoder (400) can be used in a communication system (100) and a streaming system (200). For example, an encoder (203) can be configured and operated in a similar manner to the mesh encoder (400).
[0049] The mesh encoder (400) can receive 3D mesh frames as uncompressed input and generate a bitstream corresponding to compressed 3D mesh frames. In some example implementations, the mesh encoder (400) can receive 3D mesh frames from any source, such as from... Figure 2 Received from grid or point cloud sources (201), etc.
[0050] exist Figure 4In the example, the mesh encoder (400) may include a patch generation module (406) (alternatively referred to as a graph generation module), a patch packaging module (408), a geometry image generation module (410), a texture image generation module (412), a patch information module (404), an occupancy graph module (414), a smoothing module (436), image filling modules (416) and (418), a group expansion module (420), video compression modules (422), (423) and (432), an auxiliary patch information compression module (438), an entropy compression module (434), and a multiplexer (424).
[0051] In various embodiments of this disclosure, a module can refer to a software module, a hardware module, or a combination thereof. A software module may include a computer program or part of a computer program that has predetermined functions and works with other related parts to achieve predetermined objectives (such as those functions described in this disclosure). A hardware module may be implemented using processing circuitry and / or memory configured to perform the functions described in this disclosure. Each module may be implemented using one or more processors (or processors and memory). Similarly, a processor (or multiple processors and memory) may be used to implement one or more modules. Furthermore, each module may be part of an entire module that includes the module's functions. The description herein may also be applied to the term "module" and other equivalent terms (e.g., "unit").
[0052] According to one aspect of this disclosure, and as described above, a mesh encoder (400) converts a 3D mesh frame along with some non-graphical metadata (e.g., patch or graph information) into an image-based representation (e.g., a 2D graph), the non-graphical metadata being used to assist in converting the compressed 3D mesh back into a decompressed 3D mesh. In some examples, the mesh encoder (400) may convert the 3D mesh frame into a 2D geometry or image, a texture or image, and an occupancy or image, and then use video encoding / decoding techniques to encode the geometry, texture, and occupancy images along with metadata and other compressed non-graphical data into a bitstream. Typically, and as described above, a 2D geometry image is a 2D image having 2D pixels filled with geometric values associated with 3D points projected onto the 2D pixels (the term "projection" is used to denote "mapping"), and the 2D pixels filled with geometric values may be referred to as geometry samples. A texture image is a 2D image with pixels filled with texture values associated with 3D points projected onto the 2D pixels, and the 2D pixels filled with texture values may be referred to as texture samples. An occupancy map is a 2D image with 2D pixels filled with values indicating whether a 3D point is occupied or not.
[0053] The patch generation module (406) divides the 3D mesh into a set of graphs or patches that may overlap or not overlap (e.g., a patch is defined as a continuous subset of surfaces described by the 3D mesh or point cloud), such that each patch can be described by a depth field relative to a plane in 2D space (e.g., surface flattening causes deeper 3D points on the surface to be further away from the center of the corresponding 2D graph). In some embodiments, the patch generation module (406) aims to decompose the 3D mesh into a minimum number of patches with smooth boundaries while also minimizing reconstruction errors.
[0054] The patch information module (404) can collect patch information indicating the size and shape of the patch. In some examples, the patch information can be packaged into a data frame and then encoded by the auxiliary patch information compression module (438) to generate compressed auxiliary patch information. Auxiliary patch compression can be implemented in various forms, including but not limited to various types of arithmetic encoding and decoding.
[0055] The patch or chart packaging module (408) can be configured to map extracted patches onto a 2D grid in UV space while minimizing unused space. In some exemplary embodiments, the pixels of the 2D UV space can be granularized into pixel blocks for mapping patches or charts. The block size can be predefined. For example, the block size can be M×M (e.g., 16×16). With this granularity, it can be ensured that each M×M block of the 2D UV grid is associated with a unique patch. In other words, each patch is mapped to a 2D UV space with a 2D granularity of M×M. Efficient patch packaging can directly impact compression efficiency by minimizing unused space or ensuring time consistency. Exemplary implementations of packaging patches or charts into 2D UV space are given in more detail below.
[0056] The geometry image generation module (410) can generate a 2D geometry image associated with the geometry of a 3D mesh at a given patch location in a 2D grid. The texture image generation module (412) can generate a 2D texture image associated with the texture of the 3D mesh at a given patch location in a 2D grid. As described above, the geometry image generation module (410) and the texture image generation module (412) essentially utilize the 3D-to-2D mapping calculated during the above-described packaging process to store the geometry and texture of the 3D mesh as a 2D image.
[0057] In some implementations, to better handle situations where multiple points are projected onto the same sample (e.g., patches overlap in the 3D space of a grid), the 2D image can be layered. In other words, each patch can be projected onto two images, for example, called layers, such that multiple points can be projected onto the same point in different layers.
[0058] In some exemplary embodiments, the geometric image can be represented by a monochrome frame of width x height (W×H). Therefore, three geometric images with three luminance or chrominance channels can be used to represent 3D coordinates. In some exemplary embodiments, the geometric image can be represented by a 2D image with three channels (RGB, YUV, YCrCb, etc.) having a specific color depth (e.g., 8-bit, 12-bit, 16-bit, etc.). Therefore, one geometric image with three color channels can be used to represent 3D coordinates.
[0059] To generate texture images, the texture generation process utilizes reconstructed / smoothed geometry to compute colors associated with sampling points from the original 3D mesh (see [link to original text]). Figure 3 The "sampling" will, for example, generate 3D surface points that are not in the vertices of the original 3D mesh.
[0060] The occupancy map module (414) can be configured to generate an occupancy map describing the filling information at each cell. For example, as described above, the occupancy map may include a binary map indicating whether each cell of a 2D raster belongs to empty space or to a 3D mesh. In some exemplary embodiments, the occupancy map may use binary information to describe whether each pixel is filled. In some other exemplary embodiments, the occupancy map may use binary information to describe whether each pixel block (e.g., each M×M block) is filled.
[0061] The occupancy graph generated by the occupancy graph module (414) can be compressed using either lossless or lossy encoding / decoding. When using lossless encoding / decoding, the entropy compression module (434) can be used to compress the occupancy graph. When using lossy encoding / decoding, the video compression module (432) can be used to compress the occupancy graph.
[0062] Note that the patch packing module (408) may leave some blank space between 2D patches that are packed into image frames. Image padding modules (416) and (418) may fill the blank space (referred to as padding) to generate image frames suitable for 2D video and image codecs. Image padding, also known as background padding, fills unused space with redundant information. In some examples, well-implemented background padding increases the bit rate minimally while avoiding the introduction of significant codec distortion around patch boundaries.
[0063] The video compression modules (422), (423), and (432) can encode 2D images (such as filled geometric images, filled texture images, and occupancy maps) based on suitable video codec standards (such as HEVC, VVC, etc.). In some exemplary embodiments, the video compression modules (422), (423), and (432) are separate components that operate independently. Note that in some other exemplary embodiments, the video compression modules (422), (423), and (432) can be implemented as a single component.
[0064] In some example implementations, the smoothing module (436) can be configured to generate a smoothed image of the reconstructed geometry. The smoothed image can be provided to the texture image generation (412). The texture image generation (412) can then adjust the generation of the texture image based on the reconstructed geometry. For example, when the patch shape (e.g., geometry) is slightly distorted during encoding and decoding, this distortion can be taken into account when generating the texture image to correct for the distortion in the patch shape.
[0065] In some embodiments, group expansion (420) is configured to fill pixels around object boundaries with redundant low-frequency content in order to improve encoding / decoding gain and the visual quality of the reconstructed 3D mesh.
[0066] The multiplexer (424) can be configured to multiplex compressed geometric images, compressed texture images, compressed occupancy maps, and compressed auxiliary patch information into a compressed bitstream (or encoded bitstream).
[0067] Figure 5 A block diagram of an exemplary mesh decoder (500) for decoding a compressed bitstream corresponding to a 3D mesh frame, according to some embodiments of the present disclosure, is shown. In some exemplary embodiments, the mesh decoder (500) may be used in a communication system (100) and a streaming system (200). For example, a decoder (210) may be configured to operate in a manner similar to the mesh decoder (500). The mesh decoder (500) receives a compressed bitstream and generates a reconstructed 3D mesh based on the compressed bitstream, which includes, for example, compressed geometric images, compressed texture images, compressed occupancy maps, and compressed auxiliary patch information.
[0068] exist Figure 5 In the example, the mesh decoder (500) may include a demultiplexer (532), a video decompression module (534) and (536), a occupancy graph decompression module (538), an auxiliary patch information decompression module (542), a geometry reconstruction module (544), a smoothing module (546), a texture reconstruction module (548), and a color smoothing module (552).
[0069] The demultiplexer (532) can receive the compressed bitstream and separate it into a compressed texture image, a compressed geometric image, a compressed occupancy map, and compressed auxiliary patch information.
[0070] Video decompression modules (534) and (536) can decode compressed images and output decompressed images according to suitable standards (e.g., HEVC, VVC, etc.). For example, video decompression module (534) can decode compressed texture images and output decompressed texture images. Video decompression module (536) can further decode compressed geometric images and output decompressed geometric images.
[0071] The occupancy map decompression module (538) can be configured to decode the compressed occupancy map according to a suitable standard (e.g., HEVC, VVC, etc.) and output the decompressed occupancy map.
[0072] The auxiliary patch information decompression module (542) can be configured to decode the compressed auxiliary patch information according to a suitable decoding algorithm and output the decompressed auxiliary patch information.
[0073] The geometry reconstruction module (544) can be configured to receive a decompressed geometry image and generate a reconstructed 3D mesh geometry based on the decompressed occupancy map and decompressed auxiliary patch information.
[0074] The smoothing module (546) can be configured to smooth inconsistencies at patch edges. The smoothing process can be designed to mitigate potential discontinuities that may occur at patch boundaries due to compression artifacts. In some example implementations, a smoothing filter can be applied to pixels located on patch boundaries to mitigate distortion that may be caused by compression / decompression.
[0075] The texture reconstruction module (548) can be configured to determine the texture information of points in a 3D mesh based on the decompressed texture image and smoothed geometry.
[0076] The color smoothing module (552) can be configured to smooth out inconsistencies in shading. In 2D video, non-adjacent patches in 3D space are often packed next to each other. In some examples, pixel values from non-adjacent patches can be confused by block-based video codecs. The goal of color smoothing can be to reduce visible artifacts appearing at patch boundaries.
[0077] Figure 6 A block diagram of an example video decoder (610) according to an embodiment of the present disclosure is shown. The video decoder (610) may be used in a grid decoder (500). For example, video decompression modules (534) and (536), and occupancy graph decompression module (538) may be similarly configured as video decoders (610).
[0078] The video decoder (610) may include a parser (620) to reconstruct symbols (621) from a compressed image, such as an encoded video sequence. The categories of these symbols may include information for managing the operation of the video decoder (610). The parser (620) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on video coding techniques or standards and may follow various principles, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (620) may extract a set of subgroup parameters from the encoded video sequence for use in the video decoder, based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (620) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.
[0079] The parser (620) can perform entropy decoding / parsing operations on the image sequence received from the buffer memory to create symbols (621).
[0080] Depending on the type of encoded video frames or a subset of encoded video frames (e.g., inter-frame and intra-frame frames, inter-frame and intra-frame blocks) and other factors, the reconstruction of the symbol (621) may involve multiple different units. Which units are involved and how they are involved can be controlled by the subgroup control information parsed from the encoded video sequence by the parser (620). For brevity, the flow of such subgroup control information between the parser (620) and the various units described below is not described.
[0081] In addition to the functional blocks already mentioned, the video decoder (510) can be conceptually subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with one another. However, for the purpose of describing the disclosed subject matter only, they are conceptually subdivided into the functional units described below.
[0082] The video decoder (610) may include a scaler / inverse transform unit (651). The scaler / inverse transform unit (651) receives quantization transform coefficients as symbols (621) and control information from the parser (620), including the transform method used, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (651) may output blocks including sample values, which may be input to the aggregator (655).
[0083] In some cases, the output samples of the scaler / inverse transform unit (651) may belong to intra-coded blocks that do not use predictive information from previously reconstructed images, but can use predictive information from previously reconstructed portions of the current image. Such predictive information may be provided by the intra-picture prediction unit (652). In some cases, the intra-picture prediction unit (652) uses reconstructed information extracted from the current picture buffer (658) to generate surrounding blocks of the same size and shape as the block being reconstructed. For example, the current picture buffer (658) buffers partially reconstructed and / or fully reconstructed current images. In some cases, the aggregator (655) adds the predictive information generated by the intra-picture prediction unit (652) to the output sample information provided by the scaler / inverse transform unit (651) based on each sample.
[0084] In other cases, the output samples of the scaler / inverse transform unit (651) may belong to inter-frame coding and latent motion compensation blocks. In this case, the motion compensation prediction unit (653) can access the reference image memory (657) to extract samples for prediction. After motion compensation of the extracted samples according to the symbols (621), these samples can be added by the aggregator (655) to the output of the scaler / inverse transform unit (651) (referred to in this case as residual samples or residual signals) to generate output sample information. The motion compensation prediction unit (653) can obtain the predicted samples from the address in the reference image memory (657) under motion vector control, and the motion vector is available to the motion compensation prediction unit (653) in the form of the symbols (621), which, for example, include X, Y and reference image components. Motion compensation may also include interpolation of sample values extracted from the reference image memory (657) when using subsample precise motion vectors, motion vector prediction mechanisms, etc.
[0085] The output samples of the aggregator (655) can be employed by various loop filtering techniques in the loop filter unit (656). Video compression techniques may include in-loop filtering techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video stream), and these parameters can be used as symbols (621) from the parser (620) in the loop filter unit (656). Video compression techniques may also respond to metadata obtained during decoding of a previous (in decoding order) portion of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0086] The output of the loop filter unit (656) can be a sample stream, which can be output to a display device and stored in a reference image memory (657) for subsequent inter-frame image prediction.
[0087] Once fully reconstructed, certain encoded images can be used as reference images for future predictions. For example, once the encoded images corresponding to the current image have been fully reconstructed and the encoded images (by, for example, the parser (620)) are identified as reference images, the current image buffer (658) can become part of the reference image memory (657), and a new current image buffer can be reallocated before the reconstruction of subsequent encoded images begins.
[0088] The video decoder (610) can perform decoding operations according to a predetermined video compression technique, such as that specified in the ITU-T H.265 standard. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the configuration file recorded in the video compression technique or standard. Specifically, the configuration file may select certain tools from all available tools in the video compression technique or standard as the only tools available under said configuration file. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technique or standard. In some cases, the hierarchy limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in megasamples per second, for example), maximum reference picture size, etc. In some cases, the limitations set by the hierarchy can be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata managed by the HRD buffer, which is represented by signals in the encoded video sequence.
[0089] Figure 7A block diagram of a video encoder (703) according to an embodiment of the present disclosure is shown. The video encoder (703) can be used in a mesh encoder (400) for compressing 3D meshes or point clouds. In some exemplary embodiments, video compression modules (422) and (423) and video compression module (432) are configured similarly to encoder (703).
[0090] The video encoder (703) can receive 2D images (such as filled geometric images, filled texture images, etc.) and generate compressed images.
[0091] According to an embodiment, the video encoder (703) can encode and compress images of a source video sequence (images) into an encoded video sequence (compressed images) in real time or under any other time constraints required by the application. Implementing an appropriate encoding rate is a function of the controller (750). In some embodiments, the controller (750) controls and is functionally coupled to other functional units described below. For simplicity, coupling is not shown in the figures. Parameters set by the controller (750) may include rate control related parameters (image skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), image size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (750) may be used with other suitable functions related to the video encoder (703) optimized for a particular system design.
[0092] In some embodiments, the video encoder (703) operates within an encoding loop. As a simplified description, in an embodiment, the encoding loop may include a source encoder (730) (e.g., responsible for creating symbols, such as a symbol stream, based on the input image to be encoded and a reference image) and a (local) decoder (733) embedded within the video encoder (703). The decoder (733) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data (since any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques considered in the disclosed subject matter). The reconstructed sample stream (sample data) is input to a reference image memory (734). Since decoding of the symbol stream produces bit-precise results independent of the decoder's location (local or remote), the contents of the reference image memory (734) also correspond bit-precisely between the local and remote encoders. In other words, the reference image samples "seen" by the encoder's prediction portion are exactly the same sample values that the decoder will "see" when using the prediction during decoding. This fundamental principle of reference image synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.
[0093] The operation of the “local” decoder (733) can be combined with, for example, the above-described method. Figure 6 The video decoder (610) is described in detail as the same as the "remote" decoder. However, a further brief reference is provided. Figure 6 When symbols are available and the entropy encoder (745) and parser (620) are able to encode / decode the symbols into an encoded video sequence without loss, the entropy decoding part of the video decoder (610), including the parser (620), may not be fully implemented in the local decoder (733).
[0094] In the various embodiments of this application, any decoder technique, other than the parsing / entropy decoding present in the decoder, may need to exist in the corresponding encoder in essentially the same functional form. For this reason, the embodiments of this application focus on decoder operation. The description of the encoder technique can be simplified because the encoder technique is the inverse of the fully described decoder technique. More detailed description is only required in certain areas, and is provided below.
[0095] During operation, in some embodiments, the source encoder (730) may perform motion-compensated predictive coding. The motion-compensated predictive coding predictively encodes the input image, referencing one or more previously encoded images from the video sequence designated as "reference images." In this manner, the encoding engine (732) encodes the differences between pixel blocks of the input image and pixel blocks of the reference image, which may be selected as a predictive reference for the input image.
[0096] The local video decoder (733) can decode encoded video data of a picture that can be designated as a reference picture, based on symbols created by the source encoder (730). The operation of the encoding engine (732) can be a lossy process. When the encoded video data can be decoded by the video decoder (730), Figure 7 When the source video sequence (not shown) is decoded, the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (733) replicates the decoding process, which can be performed by the video decoder on the reference image, and allows the reconstructed reference image to be stored in a reference image cache (734). In this way, the video encoder (703) can locally store a copy of the reconstructed reference image that shares the same content (no transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0097] The predictor (735) can perform a prediction search against the encoding engine (732). That is, for a new image to be encoded, the predictor (735) can search in the reference image memory (734) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference image motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new image. The predictor (735) can operate pixel-by-pixel based on the sample blocks to find suitable prediction references. In some cases, based on the search results obtained by the predictor (735), it can be determined that the input image may have prediction references obtained from multiple reference images stored in the reference image memory (734).
[0098] The controller (750) can manage the encoding operations of the source encoder (730), including, for example, setting parameters and subgroup parameters for encoding video data.
[0099] The outputs of all the above-mentioned functional units can be entropy encoded in the entropy encoder (745). The entropy encoder (745) performs lossless compression on the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, and arithmetic coding, thereby converting the symbols into an encoded video sequence.
[0100] The controller (750) manages the operation of the video encoder (703). During encoding, the controller (750) can assign a specific encoded image type to each encoded image, but this may affect the encoding techniques applicable to the corresponding images. For example, images can typically be assigned to any of the following image types:
[0101] An intra-frame picture (I-picture) is a picture that can be encoded and decoded without using any other pictures in the sequence as a prediction source. Some video codecs allow different types of intra-frame pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will understand variations of I-pictures and their corresponding applications and characteristics.
[0102] A predictive image (P-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0103] A bidirectional predictive image (B-image) can be an image that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses at most two motion vectors and a reference index to predict sample values for each block. Similarly, multiple predictive images can use more than two reference images and associated metadata to reconstruct a single block.
[0104] Source images are typically spatially subdivided into multiple sample blocks (e.g., 4×4, 8×8, 4×8, or 16×16 sample blocks), and encoded block by block. These blocks can be predictively coded with reference to other (already coded) blocks, determined based on the coding assignments of the corresponding images applied to the blocks. For example, blocks of an I-image can be non-predictively coded, or the blocks can be predictively coded (spatial or intra-frame prediction) with reference to already coded blocks of the same image. Pixel blocks of a P-image can be predictively coded with reference to a previously coded reference image via spatial or temporal prediction. Blocks of a B-image can be predictively coded with reference to one or two previously coded reference images via spatial or temporal prediction.
[0105] The video encoder (703) can perform encoding operations according to a predetermined video coding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (703) can perform various compression operations, including predictive coding operations that utilize temporal and spatial redundancy in the input video sequence. Therefore, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0106] Video can be presented as a time-series of multiple source images. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes spatial correlations within a given image, while inter-frame image prediction utilizes (temporal or other) correlations between images. In an embodiment, a specific image being encoded / decoded is segmented into blocks, referred to as the current image. When a block in the current image resembles a reference block in a previously encoded and still buffered reference image in the video, the block in the current image can be encoded using a vector called a motion vector. This motion vector points to the reference block in the reference image, and when multiple reference images are used, the motion vector may have a third dimension that identifies the reference image.
[0107] In some embodiments, bidirectional prediction techniques can be used in inter-frame image prediction. According to bidirectional prediction, two reference images are used, such as a first reference image and a second reference image, both preceding the current image in the video in decoding order (but possibly past and future in display order). A block in the current image can be encoded using a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted using a combination of the first and second reference blocks.
[0108] In addition, merging mode techniques can be used in inter-frame image prediction to improve coding efficiency.
[0109] According to some embodiments disclosed in this application, predictions such as inter-frame image prediction and intra-frame image prediction are performed on a block-by-block basis. For example, according to the HEVC standard, images in a video image sequence are segmented into coding tree units (CTUs) for compression. The CTUs in the images have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs): one luma CTB and two chroma CTBs. Furthermore, each CTU can be further subdivided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be subdivided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In embodiments, each CU is analyzed to determine the prediction type used for the CU, such as inter-frame prediction or intra-frame prediction. Furthermore, depending on temporal and / or spatial predictability, the CU is divided into one or more prediction units (PUs). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In embodiments, prediction operations in encoding (encoding / decoding) are performed on a per-prediction-block basis. Taking a luma prediction block as an example, a prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0110] In various embodiments, the mesh encoder (400) and mesh decoder (500) described above can be implemented in hardware, software, or a combination thereof. For example, the mesh encoder (400) and mesh decoder (500) can be implemented using processing circuitry such as one or more integrated circuits (ICs) that operate with or without software, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc. In another example, the mesh encoder (400) and mesh decoder (500) can be implemented as software or firmware comprising instructions stored in a non-permanent (or non-volatile) computer-readable storage medium. When executed by processing circuitry such as one or more processors, these instructions cause the processing circuitry to perform the functions of the mesh encoder (400) and / or mesh decoder (500).
[0111] Figure 8An example is shown of mapping a 3D patch (810) in a 3D xyz coordinate system (819) to a 2D UV plane (850) in a uv coordinate system (859). In some implementations, a 3D patch (or simply a patch) can generally refer to a continuous subset of a surface described by a set of vertices corresponding to a mesh in 3D space. In a non-limiting example, a patch includes vertices having 3D coordinates, normal vectors, color, texture, and other information. In some implementations, a 2D geometric patch (or simply a geometric patch) can refer to a projected shape in the 2D UV plane, a projected shape corresponding to the patch, and a geometric patch corresponding to the patch.
[0112] In the projected shape of the UV plane, each mapped point (ui, vi) corresponds to a 3D vertex with position (xi, yi, zi), where i = 1, 2, 3, 4, etc. For example, the first vertex (811) with 3D coordinates (x1, y1, z1) is mapped to the first point (851) with 2D coordinates (u1, v1); the second vertex (812) with 3D coordinates (x2, y2, z2) is mapped to the second point (852) with 2D coordinates (u2, v2); the third vertex (813) with 3D coordinates (x3, y3, z3) is mapped to the third point (853) with 2D coordinates (u3, v3); and the fourth vertex (814) with 3D coordinates (x4, y4, z4) is mapped to the fourth point (854) with 2D coordinates (u4, v4).
[0113] Then, the encoding and decoding of the vertex geometry (i.e., the values of xi, yi, and zi) is converted into encoding 3-channel values in a 2D plane, where each 3-channel value at the (u, v) location corresponds to an associated 3D location. For example, the pixel value of the first location (851) includes a 3-channel value (or three channel values) corresponding to the 3D coordinates of the first vertex (811), where the first channel value of the 3-channel value corresponds to the value of x1, the second channel value of the 3-channel value corresponds to the value of y1, and the third channel value of the 3-channel value corresponds to the value of z1.
[0114] Therefore, a 2D plane with one or more patches of projection / mapping is called a geometric image, and can therefore be encoded using any image or video codec (e.g., a video codec that supports the 4:4:4 color format).
[0115] In some implementations, the pixel values of projected / mapped points on the UV plane can correspond to the distance from the 3D vertex to the UV plane. Therefore, multiple planes in different directions can be used for this projection to locate vertex position information. In this case, each location (or point) in the UV plane is a 1-channel value recording the distance. This image is called a depth image. Depth images can be encoded using any image or video codec (e.g., a video codec that supports YUV4:2:0 or YUV4:0:0 color formats).
[0116] In some implementations, using pixel values of points in the UV plane to represent the 3D coordinates of vertices in a 3D mesh can present problems / challenges. One problem / challenges is that the dynamic range of one or more components of a vertex's 3D coordinates may be greater than the range of pixel values with a specific bit depth. For example, the bit depth of sampled values on a 2D UV plane may need to match the dynamic range of the geometry of a 3D patch. Sometimes this dynamic range can be very large, so the required bit depth from the video codec is higher than the maximum bit depth supported by some video codecs. In other words, not every video codec can support high bit depth encoding to match the very large dynamic range of the 3D coordinates of vertices in the mesh. Another problem / challenges may include the fact that even if a video codec supports high bit depth, it may not be necessary to allocate a high bit depth to each patch because not every patch has a high dynamic range in its geometric distribution.
[0117] This disclosure describes various embodiments for encoding or decoding 3D meshes with limited geometric dynamic range, solving at least one of the aforementioned problems / challenges, thereby improving the efficiency of compressing geometric information and / or advancing efficient 3D mesh compression techniques.
[0118] Figure 9 A flowchart 900 illustrates an example method for decoding a 3D mesh with a limited geometric dynamic range, following the principles upon which the above-described embodiments are based. The example decoding method begins at S901 and may include some or all of the following steps: in S910, receiving an encoded bitstream comprising a geometric patch of a 3D mesh; in S920, extracting a first syntax element from the encoded bitstream indicating whether the geometric patch is partitioned, wherein the geometric patch comprises one or more partitions; and in S930, for each partition in the geometric patch, obtaining the dynamic range of pixel values of points in each partition, wherein the points correspond to vertices in the 3D mesh, and the dynamic range enables encoding of the geometric patch within a predetermined bit depth. The example method terminates at S999.
[0119] In some implementations, another example method for decoding a 3D mesh with a limited geometric dynamic range may include: receiving an encoded bitstream; extracting a first syntax element from the encoded bitstream indicating whether a geometric patch is partitioned, wherein the geometric patch includes one or more partitions; for partitions in the geometric patch: obtaining extreme pixel values of points in the partition from the encoded bitstream, wherein the points in the partition correspond to vertices in the 3D mesh; obtaining encoded pixel values of the points in the partition from the encoded bitstream; and determining pixel values of the points in the partition based on the extreme pixel values and the encoded pixel values, wherein the pixel values correspond to the 3D coordinates of vertices in the 3D mesh.
[0120] In some implementations, the encoded bitstream includes at least one of the following: encoded geometry or encoded metadata. For example, the encoded bitstream may be... Figure 4 The compressed bitstream may include one or more compressed geometry images / graphs, one or more compressed texture images / graphs, one or more compressed occupancy maps, and / or compressed auxiliary patch information. Some encoded bitstreams may not have any occupancy maps because occupancy map information can be inferred from the decoder side when the boundary vertices of each patch are signaled. In some implementations, a geometry patch is one of the patches corresponding to the encoded geometry map.
[0121] In some implementations, when the first syntax element indicates that the geometry patch is partitioned, the geometry patch may include more than one partition.
[0122] In some implementations, when the first syntax element indicates that the geometry patch is not partitioned, the geometry patch may include a single partition as the geometry patch itself.
[0123] In some implementations, each partition includes one or more points in the UV plane, and each partition may include multiple partitions when the geometry patch is partitioned, or a single partition when the geometry patch is not partitioned. The points in the partition correspond to vertices in the 3D mesh; and the encoded pixel values of the points in the partition are associated with the 3D coordinates of the vertices in the 3D mesh.
[0124] In some implementations, the extreme pixel values of points within a partition can be offsets from the encoded pixel values within the partition. Based on the encoded pixel values and extreme pixel values of points within a partition, the pixel values of points within the partition corresponding to the 3D coordinates of vertices in the 3D mesh can be obtained.
[0125] In some implementations, the extreme pixel value of a point in a partition can be either the minimum pixel value or the maximum pixel value of a point in the partition. For a non-limiting example, when the pixel value has three channels (e.g., RGB), the minimum pixel value can include three minimum values for each channel, where the minimum pixel value is (r_min, g_min, b_min), where r_min is the minimum red channel value of all points in the partition, g_min is the minimum green channel value of all points in the partition, and b_min is the minimum blue channel value of all points in the partition.
[0126] In some implementations, in response to a first syntax element indicating that the geometry patch is not partitioned, the geometry patch includes a single partition. For example, in response to a first syntax element indicating that the geometry patch is not partitioned, the geometry patch always includes a single partition.
[0127] In some implementations, in response to a first syntax element indicating that a geometry patch is partitioned, the method further includes: extracting a second syntax element indicating the partition type of the geometry patch from the encoded bitstream by a device, and partitioning the geometry patch based on the partition type by the device to obtain multiple partitions.
[0128] In some implementations, for partitions among multiple partitions, the example decoding method may further include: extracting a third syntax element from the encoded bitstream by the device indicating whether the partition is further partitioned; and / or, in response to the third syntax element indicating that the partition is further partitioned, extracting a fourth syntax element from the encoded bitstream indicating the partition type of the partition; and further partitioning the partition based on the partition type of the partition.
[0129] In some implementations, in response to a first syntax element indicating that the geometry patch is not partitioned, the geometry patch includes a single partition.
[0130] In some implementations, in response to a first syntax element indicating that a geometry patch is partitioned, the method further includes: extracting a second syntax element indicating the partition depth from the encoded bitstream by a device, and / or partitioning the geometry patch based on the partition depth and a predefined partition type by the device to obtain multiple partitions.
[0131] In some implementations, the extreme pixel value of a point in a partition is the minimum pixel value; and determining the pixel value of a point in a partition based on the extreme pixel value and the encoded pixel value includes: adding the component of the minimum pixel value to the component of each pixel value in the encoded pixel value to obtain the pixel value.
[0132] In some implementations, the minimum pixel value includes {x_min, y_min, z_min}, and each encoded pixel value includes {x_coded, y_coded, z_coded}; and each pixel value includes {x, y, z}, which is determined according to x = x_min + x_coded, y = y_min + y_coded, and z = z_min + z_coded, where x, y, z, x_min, y_min, z_min, x_coded, y_coded, and z_coded are non-negative integers.
[0133] In some implementations, the example decoding method may further include obtaining the out-of-limit value of the out-of-limit point in the partition from the encoded bitstream by the device.
[0134] In some implementations, for an out-of-limit point, adding the component of the minimum pixel value and the component of each of the encoded pixel values to obtain the pixel value includes: adding the component of the out-of-limit value, the component of the minimum pixel value, and the component of the encoded pixel value corresponding to the out-of-limit point to obtain the pixel value of the out-of-limit point.
[0135] In some implementations, for an out-of-limit point, the component of the encoded pixel value corresponding to the non-zero component of the out-of-limit value is equal to 2N-1, where N is the bit depth of the encoded pixel value.
[0136] In some implementations, the extreme pixel value of a point in a partition is the maximum pixel value.
[0137] In some implementations, determining the pixel value of a point in a partition based on extreme pixel values and encoded pixel values includes subtracting a component of each encoded pixel value from the component of the maximum pixel value to obtain the pixel value.
[0138] In some implementations, the maximum pixel value includes {x_max, y_max, z_max}, and each encoded pixel value includes {x_coded, y_coded, z_coded}; and each pixel value includes {x, y, z}, wherein {x, y, z} is determined according to x = x_max - x_coded, y = y_max - y_coded, and z = z_max - z_coded, where x, y, z, x_max, y_max, z_max, x_coded, y_coded, and z_coded are non-negative integers.
[0139] In some implementations, the exemplary decoding method may further include obtaining the over-limit value of an over-limit point in a partition from the encoded bitstream by the device; and wherein, for an over-limit point, subtracting a component of each of the encoded pixel values from the component of the maximum pixel value to obtain a pixel value includes: subtracting the component of the over-limit value and a component of each of the encoded pixel values from the component of the maximum pixel value to obtain a pixel value.
[0140] In some implementations, for an out-of-limit point, the component of the encoded pixel value corresponding to the non-zero component of the out-of-limit value is equal to 2^N-1, and N is the bit depth of the encoded pixel value.
[0141] In some implementations, obtaining the extreme pixel values of points in a partition includes: obtaining the extreme pixel value difference corresponding to the extreme pixel values of points in the partition; and determining the extreme pixel values of points in the partition based on the extreme pixel value difference and the second extreme pixel value.
[0142] In some implementations, the exemplary decoding method may further include obtaining a second extreme pixel value from the encoded bitstream for a second partition or a second geometric patch.
[0143] The various steps in one or more embodiments or implementations may be applied individually or in any combination. The various embodiments of this disclosure can be applied to dynamic or static meshes. In a static mesh, there may be only one mesh frame or mesh content that does not change over time. The various embodiments of this disclosure can be extended to the encoding and decoding of depth images / attribute images / texture images, etc., where the dynamic range of sample values may exceed a certain threshold, such as 8 bits.
[0144] In some implementations, each graph / patch in the geometric image can be partitioned such that the dynamic range of each partition is within a certain dynamic range, so that the geometric image can be encoded by specifying a bit depth (such as 8 bits, 10 bits, etc.). Graphs / patches can be partitioned using one or more predefined patterns.
[0145] For non-restrictive examples, see [reference]. Figure 10A 2D geometry patch / graph can be partitioned using quadtree partitioning. A 2D geometry patch / graph has an arbitrary shape (1011) and a bounding box (1010). The dynamic range of this patch / graph may exceed the target dynamic range (e.g., 8 bits). Quadtree partitioning can be applied to the bounding box of the patch / graph (1010) to obtain a partition structure (1020) where the entire patch / graph has 4 partitions. Partitioning can be stopped when the dynamic range of each partition is within the target dynamic range (e.g., 8 bits). For example, partitioning of a partition (or sub-partition) (1021) can be stopped when it is within the target dynamic range.
[0146] When the dynamic range of each partition is not within the target dynamic range, the patch / graph can be further partitioned into smaller blocks until the dynamic range of each partition is within the target dynamic range (e.g., 8 bits). For example, a quadtree partition (1030) comprising 16 partitions (or subpartitions) is shown.
[0147] In some implementations, further partitioning may be applied only to partitions that are not within the target dynamic range, and further partitioning may not be applied to partitions that are within the target dynamic range.
[0148] This recursive partitioning algorithm may converge because each partition reduces the dynamic range. In the worst case, all or part of a patch / graph can be partitioned down to a single pixel (e.g., a minimal partition of size 1x1 containing a single pixel).
[0149] Figure 11 Some non-limiting examples are shown, including predefined partitioning patterns that allow recursive partitioning to form a partition tree. In one example, a portion or all of a 10-way partitioning structure or pattern may be predefined. Exemplary partitioning structures may include various 2:1 / 1:2 (in...) Figure 11 (in the third line) and 4:1 / 1:4 (in Figure 11 The first row shows rectangular partitions. The partition types with three sub-partitions, indicated as 1102, 1104, 1106, and 1108 in the second row, can be referred to as "T-shaped" partitions, specifically left T-shaped, top T-shaped, right T-shaped, and bottom T-shaped, respectively. In some exemplary embodiments, further subdivision is not permitted. Figure 11Any of the rectangular partitions. The partition depth can be further defined to indicate the depth of the partition from the patch or graph. In some implementations, only recursive partitioning of full square partitions in 1110 to the next level of the partition tree may be allowed. In other words, recursive partitioning may not be allowed for square partitions within a T-pattern. In some other examples, a quadtree structure can be used to partition a patch or a patch's partitions into quadtree partitions. This quadtree partitioning can be applied hierarchically and recursively to any square partition.
[0150] In some implementations, various partitioning structures / methods in video encoding / decoding can be applied to partition patches / graphs.
[0151] In some implementations of a mesh codec, one approach may include some or all of the following: first, signaling some syntax elements in the encoded bitstream to support the decoding process; second, altering the geometric image, wherein pixel values in each partition are subtracted from the minimum value in that partition, so that the image can be processed / encoded within a certain bit depth range.
[0152] In some implementations, the syntax elements for each patch / chart may include some or all of the following.
[0153] Signaling binary flags indicates whether a patch / graph has been partitioned. When the binary flags indicate that a patch / graph has been partitioned, signaling partitioning flags indicates how to partition the patch / graph. Partitioning can be applied in various forms, and the syntax elements associated with partitioning can be changed accordingly. For example, quadtrees, binary trees, and any other form of symmetric or asymmetric partitioning can be applied. Combinations of different partitioning structures can also be applied.
[0154] For the non-restrictive example, refer to the one where only quadtrees are applied. Figure 9 It can send a signal to notify the quadtree partition depth to indicate the partition mode.
[0155] In another non-restrictive example, partitions are applied only to areas where the dynamic range exceeds the limits, and not to areas where the dynamic range is within the limits. In this case, a flag can be signaled to each partition / region to indicate whether further partitioning is needed, and when that flag indicates that further partitioning is needed, another flag can be signaled to indicate its partition type.
[0156] In some implementations, the minimum value for each partition can be signaled, which includes the minimum x, y, and z values corresponding to the geometric image. When a patch / graph is not partitioned, the minimum value for the entire patch can be signaled. In this case, the entire patch / graph can be considered a single partition, and the minimum value for the entire patch can be the minimum value of a single partition.
[0157] In some implementations, some partitions can be empty (i.e., contain no sampling points) and the minimum value of empty patches / charts can be skipped. For example, Figure 10 The lower left partition (or subpartition) (1039) is an empty partition, and there is no need to signal the minimum value of partition (1039).
[0158] In some implementations, it is necessary to signal the minimum values (e.g., minimum x, y, z values for the geometry) of each partition in a patch / graph within the encoded bitstream. For each sample value in each channel of the UV plane, its dynamic range can be reduced by subtracting the signaled minimum value of the same patch or partition from the value of that channel, while still keeping the updated value non-negative. For example, let xm, ym, and zm be the minimum values of the x, y, and z coordinates in the partition, respectively. For any pixel value (x, y, z) in the patch, the updated value to be encoded can become (x - xm, y - ym, z - zm).
[0159] In some implementations, when multiple patches exist in the grid, a minimum value may exist for each patch. To signal these values, predictive encoding / decoding can be used, where the minimum value of one patch can be used as a predicted value for another patch. Encoding the difference between such minimum value pairs reduces the number of bits that need to be encoded into the bitstream and improves grid compression / decompression efficiency.
[0160] In some implementations, the signaling for the minimum value is performed on a per-defined region basis. The 2D geometric image is divided into several non-overlapping regions. For each region, the minimum value for each channel is signaled. This minimum value encoding / decoding can be accomplished through predictive encoding / decoding from previously encoded values, such as those from spatially adjacent regions. For example, the entire image can be divided into a 64x64 grid. For each 64x64 region, the minimum value (x, y, z) for that region is signaled. All samples within that region are updated by subtracting the minimum value before encoding / decoding.
[0161] In some implementations, the patch can be further divided into smaller sub-patches. For each sub-patch, the dynamic range of the sample values within it is within a certain dynamic range, and therefore can be supported by the video codec. An offset value for each channel is signaled, which is the minimum value of the sample channel in that sub-patch.
[0162] In some implementations, a special encoding / decoding process can be applied to isolated 3D points within the patch, where one or more channels have a dynamic range exceeding a given bit depth. The position of this point in the UV plane and the portion exceeding the limit are recorded and encoded separately.
[0163] In some implementations, for one or more points having a dynamic range exceeding a given bit depth, the original value may be replaced by the maximum possible value for the given bit depth. For example, when the bit depth is 8, the original value may be replaced by 255 (=2^8-1). For another example, when the bit depth is 12, the original value may be replaced by 4095 (=2^12-1). In some implementations, the final decoded value may be equal to the maximum value plus the excess portion.
[0164] For the unrestricted example, xm, ym, and zm can represent the minimum values of the x, y, and z coordinates in the partition, respectively; and the bit depth allowed for encoding the sample can be 8 bits. For the point (x1, y1, z1) in the patch, the updated value is (x1 - xm, y1 - ym, z1 - zm). When x1 - xm > 255 (= 2^8 - 1) and y1 - ym <= 255 and z1 - zm <= 255, and the projection position of this point in the UV plane is (u1, v1), in addition to normal UV plane encoding and decoding for each position in the UV plane at a depth of 8 bits, the position (u1, v1) along with additional out-of-limit information (x1 - xm - 255, 0, 0) is encoded into the bitstream. On the decoder side, after decoding the geometric image using the method described above, the additional value (x1 - xm - 255, 0, 0) should be decoded and added to the decoded 3-channel value at position (u1, v1) to recover the true geometric information of that point. That is, (x1 - xm - 255, 0, 0) + (255, y1 - ym, z1 - zm) = (x1 - xm, y1 - ym, z1 - zm). The original geometry (x1, y1, z1) can be recovered by adding the minimum value of each channel (xm, ym, zm) to further adjustments at each position.
[0165] In some implementations, when the dynamic range of the entire geometric image is known, such as [0, 2^m] and the bit depth of each patch is n and m > n, a conventional m-bit video codec can be used to encode the lower n bits of the dynamic range of that value with other conventional values. The higher (m-n) bits of the geometric information of that value can be encoded separately. For a non-limiting example with a bit depth of 8 bits, when the dynamic range of the entire geometric image is 14 bits, the dynamic range of the higher bits of information may be limited to a 6-bit range since 6 = 14 - 8.
[0166] In some implementations, maximum values can be used for encoding / decoding instead of minimum values. The maximum value for each partition within a patch can be signaled in the bitstream (e.g., the maximum x, y, z values for the geometric image, respectively). For each sample value in each channel of the UV plane, its dynamic range can be reduced by the distance from the value of that channel to the signaled maximum value of the same patch or partition, while still keeping the updated value non-negative. For example, xm, ym, and zm are the maximum values of the partitions on the three coordinate axes, respectively. For any point (x, y, z) in the patch, the updated value to be encoded can become (xm - x, ym - y, zm - z).
[0167] In some implementations, minimum values can be marked for some patches, and maximum values can be marked for some other patches. In some implementations, for a single patch, minimum values can be marked for some partitions within the patch, and maximum values can be marked for some other partitions within the patch.
[0168] In some implementations, a maximum signaling value can be sent for each defined region.
[0169] In some implementations, the patch can be further divided into smaller sub-patches. For each sub-patch, the dynamic range of the sample values within it is within a certain dynamic range, and therefore can be supported by the video codec. An offset value for each channel is signaled, and this offset value can be the maximum value of the channel for the samples in that sub-patch.
[0170] In some implementations, a special encoding / decoding process is applied to isolated 3D points within the patch, where one or more channels have a dynamic range exceeding a given bit depth. The position of this point in the UV plane and the out-of-range portion are recorded and encoded separately, and simultaneously, the original value may be replaced with the maximum possible value for a given bit depth.
[0171] For the non-restricted example, when xm, ym, and zm are the maximum values of the patch on the three coordinate axes and the allowed bit depth for encoding the sample is 8 bits; and for the point (x1, y1, z1) in the patch, the updated value is (xm - x1, ym - y1, zm - z1). When xm - x1 > 255 and ym - y1 <= 255 and zm - z1 <= 255, and the projection position of this point in the UV plane is (u1, v1), in addition to normal UV plane encoding and decoding for each position in the UV plane at a depth of 8 bits, the position (u1, v1) is also encoded into the bitstream along with additional out-of-limit information (xm - x1 - 255, 0, 0). On the decoder side, after decoding the geometric image using the method described above, the additional value (xm-x1-255, 0, 0) can be decoded and added to the decoded 3-channel value at position (u1, v1) to correctly recover the true geometric information of that point. That is, (xm-x1-255, 0, 0) + (255, ym-y1, zm-z1) = (xm-x1, ym-y1, zm-z1). The original geometry (x1, y1, z1) = (xm, ym, zm) - (xm-x1, ym-y1, zm-z1) can be recovered by further adjusting the value by subtracting this additional value from the maximum value of each channel (xm, ym, zm) to each position.
[0172] In this disclosure, the symbol "m^n" represents the exponentiation operation corresponding to the base m and the exponent or power n, i.e., m n .
[0173] The techniques disclosed in this disclosure can be used individually or in any combination in any order. Furthermore, each of these techniques (e.g., methods, embodiments), encoders, and decoders can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In some examples, one or more processors execute a program stored on a non-volatile computer-readable medium.
[0174] The above-described technology can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 12 A computer device (1300) is shown, which is adapted to implement certain embodiments of the disclosed subject matter.
[0175] The computer software can be encoded using any suitable machine code or computer language, and code including instructions can be created through mechanisms such as assembly, compilation, and linking. These instructions can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through decoding, microcode, etc.
[0176] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0177] Figure 12 The components shown for the computer device (1300) are exemplary in nature and are not intended to limit the scope or functionality of the computer software used to implement the embodiments of this application. Nor should the configuration of the components be construed as having any dependency or requirement on any component or combination thereof shown in the exemplary embodiments of the computer device (1300).
[0178] The computer device (1300) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to input from one or more human users through tactile input (e.g., keyboard input, swiping, data glove movement), audio input (e.g., sound, applause), visual input (e.g., gestures), and olfactory input (not shown). The human-machine interface device may also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from still cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0179] The human-machine interface input device may include one or more of the following (only one is shown): keyboard (1301), mouse (1302), touchpad (1303), touch screen (1310), data glove (not shown), joystick (1305), microphone (1306), scanner (1307), camera (1308).
[0180] The computer device (1300) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1310), data gloves (not shown), or joystick (1305), but may also include tactile feedback devices that are not used as input devices), audio output devices (e.g., speakers (1309), headphones (not shown)), visual output devices (e.g., screens (1310) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touchscreen input functionality, each of which may or may not have tactile feedback functionality—some of which may output two-dimensional or more three-dimensional visual outputs by means such as stereoscopic image output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0181] The computer device (1300) may also include human-accessible storage devices and related media, such as optical media including high-density read-only / rewritable optical discs (CD / DVD ROM / RW) (1320) or similar media (1321) with CD / DVD, thumb drives (1322), removable hard disk drives or solid-state drives (1323), conventional magnetic media such as magnetic tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.
[0182] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0183] The computer device (1300) may also include an interface (1355) to one or more communication networks (1354). The network may be wireless, wired, or optical. The network may also be a local area network (LAN), wide area network (WAN), metropolitan area network (MAN), vehicular and industrial network, real-time network, latency-tolerant network, etc. Examples of networks may include Ethernet, wireless LAN, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter for connection to certain general-purpose data ports or peripheral buses (1349) (e.g., a USB port on the computer device (1300); other systems are typically integrated into the core of the computer device (1300) via a system bus as described below (e.g., an Ethernet interface integrated into a PC computer system or a cellular network interface integrated into a smartphone computer system). By using any of these networks, the computer device (1300) can communicate with other entities. The communication can be unidirectional, used only for receiving (e.g., wireless television), unidirectional, used only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, such as through a local area or wide area digital network to other computer systems. Each of the above networks and network interfaces can use certain protocols and protocol stacks.
[0184] The aforementioned human-computer interface device, human-accessible storage device, and network interface can be connected to the core (1340) of the computer device (1300).
[0185] The core (1340) may include one or more central processing units (CPU) (1341), graphics processing units (GPUs) (1342), dedicated programmable processing units in the form of field-programmable gate arrays (FPGAs) (1343), task-specific hardware accelerators (1344), etc. These devices, along with read-only memory (ROM) (1345), random access memory (1346), internal mass storage (e.g., internal non-user-accessible hard disk drives, solid-state drives, etc.) (1347), etc., can be connected via a system bus (1348). In some computer systems, the system bus (1348) can be accessed as one or more physical connectors to allow for expansion with additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the core's system bus (1348) or connected via a peripheral bus (1349). Peripheral bus architectures include external controller interfaces (PCI), universal serial buses (USB), etc. In one example, a screen (1310) may be connected to a graphics adapter (1350). Peripheral bus architectures include PCI, USB, etc.
[0186] The CPU (1341), GPU (1342), FPGA (1343), and accelerator (1344) can execute certain instructions, which, when combined, constitute the aforementioned computer code. This computer code can be stored in ROM (1345) or RAM (1346). Transient data can also be stored in RAM (1346), while permanent data can be stored, for example, in internal mass storage (1347). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1341), GPUs (1342), mass storage (1347), ROM (1345), RAM (1346), etc.
[0187] The computer-readable medium may contain computer code for performing various computer-implemented operations. The medium and computer code may be specifically designed and constructed for the purposes of this application, or they may be media and code well-known and usable by those skilled in the art of computer software.
[0188] By way of example and not limitation, a computer system having an architecture (1300), particularly a core (1340), can provide functionality as a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in one or more tangible computer-readable media. Such computer-readable media can be media associated with the aforementioned user-accessible mass storage, as well as specific memory of the non-volatile core (1340), such as internal mass storage (1347) or ROM (1345). Software implementing various embodiments of this application can be stored in such a device and executed by the core (1340). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the core (1340), particularly the processor therein (including a CPU, GPU, FPGA, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in RAM (1346) and modifying such data structures according to software-defined processes. Alternatively or as an alternative, the computer system may provide logic hardwired or otherwise incorporated into circuitry (e.g., an accelerator (1344)) that may replace or operate with the software to perform the specific process or a specific portion of the specific process described herein. References to software may include logic, and vice versa, where appropriate. References to computer-readable media may include, where appropriate, circuitry storing the execution of software (such as an integrated circuit (IC)), circuitry containing the execution logic, or both. This application includes any suitable combination of hardware and software.
[0189] While this application has described several exemplary embodiments, various modifications, arrangements, and equivalent substitutions of the embodiments are all within the scope of this application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of this application and are thus within its spirit and scope.
Claims
1. A method for decoding geometric patches of a three-dimensional mesh, characterized in that, include: Receive an encoded bitstream, the encoded bitstream comprising geometric patches of a three-dimensional mesh; Extract a first syntax element from the encoded bitstream, wherein the first syntax element indicates whether the geometry patch is partitioned, the geometry patch comprising one or more partitions; and For out-of-limit points in one or more partitions, where the out-of-limit points correspond to isolated vertices in the 3D mesh, the following processing is performed: From the encoded bitstream, obtain the maximum pixel value of the partition where the exceedance point is located; From the encoded bitstream, obtain the encoded pixel value of the out-of-limit point; From the encoded bitstream, the over-limit value of the over-limit point is obtained, wherein the three channels in the over-limit point have a dynamic range exceeding a predetermined bit depth, and the component of the encoded pixel value corresponding to the non-zero component of the over-limit value is equal to 2. N -1, N is the predetermined bit depth; For each channel, the component of the excess value and the component of the encoded pixel value are subtracted from the component of the maximum pixel value to obtain the pixel value of the excess point, the pixel value corresponding to the three-dimensional coordinates of the isolated vertex.
2. The method according to claim 1, characterized in that: The encoded bitstream includes at least one of the following: Encoded geometry, or Encoded metadata.
3. The method according to claim 1, characterized in that: In response to the first syntax element indicating that the geometry patch is not partitioned, the geometry patch comprising a single partition; as well as In response to the first syntax element indicating that the geometry patch is partitioned, the method further includes: Extract a second syntax element from the encoded bitstream, the second syntax element being used to indicate the partition type of the geometry patch, and Based on the partition type, the geometry patch is partitioned to obtain multiple partitions.
4. The method according to claim 3, characterized in that, Further includes: For the partitions among the multiple partitions: Extract a third syntax element from the encoded bitstream, the third syntax element being used to indicate whether the partition is further partitioned; as well as In response to the third syntax element indicating that the partition is further partitioned, a fourth syntax element is extracted from the encoded bitstream, the fourth syntax element indicating the partition type of the partition; as well as Based on the partition type of the partition, the partition is further partitioned.
5. The method according to claim 1, characterized in that: In response to the first syntax element indicating that the geometry patch is not partitioned, the geometry patch comprising a single partition; as well as In response to the first syntax element indicating that the geometry patch is partitioned, the method further includes: From the encoded bitstream, extract the fifth syntax element indicating the partition depth, and The geometry patch is partitioned based on the partition depth and the predefined partition type to obtain multiple partitions.
6. An apparatus for decoding geometric patches of a three-dimensional mesh, characterized in that, The device includes: Memory for storing instructions; and A processor communicating with the memory, wherein when the processor executes the instructions, the processor is configured to cause the device to perform the method of any one of claims 1 to 5.
7. A method for encoding geometric patches of a three-dimensional mesh, characterized in that, include: Determine whether to partition the geometric patch of the 3D mesh, resulting in one or more partitions; For out-of-limit points in one or more partitions, where the out-of-limit points correspond to isolated vertices in the 3D mesh, the following processing is performed: Determine the maximum pixel value of the partition where the exceedance point is located; Determine the encoded pixel value of the point exceeding the limit; Determine the excess value of the excess point, wherein three channels in the excess point have a dynamic range exceeding a predetermined bit depth, and the component of the encoded pixel value corresponding to the non-zero component of the excess value is equal to 2. N -1, N is the predetermined bit depth; Specifically, for each channel, the pixel value of the out-of-limit point is obtained by subtracting the component of the out-of-limit value and the component of the encoded pixel value from the component of the maximum pixel value, and the pixel value corresponds to the three-dimensional coordinates of the isolated vertex.
8. An apparatus for encoding geometric patches of a three-dimensional mesh, characterized in that, The device includes: Memory for storing instructions; and A processor communicating with the memory, wherein when the processor executes the instructions, the processor is configured to cause the device to perform the method of claim 7.
9. A method for storing a bitstream, characterized in that, Generate a bitstream by performing the method of claim 7; and store the bitstream.
10. A non-volatile computer-readable storage medium, characterized in that, The system stores instructions and a bitstream, wherein when the instructions are executed by a processor, the instructions are configured to cause the processor to perform the method of claim 7 to generate the bitstream.
Citation Information
Patent Citations
Patch splitting for improving video-based point cloud compression performance
US11216984B2