Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
The V-Mesh compression method addresses the challenges of generating and processing point cloud data by optimizing encoding and decoding processes, enabling efficient transmission and reception for high-quality point cloud services in VR, AR, and MR applications.
Patent Information
- Application Number
- PCT/KR2025/004715
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-09
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-16
AI Technical Summary
The sheer number of points in 3D space makes it difficult to generate point cloud data, requiring significant processing power for transmission and reception, and existing methods face challenges in latency and encoding/decoding complexity.
A method for encoding and decoding point cloud data using a base mesh, displacement, and attribute in a bitstream, employing video-based dynamic mesh compression techniques like V-Mesh, which includes preprocessing, intra-frame and inter-frame encoding, and decoding processes to optimize transmission and reception.
This approach provides high-quality point cloud services with reduced latency and improved encoding/decoding efficiency, supporting applications such as autonomous driving and immersive experiences in VR, AR, and MR.
Smart Images

Figure KR2025004715_16102025_PF_FP_ABST
Abstract
Description
Mesh data encoding device, mesh data encoding method, mesh data decoding device, and mesh data decoding method
[0001] The embodiments provide a method for providing Point Cloud content to provide users with various services such as Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and autonomous driving services.
[0002] A point cloud is a collection of points in 3D space. The sheer number of points in 3D space makes it difficult to generate point cloud data.
[0003] There is a problem that a lot of processing power is required to transmit and receive point cloud data.
[0004] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device and reception method for efficiently transmitting and receiving point clouds in order to solve the problems described above.
[0005] The technical problem according to the embodiments is to provide a point cloud data transmission device, transmission method, point cloud data reception device, and reception method for resolving latency and encoding / decoding complexity.
[0006] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0007] In order to achieve the above-described purpose and other advantages, a decoding method according to embodiments may include a step of decoding a base mesh in a bitstream; a step of decoding a displacement in the bitstream; and a step of decoding an attribute in the bitstream. An encoding method according to embodiments may include a step of encoding a base mesh of mesh data; a step of encoding a displacement of the mesh data; and a step of encoding an attribute of the mesh data.
[0008] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide a high-quality point cloud service.
[0009] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can achieve various video codec methods.
[0010] The point cloud data transmission method, transmission device, point cloud data reception method, and reception device according to the embodiments can provide general-purpose point cloud content such as autonomous driving services.
[0011] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0012] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0013] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0014] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0015] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0016] Figure 5 shows a displacement generation process according to embodiments.
[0017] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0018] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0019] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0020] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0021] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0022] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0023] Fig. 12 shows a V-MES and Fig. 13 shows a point cloud data transmission device according to embodiments.
[0024] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0025] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0026] Fig. 15 may represent a dynamic mesh encoder according to embodiments.
[0027] Figure 16 can show a mesh refinement method according to real examples.
[0028] Figure 17 may represent an example of signaling parameters for determining a mesh subdivision method according to embodiments.
[0029] Figure 18 may represent an example of signaling parameters for determining a mesh refinement method according to embodiments.
[0030] Fig. 19 may show an example of a midpoint segmentation method according to embodiments.
[0031] Fig. 20 may show an example of a Butterfly segmentation method according to embodiments.
[0032] Fig. 21 may represent an LS3 segmentation method according to embodiments.
[0033] Fig. 22 may represent a normal vector-based segmentation method according to embodiments.
[0034] Figure 23 can show an example of a method for calculating a normal vector of a face according to embodiments.
[0035] Figure 24 can represent a normal vector-based segmentation method (without applying plane constraints) according to embodiments.
[0036] Fig. 25 may represent an example of a segmentation execution unit according to embodiments.
[0037] Fig. 26 may represent an example of a segmentation execution unit according to embodiments.
[0038] Fig. 27 may represent a normal vector-based segmentation method to which plane constraints are applied according to embodiments.
[0039] Fig. 28 can represent encoding of a displacement vector through a video codec according to embodiments.
[0040] Fig. 29 can represent the encoding of a displacement vector through an arithmetic coding method according to embodiments.
[0041] Fig. 30 can represent lifting transformation according to embodiments.
[0042] Fig. 31 can represent a normal vector-based lifting transformation prediction according to embodiments.
[0043] Figure 32 may represent normal vector-based prediction weights according to embodiments.
[0044] Fig. 33 may represent a dynamic mesh decoder according to embodiments.
[0045] Figure 34 can represent the normal vector assignment of the restored base mesh according to the embodiments.
[0046] Figure 35 can represent the calculation of normal vectors of restored base meshes according to embodiments.
[0047] Fig. 36 can represent decoding of a displacement vector based on a video codec according to embodiments.
[0048] Fig. 37 can represent the decoding of a displacement vector based on an arithmetic coding method according to embodiments.
[0049] Fig. 38 can represent the lifting inverse transformation according to embodiments.
[0050] Fig. 39 may represent an atlas sequence parameter set according to embodiments.
[0051] Fig. 40 may represent an atlas frame parameter set according to embodiments.
[0052] Figure 41 may represent a mesh patch data unit according to embodiments.
[0053] Fig. 42 may represent an encoding method according to embodiments.
[0054] Figure 43 may represent a decryption method according to embodiments.
[0055] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0056] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0057] Figure 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0058] The system of FIG. 1 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a reception unit (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or a dynamic mesh video encoder (hereinafter, referred to as an encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to a receiving device (110) or a dynamic mesh video decoder (hereinafter, decoder) (113).
[0059] The system of Fig. 1 can perform video-based dynamic mesh compression and decompression.
[0060] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.
[0061] First, the method of compressing dynamic mesh data starts with the V-PCC (Video-based point cloud compression) standard technology. Point cloud data is data that contains color information at the vertex coordinates (X, Y, Z). Mesh data refers to data in which connectivity information between vertices is added to this vertex information. When creating content, it can be created in the form of mesh data from the beginning. By adding connectivity information to point cloud data, it can be converted into mesh data and used.
[0062] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.
[0063] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.
[0064] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.
[0065] A mesh compression system may include a transmitting device and a receiving device. The transmitting device can encode mesh video to output a bitstream, which can be delivered to the receiving device via digital storage media or a network in the form of a file or streaming segment. The digital storage media may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, or SSD.
[0066] The transmitting device may roughly include a mesh video acquisition unit, a mesh video encoder, and a transmitting unit. The receiving device may roughly include a receiving unit, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / video / picture / frame encoding device, and the decoder may be referred to as a mesh video / video / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device and the receiving device may further include separate internal or external modules / units / components for a feedback process.
[0067] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by its vertices in 3D space and connection information that describes how those vertices are connected. It can also contain vertex properties such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included as a mesh property. The mapping is typically described as a set of parametric coordinates, called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement.
[0068] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc. into a mesh data type with the properties described above through a series of processes and generating a video composed of such mesh data. The mesh video may have properties of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video with properties and connection information that change over time can be expressed as a dynamic mesh video.
[0069] A mesh video encoder can encode an input mesh video into one or more video streams. A single video can include multiple frames, and a single frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and the mesh video can be used interchangeably with the mesh image / frame / picture. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. A mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0070] The encapsulation processing unit (file / segment encapsulation module) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from the metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder, or may be configured as a separate component / module. The encapsulation processing unit may encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The encapsulation processing unit may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. According to an embodiment, the encapsulation processing unit may encapsulate mesh video-related metadata itself in a file.
[0071] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit, or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process the same for transmission.
[0072] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and can include an element for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.
[0073] The receiver can receive mesh video data transmitted by a mesh video transmission device. Depending on the transmission channel, the receiver can receive mesh video data via a broadcast network, via broadband, or via digital storage media.
[0074] The receiving processing unit can perform processing on the received mesh video data according to a transmission protocol. The receiving processing unit can be included in the receiving unit, or can be configured as a separate component / module. In response to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmitting processing unit. The receiving processing unit can transfer the acquired mesh video data to the decapsulation processing unit, and transfer the acquired mesh video-related metadata to a metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
[0075] A decapsulation processing unit (file / segment decapsulation module) can decapsulate mesh video data in file format received from a receiving processing unit. The decapsulation processing unit can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder, or may be configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation processing unit may be in the form of a box or track within a file format. If necessary, the decapsulation processing unit may receive metadata required for decapsulation from the metadata processing unit. Mesh video related metadata can be passed to a Mesh video decoder for use in the Mesh video decoding process, or passed to a renderer for use in the Mesh video rendering process.
[0076] A mesh video decoder can receive a bitstream and perform operations corresponding to those of a mesh video encoder to decode video / images. The decoded mesh video can be displayed via a display unit. Users can view all or part of the rendered result via a VR / AR display or a general display.
[0077] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
[0078] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
[0079] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0080] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.
[0081] This document relates to dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and properties that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-view video, and AR / VR.
[0082] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0083] In this document, picture / frame can generally mean a unit representing one video of a specific time period.
[0084] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0085] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0086] The encoding process of Figure 1 is as follows.
[0087] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. The V-Mesh compression process receives the following data as input and performs compression.
[0088] Input mesh: Contains the 3D coordinates (geometry) of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.
[0089] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the properties (color, normal, displacement, etc.) of the mesh, and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video (or attribute for short). The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).
[0090] Material Library File: Contains information about the material properties used in a mesh, particularly information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
[0091] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0092] Base mesh: The input mesh is simplified (decimated) through a preprocessing process, and the objects of the input mesh are expressed using the minimum number of vertices determined by the user's standards.
[0093] Displacement: This is displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.
[0094] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.
[0095] Referring to FIGS. 2 to 7, a method for encoding mesh position information (vertex) is described, and referring to FIGS. 6-10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0096] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0097] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The encoder of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. The transmitting device of Fig. 1 may be broadly referred to as an encoder, and the dynamic mesh video encoder of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing (200) and an encoding (201) process as in Fig. 2. The pre-processor of Fig. 2 may be located in front of the encoder of Fig. 2. The pre-processor and the encoder of Fig. 2 may be referred to as a single encoder.
[0098] The preprocessor can receive a static dynamic mesh and / or an attribute map. The preprocessor can generate a base mesh and / or displacement through preprocessing. The preprocessor can receive feedback information from the encoder and generate the base mesh and / or displacement based on the feedback information.
[0099] The encoder can receive a base mesh, displacement mesh, static dynamic mesh, and / or attribute map. The encoder can encode mesh-related data to generate a compressed bitstream.
[0100] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0101] Figure 3 shows the configuration and operation of the pre-processor of Figure 2.
[0102] Fig. 3 shows a process of performing preprocessing on an input mesh. The preprocessing process (200) can be broadly divided into four steps: 1) Group of Frame (GoF) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). The preprocessor (200) can receive an input mesh, generate a displacement and / or base mesh, and transmit the generated displacement and / or base mesh to the encoder (201). The preprocessor (200) can transmit GoF information related to GoF generation to the encoder (201).
[0103] Below, each step of Figure 3 is described.
[0104] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, number of texture coordinates, vertex connection information, and texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame performs intra-frame encoding.
[0105] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.
[0106] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0107] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.
[0108] Fitting subdivision surface: This is the process of performing subdivision on a simplified mesh. The subdivision method can be a user-defined method, such as the mid-edge method. The fitting process ensures that the input mesh and the subdivision mesh are similar to each other.
[0109] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0110] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex midway between the edges between the vertices.
[0111] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decrypted base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents a position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0112] Figure 5 shows a displacement generation process according to embodiments.
[0113] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.
[0114] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a reconstructed base mesh and generate a subdivided reconstructed base mesh. The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and transform a coordinate system of the mesh into a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.
[0115] The method and device for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data (which may be referred to as a point cloud for short) according to the embodiments can refer to data including vertex coordinates and color information. The term "point cloud" includes mesh data, and in this document, point cloud and mesh data can be used interchangeably.
[0116] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).
[0117] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.
[0118] FIG. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments.
[0119] The encoding process of Fig. 6 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an intra-frame method. The encoder of Fig. 6 may include a preprocessor (200) and / or an encoder (201).
[0120] The preprocessor can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivided mesh. The quantizer can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder can encode the static mesh. The static mesh encoder can generate a bitstream including the encoded base mesh. The static mesh decoder can decode the encoded static mesh. The inverse quantizer can inversely quantize the quantized static mesh. The displacement calculation unit can receive the reconstructed static mesh and generate displacement, which is a position difference, based on the fitted subdivided mesh. The forward linear lifting unit can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. The image packing unit can pack an image based on the quantized lifting coefficients. A video encoder can encode a packed image. A video decoder decodes the encoded video. An image unpacker can unpack a packed image. A dequantizer can inversely quantize an image. An inverse linear lifting unit applies inverse lifting to the image to generate a reconstructed displacement. A mesh restoration unit restores a warped mesh using the reconstructed displacement and the reconstructed base mesh. An attribute transfer unit receives an input mesh and / or an input attribute map, and generates an attribute map based on the reconstructed warped mesh. A push-pull padding unit can pad data in the attribute map based on a push-pull method. A color space transformation unit can transform the space of a color component, which is an attribute. A video encoder can encode an attribute. A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.
[0121] Figure 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments.
[0122] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a preprocessor (200) and / or an encoder (201).
[0123] Among the encoding operations of Fig. 7, the corresponding configuration of the encoding operation of Fig. 6 refers to the description of Fig. 7. For the inter-frame-based encoding of Fig. 7, the motion encoder can encode motion based on the restored quantized reference base mesh. The base mesh restoration unit can restore the base mesh based on the restored quantized reference base mesh.
[0124] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.
[0125] The encoding method according to the embodiments includes base mesh encoding (intra encoding). When performing intra frame encoding on the current input mesh frame, the base mesh generated in the preprocessing process can be encoded using a static mesh compression technique after going through a quantization process. In the V-Mesh compression method, for example, Draco technology is applied, and the vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh are compressed.
[0126] The encoding method according to the embodiments may include motion field encoding (inter encoding). Inter frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field, may be calculated and this information may be encoded. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may be encoded as a value. Alternatively, the predicted motion field may be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and this predicted motion field The residual motion field, which is the difference between the value and the motion field value of the current vertex, can be encoded. This value can be encoded using entropy coding. The process of encoding the displacement and attribute map, excluding the motion field encoding process of inter-frame encoding, is the same as the structure of the intra-frame encoding method except for the base mesh encoding.
[0127] Figure 8 shows a lifting conversion process for displacement according to embodiments.
[0128] Figure 9 illustrates a process of packing transformation coefficients according to embodiments into a 2D image.
[0129] Figures 8-9 show the process of converting the displacement of the encoding process of Figures 6-7 and the process of packing the conversion coefficients, respectively.
[0130] The encoding method according to the embodiments includes displacement encoding.
[0131] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information.
[0132] Fig. 8 shows the process of transforming displacement information using lifting transform in V-Mesh. The transform coefficients generated through the transform process are quantized and then packed into a 2D image as shown in Fig. 9. The transform coefficients are organized into one block for every 256 (=16X16) units, and each block can be packed in z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. The transform coefficients can be packed by sorting them with Morton code within a block. The packed images generate a displacement video for each GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0133] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. The first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices (LoD0) of the base mesh. The first subdivision mesh may be generated by dividing the second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes the base mesh vertices (LoD0), LoD1 including the vertices additionally generated from LoD0, and the vertices additionally divided from LoD1. LoD is a level indicating the degree of detail (Level of Detail), and as the level index increases, the distance between vertices becomes closer and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When a vertex is further divided through subdivision, considering the previous vertices v1, v2 and the subdivided vertex v, the mesh can be encoded based on a prediction and / or update method. Instead of still encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated, and the mesh can be encoded using the residual value to reduce the size of the bitstream. The prediction process means the operation of predicting the current vertex v using the previous vertices v1, v2. Since adjacent subdivision meshes have similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted as the residual for the previous vertex position information, and the previous vertex position information is updated using the residual.
[0134] Referring to Figure 9, the vertices have coefficients generated through the lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.
[0135] Fig. 10 shows an attribute transfer process of a V-MESH compression method according to embodiments.
[0136] Figure 10 shows the detailed operation of attribute transfer of encoding such as Figures 6-7.
[0137] Encoding according to embodiments includes attribute map encoding.
[0138] Information about the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has vertex position information, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is created through the attribute transfer process.
[0139] Attribute transfer first checks whether each point P(u, v) in the 2D texture domain belongs to a texture triangle of the reconstructed deformed mesh, and if it is in the texture triangle T, calculates the barycentric coordinate (α, β γ) of P(u, v) according to the triangle T. Then, using the 3D vertex position and (α, β γ) of triangle T, calculate the 3D coordinate M(x, y, z) of P(u, v). Find the vertex coordinate M'(x', y', z') that corresponds to the position most similar to the calculated M(x, y, z) in the input mesh domain and the triangle T' that contains this point. And in this triangle T', the center of mass coordinates (α', β', γ') of M'(x', y', z') are calculated. Using the texture coordinates corresponding to the three vertices of triangle T' and (α', β', γ'), the texture coordinates (u', v') are calculated, and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.
[0140] The new attribute map generated through attribute transfer is grouped into GoF units to form an attribute map video, which is then compressed using a video codec.
[0141] Referring to Figure 10, the reference relationship between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.
[0142] The decoding process of Fig. 1 can perform the reverse process of the corresponding encoding process of Fig. 1. The specific decoding process is as follows.
[0143] Fig. 11 illustrates an intra-frame decoding process of a V-MESH compression method according to embodiments.
[0144] Fig. 11 shows the configuration and operation of a decoder of a receiving device such as Fig. 1.
[0145] Figure 11 illustrates the intra decoding process of V-Mesh technology. First, the input bitstream can be separated into a mesh sub-stream, a displacement sub-stream, an attribute map sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0146] The mesh sub-stream is decoded by the decoder of the static mesh codec used in encoding, such as Google Draco, and as a result, the connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be restored. The displacement sub-stream is decoded into a displacement video by the decoder of the video compression codec used in encoding, and goes through the image unpacking, inverse quantization, and inverse transform processes to restore displacement information for each vertex. Inverse quantization is applied to the restored base mesh, and this result is combined with the restored displacement information to generate the final decoded mesh.
[0147] The attribute map sub-stream is decoded through the decoder of the video compression codec used in encoding, and then restored to the final attribute map through processes such as color format conversion.
[0148] The restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0149] Referring to FIG. 11, the bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. The term "substream" is interpreted as referring to a portion of the bitstream included in the bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).
[0150] The decoder performs the following decoding operations within the frame. The static mesh decoder decodes the mesh to generate a reconstructed quantized base mesh, and the inverse quantizer applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder decodes the displacement, the unpacker unpacks the decoded video image, and the inverse quantizer inversely quantizes the quantized image. The linear lifting inverse transform applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit generates a warped mesh based on the base mesh and the displacement. The video decoder decodes the attribute map, and the color transformation unit transforms the color format and / or space to generate the decoded attribute map.
[0151] Figure 12 shows the inter-frame decoding process of the V-MESH compression method.
[0152] Fig. 12 shows the configuration and operation of the decoder of the receiving device of Fig. 1.
[0153] Figure 12 illustrates the inter-decoding process of V-Mesh technology. First, the input bitstream can be separated into a motion sub-stream, a displacement sub-stream, an attribute sub-stream, and a sub-stream containing mesh patch information, such as V3C / V-PCC.
[0154] The motion sub-stream is decoded through entropy decoding and inverse prediction processes, and the reconstructed motion information is combined with the already reconstructed reference base mesh to generate a reconstructed quantized base mesh for the current frame. The result of applying inverse quantization to this is combined with the displacement information reconstructed in the same way as the intra decoding described above to generate the final decoded mesh. The attribute map sub-stream is decoded in the same way as the intra decoding. The reconstructed decoded mesh and the decoded attribute map can be utilized by the receiver as the final mesh data that can be utilized by the user.
[0155] Referring to Fig. 12, the bitstream includes motion, displacement, and attribute maps. Since inter-frame decoding is performed, a process of decoding inter-frame motion information is further included. The motion is decoded, and a restored quantized base mesh for the motion is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operations in Fig. 12, which are identical to those in Fig. 11, refer to the description in Fig. 11.
[0156] Fig. 13 shows a point cloud data transmission device according to embodiments.
[0157] Fig. 13 corresponds to the transmitting device (100), the dynamic mesh video encoder (102), the encoder (preprocessor and encoder) of Fig. 2, and / or the transmitting encoding device corresponding thereto of Fig. 13. Each component of Fig. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0158] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13.
[0159] The mesh preprocessor receives the original mesh as input and generates a simplified mesh (decimated mesh). Simplification can be performed based on the target number of vertices or the target number of polygons that constitute the mesh. Parameterization can be performed on the simplified mesh to generate texture coordinates and texture connection information per vertex. Additionally, quantization of floating-point mesh information into fixed-point information can be performed. This result can be encoded as a base mesh through a static mesh encoding unit. The mesh preprocessor can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The subdivided mesh can be fitted by adjusting the vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.
[0160] When performing intra-encoding on the corresponding mesh frame, the base mesh generated through the mesh preprocessing unit can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to the multiplexing unit.
[0161] When performing inter encoding for the corresponding mesh frame, a motion vector encoding unit is performed, which can calculate a motion vector between the base mesh and the reference reconstruction base mesh as input and encode the value. The motion vector encoding unit can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexing unit.
[0162] The encoded base mesh and motion vectors can be used to generate a restored base mesh through the base mesh restoration unit.
[0163] The displacement vector calculator can perform mesh refinement on the restored base mesh. The displacement vector can be calculated as the difference in vertex positions between the refined restored base mesh and the fitted subdivision mesh generated in the preprocessing unit. As a result, the number of displacement vectors can be calculated equal to the number of vertices in the refined mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0164] A displacement vector video generator can transform displacement vectors for effective encoding. The transformation can be performed by a lifting transformation, a wavelet transformation, etc., depending on the embodiment. In addition, quantization can be performed on the transformed displacement vector values, i.e., the transform coefficients. Different quantization parameters can be applied to each axis of the transform coefficients, and the quantization parameters can be derived according to the promise of the encoder / decoder. The transformed and quantized displacement vector information can be packed into a 2D image. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh.
[0165] A displacement vector video encoder can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer.
[0166] The displacement vector restored through the displacement vector restorer and the base mesh restored through the base mesh restorer and refined are restored through the mesh restorer, and the restored mesh has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0167] The texture map of the original mesh can be regenerated as a texture map for the restored mesh through the texture map video generation unit. The color information per vertex of the texture map of the original mesh can be assigned to the texture coordinates of the restored mesh. The regenerated texture maps for each frame can be bundled into GoF units to generate a texture map video.
[0168] The generated texture map video can be encoded using a video compression codec through a texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to a multiplexing unit.
[0169] The generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be multiplexed into a single bitstream and transmitted to a receiver via a transmitter. Alternatively, the generated motion vector bitstream, base mesh bitstream, displacement vector bitstream, and texture map bitstream may be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via a transmitter.
[0170] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (displacement), and a texture map (attribute map). A transmitting device according to inter-encoding can generate a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map). A texture map obtained from a data input unit is generated and encoded based on a restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the segmented mesh. The base mesh is generated by preprocessing, simplifying, and encoding the original mesh. Motion is generated as a motion vector for the mesh of the current frame based on the reference base mesh of the previous frame.
[0171] Fig. 14 shows a point cloud data receiving device according to embodiments.
[0172] Fig. 14 corresponds to the receiving device (110), the dynamic mesh video decoder (113), the decoder of Figs. 11-12, and / or the receiving decoding device corresponding thereto of Fig. 1. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.
[0173] The bitstream of the received Mesh is demultiplexed into a compressed motion vector bitstream or base mesh bitstream, displacement vector bitstream, and texture map bitstream after file / segment decapsulation.
[0174] If the current mesh has inter-frame encoding applied according to the frame header information, the motion vector decoding unit can perform decoding on the motion vector bitstream. The final motion vector can be reconstructed by adding the previously decoded motion vector to the residual motion vector decoded from the bitstream using the previously decoded motion vector as a predictor.
[0175] If the current mesh has been encoded within the screen according to the frame header information, the base mesh bitstream can restore the connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh through the static mesh encoding unit.
[0176] In the base mesh restoration unit, if the current mesh has inter-frame encoding applied, the current base mesh can be restored by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. If the current mesh has intra-frame encoding applied, the static mesh decoding unit can perform inverse quantization on the decoded mesh to generate a restored base mesh.
[0177] The displacement vector bitstream can be decoded as a video bitstream using a video codec in a displacement vector video decoding unit.
[0178] In the displacement vector restoration unit, displacement vector transformation coefficients are extracted from the decoded displacement vector video, and the displacement vector is restored through the inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, a process of inverse transformation to the Cartesian coordinate system can be performed.
[0179] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Subdivision can generate vertex connection information, texture coordinates, and texture coordinate connection information, including the added vertices. The subdivided restored base mesh can be combined with the restored displacement vector to generate the final restored mesh.
[0180] The texture map bitstream can be decoded as a video bitstream using a video codec in a texture map video decoding unit. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of the corresponding vertex.
[0181] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.
[0182] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data. A receiving device according to inter-decoding can receive a motion vector (motion), a base mesh, a displacement vector (displacement), and a texture map (attribute map), and decode the restored mesh and the restored texture map to render the mesh data.
[0183] A point cloud data encoding device and method according to embodiments can encode mesh data and transmit a bitstream including the encoded mesh data. A point cloud data decoding device and method according to embodiments can receive a bitstream including mesh data and decode the mesh data. The point cloud data encoding / decoding method / device according to embodiments may be referred to as the method / device according to embodiments. The point cloud data encoding / decoding method / device according to embodiments may also be referred to as the mesh data encoding / decoding method / device according to embodiments. In addition, the term encoding / decoding method / device may be used in this document for short.
[0184] The encoding method and device according to the embodiments include a transmitting device (100) of Fig. 1, an obtaining unit (101), an encoder (102), an encapsulator (103), a transmitting unit (104), a pre-processor of Figs. 2-3 and 5, an encoder of Figs. 6-7, 13 and 15, bitstream and syntax generation of Figs. 39 to 41, and an encoding operation of Fig. 42.
[0185] The decryption method and device according to the embodiments include a receiving device (110), a receiving unit (111), a decapsulator (112), a decoder (113), a receiver (114) of Fig. 1, a decoder of Figs. 11-12, 14, and 33, bitstream and syntax parsing of Figs. 39 to 41, and a decoding operation of Fig. 43.
[0186] The decryption method and device can follow the reverse process of the operation of the encryption method and device.
[0187] The encoding device and decoding device may refer to a device including a memory and a processor.
[0188] The method and apparatus according to the embodiments may include a normal-based subdivision method for V-DMC.
[0189] The embodiments relate to V-DMC (Video-based Dynamic Mesh Compression), a method for compressing 3D dynamic mesh data using a 2D video codec.
[0190] Currently, when V-DMC receives an input mesh, it generates a base mesh, which is a more simplified mesh, through a simplification process. The generated base mesh is encoded by a base mesh codec and transmitted as a base mesh sub-bitstream. Then, the preprocessing unit of the encoder generates a fitted mesh that is processed to be similar to the original by performing subdivision based on the base mesh. In addition, the base mesh is subdivided to generate a reconstructed mesh and the coordinate difference with the fitted mesh is calculated; this data is the displacement. In conclusion, the V-DMC encoder performs subdivision to obtain displacements, and currently uses the midpoint method, which simply adds a vertex to the midpoint of two points during the subdivision process. Instead of this method, embodiments further include a subdivision method based on normal vectors so that subdivision can be performed more similarly to the original mesh. The subdivision is performed closer to the original than the method of simply subdividing by midpoints, and as a result, the displacements, which are the coordinate differences from the fitted mesh, are calculated as smaller values, resulting in a bit saving effect in the displacement sub-bitstream.
[0191] The embodiments relate to V-DMC (Video-based Dynamic Mesh Compression), a method for compressing 3D dynamic mesh data using a 2D video codec, and further include a method for encoding / decoding in units of displacement vectors in the displacement vector transformation and quantization steps, as well as syntax and semantics information related thereto. In addition, the operations of a transmitter (encoder) and a receiver (decoder) applying the same are described.
[0192] Recently, V-DMC technology has been actively standardizing since the CfP Response in April 2022. The mesh subdivision process in the technology applied to V-DMC so far has been performed using a midpoint subdivision method that creates a new vertex at the average value of the positions of two adjacent vertices and connects them. This existing method has a drawback in that it does not consider the characteristics of each vertex. The embodiments further include a method of performing mesh subdivision based on the normal vector of a vertex as one of the subdivision methods that considers the characteristics of each vertex. In addition, a signaling method for related parameters that enable the normal vector-based subdivision method to be applied to the V-DMC technology is further included.
[0193] The encoding / decoding method according to the embodiments further includes a normal vector-based segmentation method, which is a method of creating and connecting new vertices at positions corrected and / or derived using normal vector components during the mesh segmentation process of a dynamic mesh encoding / decoder.
[0194] The encoding / decoding method according to the embodiments provides a result of a fitted mesh that is more similar to the original as mesh fitting is performed after the simplified and refined mesh through the mesh preprocessing unit becomes more similar to the original.
[0195] The encoding / decoding method according to the embodiments provides the effect of reducing the size of the displacement vector by making the base mesh similar to the original through a normal vector-based subdivision method through a restoration process that subdivides the base mesh.
[0196] The encoding / decoding method according to the embodiments further includes a method of signaling in units of Atlas Sequence, Atlas Frame, and Mesh patch for applying a normal vector-based segmentation method in V-DMC technology.
[0197] The encoding / decoding method according to the embodiments further includes a method of performing prediction by giving weights based on a normal vector through a lifting transformation prediction unit.
[0198] V-DMC according to the embodiments may also be referred to as V-Mesh, and the expressions are used with the same meaning.
[0199] Fig. 15 may represent a dynamic mesh encoder according to embodiments.
[0200] The encoder of FIG. 15 may be configured as a device including memory and a processor. The processor may be configured to perform the operations of each component illustrated in FIG. 15. Each component of FIG. 15 may correspond to a combination of software, hardware, and / or a processor.
[0201] The mesh simplification unit receives a source mesh and generates a base mesh through a simplification process. The input mesh can be simplified to a target number of vertices or a target number of faces.
[0202] The mesh parameterization unit performs parameterization, which generates texture coordinates (e.g. UV coordinates) and texture connection information per vertex of the input mesh.
[0203] The mesh fitting part adjusts the vertex positions so that the subdivided mesh resembles the original mesh.
[0204] Depending on the embodiment, the mesh simplification unit, mesh parameterization unit, mesh refinement unit, and mesh fitting unit may be omitted, and if these processes are omitted, the original mesh may be applied as input to the mesh quantization unit.
[0205] At this time, coordinate information of the original mesh can be applied as input to the displacement vector calculation unit, and depending on the embodiment, the displacement vector encoding process (e.g., displacement vector calculation unit, displacement vector coordinate system conversion unit, displacement vector encoding unit) can be omitted.
[0206] The mesh quantization unit can quantize floating-point geometric information (e.g., x, y, z coordinate information) or / and texture coordinates (e.g., u, v) or / and normal information (e.g., nx, ny, nz) into fixed-point information.
[0207] Depending on the embodiment, quantization for certain components may be omitted.
[0208] The static mesh encoding unit performs encoding on the connection information, vertex geometry information, vertex texture coordinates, normal information, etc. of the base mesh.
[0209] The motion vector encoding unit can perform motion vector encoding by calculating a motion vector using the reference restoration base mesh and the current base mesh as input.
[0210] The motion vector encoding unit can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and perform entropy encoding on a residual motion vector obtained by subtracting a predicted motion vector from a current motion vector.
[0211] Depending on the embodiment, encoding of motion vectors may be performed on a vertex basis or a subgroup basis.
[0212] The displacement vector calculation unit calculates displacement vectors between the refined restored base mesh generated by performing refinement on the restored current base mesh and the fitted refined mesh. The displacement vector calculation unit can calculate a number of displacement vectors equal to the number of vertices of the refined mesh.
[0213] The displacement vector coordinate system transformation unit can transform the displacement vector of each vertex calculated in (x, y, z) space into a coordinate system based on the normal vector of each vertex (e.g., normal, tangential, bi-tangential).
[0214] Whether or not to perform displacement vector coordinate system transformation is determined by an agreement between the encoder and decoder, or the coordinate system transformation flag (applyLocalCoord) is transmitted in units such as sequence, GOF (Group Of Frame), frame, and submesh to determine whether or not to perform coordinate system transformation.
[0215] Depending on the embodiment, when a coordinate system transformation is applied (e.g., normal, tangential, bi-tangential), encoding may be performed only on the normal component of the coordinate system, and this may be performed only on the normal component at all times according to an agreement between the encoder and decoder, or the encoder may decide to signal a 1-bit flag (e.g., onlyNormFlag).
[0216] At this time, the normal vector can be calculated for each subdivided vertex based on geometric information or / and connection information of the surrounding vertices.
[0217] The base mesh restoration unit performs restoration of the base mesh according to the encoding type of the current mesh (e.g., inter-screen encoding or intra-screen encoding).
[0218] When in-screen encoding is performed, the current base mesh can be restored by performing inverse quantization on the quantized base mesh through the mesh quantization unit.
[0219] When inter-screen encoding is performed, the current base mesh can be generated by adding the restored motion vectors to the reference restored base mesh.
[0220] At this time, if the motion vector is not quantized, the motion vector restoration process is omitted and the current base mesh can be restored using the motion vector calculated by the motion vector encoding unit.
[0221] In the base mesh inverse quantization unit, inverse quantization is performed using inputs such as restoration geometric information (e.g., x, y, z) of the restoration base mesh or / and texture coordinates (e.g., u, v) or / and normal information (e.g., nx, ny, nz).
[0222] Depending on the embodiment, dequantization for certain components may be omitted.
[0223] The displacement vector restoration unit packs a 2D image / video, decodes the encoded bitstream through a 2D video encoder using a 2D video decoder, and performs inverse packing. In addition, the quantized transform coefficients on which inverse packing has been performed are inversely quantized and inversely transformed to restore the displacement vector.
[0224] The mesh restoration unit refines the base mesh restored by the base mesh dequantization unit to generate refined vertex position information, texture coordinates, and connection information. Restored vertex position information is generated by adding the restored displacement vector to the refined vertex position information.
[0225] Figure 16 can show a mesh refinement method according to real examples.
[0226] The mesh subdivision section of Fig. 15 can generate additional vertices by performing subdivision on the base mesh using the same principle as Fig. 16.
[0227] At this time, the generated vertices may include information such as geometric information, connection information, and texture coordinates calculated according to the subdivision method.
[0228] Mesh subdivision can be performed multiple times, and the number of mesh subdivisions can be determined by signaling / parsing the mesh subdivision count parameter (subdivision_iteration_count) in units such as a promise between the decoder / decoder or a sequence, GOF (Group Of Frame), frame, or submesh.
[0229] Depending on the embodiment, the number of times the current mesh refinement has been performed can be defined as the Level of Details.
[0230] According to the embodiment, the vertices of the base mesh are subdivided R0, the newly generated vertices are subdivided 1 time, R1, . . . , and the newly generated vertices are subdivided n times, R n When defined as , the vertex LoD contained in any subdivision level k k can be defined as: LoD k =(R0U R1. . UR k ).
[0231] The subdivision method decision unit determines the mesh subdivision method to be performed at the current subdivision level.
[0232] The above mesh subdivision method may include a midpoint subdivision method, a butterfly subdivision method, a loop subdivision method, an LS3 (Least Squares Subdivision Surfaces) method, a normal vector-based subdivision method, etc. The mesh subdivision method to be performed can be determined by signaling / parsing an N-bit mesh subdivision method parameter (subdivision_method) in units such as an agreement between the decoder / decoder or a sequence, GOF, frame, submesh, or subdivision level.
[0233] Figure 17 may represent an example of signaling parameters for determining a mesh subdivision method according to embodiments.
[0234] FIG. 17 illustrates an example of signaling an N-bit mesh subdivision method parameter (subdivision_method) in units of / decoder promise or sequence, GOF, frame, submesh, subdivision level, etc., as described in FIG. 16.
[0235] The normal vector-based segmentation method according to the embodiments may include a method that does not apply a planar constraint and / or a method that applies a planar constraint.
[0236] Figure 18 may represent an example of signaling parameters for determining a mesh refinement method according to embodiments.
[0237] As shown in Fig. 18, depending on the embodiment, the mesh subdivision method to be performed can be determined by signaling / parsing the M-bit mesh subdivision method parameter (subdivision_method) and the 1-bit plane constraint flag (plane_constraint_flag).
[0238] The planar constraint flag may be omitted if the current subdivision method being performed is not a normal vector-based subdivision method.
[0239] N-bit and M-bit can be the same or different.
[0240] The subdivision performing unit performs mesh subdivision according to the subdivision method determined in the subdivision method determining unit.
[0241] At this time, new vertices can be created and existing vertices can be adjusted through mesh subdivision.
[0242] One or more new vertices can be created per edge of an existing mesh.
[0243] Fig. 19 may show an example of a midpoint segmentation method according to embodiments.
[0244] In an embodiment, when mesh subdivision is performed using a midpoint subdivision method as in FIG. 19, a new vertex generated through mesh subdivision can be calculated as in the following equation 1.
[0245] Formula 1: v new =1 / 2(v0+v1)
[0246] In the formula, v0 and v1 represent the two end vertices of the existing mesh edge that are the target of new vertex generation, and v new means a new vertex is created.
[0247] Fig. 20 may show an example of a Butterfly segmentation method according to embodiments.
[0248] In an embodiment, when mesh subdivision is performed using a butterfly subdivision method as in FIG. 20, a new vertex generated through mesh subdivision can be calculated as in Equation 2.
[0249] Equation 2: v new =1 / 2(v0+v1)+1 / 8(v2+v3)-1 / 16(v4+v5+v6+v7)
[0250] In Equation 2, v0 and v1 represent the two end vertices of the existing mesh edge that is the target of new vertex generation, v2, v3, v4, v5, v6, and v7 represent the vertices adjacent to the edge, and v new means a new vertex is created.
[0251] In an embodiment, when mesh subdivision is performed using the Loop subdivision method, new vertices generated through mesh subdivision can be calculated as in Equations 3 and 4, and existing vertices can be adjusted as in Equations 5 and 6.
[0252] Formula 3:V new 0 =3 / 8*(v0+v1)+1 / 8*(v2+v6) Formula 4: V new 1 =1 / 2*(v1+v6)
[0253] The calculation formula for a new vertex generated from an edge that is not a boundary of the mesh is as shown in Equation 3. v0 and v1 represent the vertices at both ends of the edge, v2 and v3 represent vertices adjacent to vertices v0 and v1, and v new 0 means a new vertex is created.
[0254] The calculation formula for a new vertex generated at an edge, which is the boundary of the mesh, is as shown in Equation 4. v1 and v6 represent the vertices at both ends of the edge, and v new 1 means a new vertex is created.
[0255]
[0256] k: the number of neighboring vertices of v0
[0257] = neighboring vertices of v0,
[0258]
[0259] Formula 6: v1=3 / 4*v1+1 / 8*(v6+v7)
[0260] The calculation formula for adjusting vertices that are not the boundary of the mesh is as shown in Equation 5. v0 represents the vertex to be adjusted, and v1, v2, v3, v4, v5, and v6 represent vertices adjacent to vertex v0. The calculation formula for adjusting vertices that are the boundary of the mesh is as shown in Equation 6. v1 represents the vertex to be adjusted, and v6 and v7 represent vertices adjacent to vertex v1.
[0261] Fig. 21 may represent an LS3 segmentation method according to embodiments.
[0262] In some embodiments, when mesh subdivision is performed using the LS3 (Least Squares Subdivision Surfaces) method, mesh subdivision may be performed through the process of FIG. 21.
[0263] By performing mesh subdivision using the LS3 method, new vertices can be created and existing vertices can be adjusted. The mesh partitioning unit of Fig. 21 can create new vertices by partitioning an existing mesh. The mesh smoothing unit of Fig. 21 can smooth the surface of the mesh by filtering each vertex of the mesh to which new vertices have been added in the mesh partitioning unit.
[0264] The mesh projection part of Fig. 21 can generate a refined mesh by calculating a mathematical surface for a mesh smoothed in a mesh relaxation part and projecting each vertex of the smoothed mesh in the direction of the calculated mathematical surface.
[0265] Fig. 22 may represent a normal vector-based segmentation method according to embodiments.
[0266] In some embodiments, when mesh refinement is performed using a normal vector-based refinement method and no plane constraints are applied, mesh refinement may be performed through the process of FIG. 22.
[0267] The normal vector calculation unit of Fig. 22 can calculate a normal vector value for each vertex of the mesh. In some embodiments, the normal vector of a vertex can be calculated by adding and normalizing the normal vectors of faces adjacent to the vertex.
[0268] Figure 23 can show an example of a method for calculating a normal vector of a face according to embodiments.
[0269] The normal vector of a surface can be calculated by calculating two vectors (v1-v0, v2-v0) through the three vertices that make up the surface, as shown in Figure 23, and taking the outer product of the two vectors.
[0270] In some embodiments, if there is normal vector information of a vertex calculated at a previous subdivision level, the normal vector of the vertex that does not include normal vector information can be calculated by interpolating the normal vector information of the vertex by averaging or distance-based weighting.
[0271] In some embodiments, when the vertices of the input mesh include normal vector information, a normal vector-based segmentation method can be performed using the normal vector information.
[0272] The normal vector-based segmentation performing unit of Fig. 22 can perform mesh segmentation using the normal vector of each vertex calculated in the normal vector calculation unit.
[0273] Fig. 24 can represent a normal vector-based segmentation method (without applying plane constraints) according to embodiments.
[0274] As shown in Fig. 24, a new vertex generated through normal vector-based mesh subdivision can be calculated as in Equation 7 by weighting the vertex value calculated by the midpoint subdivision method and the correction vector value derived through the normal vector of the vertex.
[0275]
[0276]
[0277] The method according to the embodiments can generate a new vertex Vnew using a normal component for a midpoint.
[0278] Equation 7 can be expressed as:
[0279] refineVector = Dot( ( d0 * normals[ v0 ] + d1 * normals[ v1 ] ) , nv ) * nv *Min( pow( 2.0, subdivisionIterationIndex - 4 ), 0.25 )
[0280] A method according to embodiments may calculate two normal components for two vertices of a base mesh, respectively, and add up the normal components by applying weights to them. The weights are determined based on the values of counts for subdivisions. The subdivision counts may indicate the degree of subdivision of the base mesh. The subdivision information may indicate the level of LoD (level of detail). In order to subdivide the base mesh so as to fit the original mesh, weights are calculated according to the subdivision level, and normals of the vertices of the base mesh are weighted and added using the calculated weights to calculate new vertices of the base mesh.
[0281] positionArrayOut[ e ][ d ] = ( positionArrayOut[ v0 ][ d ] +positionArrayOut[ v1 ][ d ] ) / 2 + refineVector[ d ] * normalAligned
[0282] A new vertex is calculated by adding the midpoint of the vertices to the weighted sum of the normals of the vertices of the base mesh.
[0283] In Equation 7, v0 and v1 are the two end vertices of the existing mesh edge that are the target of new vertex generation. , means the normal vector of vertices v0, v1, and v mid means the vertex value calculated by the Midpoint subdivision method, Is , It means the correction vector induced through v new means a new vertex is created.
[0284] The weight parameter w can be determined based on an agreement between the encoder and decoder, or can be signaled / parsed in units such as sequence, GOF, frame, submesh, or granularity level, or can be calculated based on vertex geometry, connection information, etc.
[0285] In some embodiments, the weight parameter w may be calculated as in Equation 8 based on the current refinement level (the number of times the current mesh refinement has been performed).
[0286] Equation 8: w=1 / (2currentLOD+n)=2-(currentLOD+n)
[0287] In Equation 8, currentLOD represents the current subdivision level, and n can be determined as any positive value.
[0288] w=2^(subdivisionIterationIndex - 4)
[0289] Fig. 25 may represent an example of a segmentation execution unit according to embodiments.
[0290] According to an embodiment, the subdivision performing unit may perform mesh subdivision by calculating the normal vector of a vertex through a normal vector calculation unit at each subdivision level as shown in FIG. 25, and using the normal vector of each subdivision level vertex calculated through the normal vector-based subdivision performing unit.
[0291] According to an embodiment, the subdivision performing unit may perform a normal vector calculation unit only for subdivision level s from subdivision level s to subdivision level e, and the normal vector calculation unit may be omitted for subdivision levels (s+1) to e. In addition, mesh subdivision may be performed using a midpoint subdivision method for subdivision levels s to (e-1), and mesh subdivision may be performed for subdivision level e by correcting new vertices generated in subdivision levels s to e using normal vectors calculated in the subdivision level s through a normal vector-based subdivision performing unit.
[0292] At this time, the normal vector-based subdivision performing unit of the above subdivision level e can perform mesh subdivision using a weight parameter w as in Equation 8, depending on which subdivision level the new vertices were generated from.
[0293] The above-mentioned subdivision levels s and e can be determined by agreement between the encoder / decoder, or by signaling / parsing in units such as sequence, GOF, frame, and submesh.
[0294] Fig. 26 may represent an example of a segmentation execution unit according to embodiments.
[0295] Fig. 26 is an embodiment of performing mesh subdivision for subdivision levels 0 to 2, performing normal vector calculation only for subdivision level 0, performing midpoint subdivision for subdivision levels 0 to 1, and performing mesh subdivision for subdivision level 2 by correcting new vertices generated at subdivision levels 0 to 2 using normal vectors of vertices calculated at subdivision level 0.
[0296] Fig. 27 may represent a normal vector-based segmentation method to which plane constraints are applied according to embodiments.
[0297] When a plane constraint is applied, mesh refinement is performed using a normal vector-based refinement method, and when a plane constraint is applied, mesh refinement can be performed through the process of Fig. 27.
[0298] Fig. 28 can represent encoding of a displacement vector through a video codec according to embodiments.
[0299] The normal vector calculation unit can calculate normal vector values for each vertex and face of the mesh. In some embodiments, the normal vector of a face can be calculated by calculating two vectors (v1-v0, v2-v0) through three vertices constituting the face, as shown in Fig. 23, and taking the cross product of the two vectors.
[0300] In some embodiments, the normal vector of a vertex can be calculated by adding and normalizing the normal vectors of faces adjacent to the vertex.
[0301] In some embodiments, if there is normal vector information of a vertex calculated at a previous subdivision level, the normal vector of the vertex that does not include normal vector information can be calculated by interpolating the normal vector information of the vertex by averaging or distance-based weighting.
[0302] In some embodiments, when the vertices of the input mesh include normal vector information, a normal vector-based segmentation method can be performed using the normal vector information.
[0303] The plane constraint parameter determination part of Fig. 27 is a parameter according to the plane constraint. and can be determined. At this time, the parameter is a parameter equal to the number of faces adjacent to the mesh edges that are the target of new vertex generation. i It can be composed of.
[0304] Depending on the embodiment, the parameter i can be calculated as in Equation 9 so that it is proportional to the area of the surfaces adjacent to the above-mentioned main line.
[0305]
[0306] In Equation 9, k represents the number of faces adjacent to the mesh edge that is the target of new vertex generation, and a i means the area of one adjacent side, i means the parameter for one adjacent side.
[0307] According to an embodiment, the above parameters can be calculated as in formula 10.
[0308]
[0309]
[0310]
[0311] In Equation 10, v0 and v1 are the two end vertices of the existing mesh edge that are the target of new vertex generation, and v norm A denotes a vertex calculated by a normal vector-based subdivision method without plane constraints, k denotes the number of faces adjacent to the edge, and A i refers to the equation of the surfaces adjacent to the above-mentioned edge.
[0312] According to an embodiment, the above parameters or / and It can be determined by an agreement between the encoder / decoder, or the encoder can calculate the optimal value for each unit such as sequence, GOF, frame, submesh, vertex, etc. and signal / parse it to determine the optimal value.
[0313] The normal vector-based subdivision performing unit of Fig. 28 can perform mesh subdivision using the normal vectors of each vertex and face calculated by the normal vector calculation unit and the parameters determined by the plane constraint parameter determination unit.
[0314] When a planar constraint is applied, a new vertex generated through normal vector-based mesh refinement can be calculated from a matrix derived through the equation of the vertex of the mesh edge that is the target of new vertex generation and the adjacent face, and the planar constraint parameters, as in Equation 11.
[0315]
[0316]
[0317]
[0318] α in Equation 11 i , is the plane constraint parameter, A i is the equation of the mesh edges and adjacent faces that are the target of new vertex generation, v i P is the vertex at both ends of the edge i is the vertex v iThe matrix derived from is Q, the equation of the vertex of the edge and the adjacent surface, the matrix derived through the plane constraint parameter, and v new refers to a new vertex that is created. The equation of a surface can be calculated through the vertices included in the surface and the normal vector of the surface.
[0319] Fig. 29 can represent the encoding of a displacement vector through an arithmetic coding method according to embodiments.
[0320] In the displacement vector encoding unit, encoding of displacement vectors can be performed using a 2D video codec such as H.264, HEVC, or VVC, or an arithmetic coding method.
[0321] At this time, the encoding method of the displacement vector can be determined by an agreement between the sub-encoder and the decoder, or can be determined by signaling the encoding method determined by the encoder as an index (dispEncType), and an index (profilesetIdx) for the profile information defined between the sub-encoder and the decoder can be signaled according to the determined encoding method.
[0322] According to an embodiment, when a displacement vector is encoded through a 2D video codec, displacement vector encoding can be performed through the process of FIG. 28.
[0323] The displacement vector transform coefficient quantization unit can perform quantization on the displacement vector transform coefficient.
[0324] Depending on the embodiment, the transform coefficients may be quantized with different quantization parameters for each axis and / or each level of refinement, and the quantization parameters may be determined by a contract between the decoder and the decoder or derived from the decoder and the decoder.
[0325] The displacement vector transformation coefficient packing part of Fig. 28 can pack K displacement vector transformation coefficients or quantized displacement vector transformation coefficients into an image of size WХH.
[0326] The displacement vector video encoding unit of Fig. 28 can encode displacement vector transform coefficients packed into an image through a 2D video codec.
[0327] Fig. 30 can represent lifting transformation according to embodiments.
[0328] In an embodiment, when a displacement vector is encoded using an arithmetic coding method, displacement vector encoding may be performed using the process of FIG. 30.
[0329] The displacement vector prediction unit can use the transform coefficients of the reference frame stored in the buffer as predictors to generate residual transform coefficients for the current transform coefficients.
[0330] Depending on the embodiment, the order of the displacement vector transform coefficient quantization unit and the displacement vector prediction unit may be changed, and the displacement vector prediction unit may be omitted.
[0331] In the displacement vector conversion section of FIGS. 29 to 30, displacement vectors in the (x, y, z) or (normal, tangential, bi-tangential) coordinate system can be converted through the displacement vector conversion section.
[0332] In an embodiment, when a coordinate system transformation is performed into a (normal, tangential, bi-tangential) coordinate system, a 1D scalar displacement vector of the normal component is applied as an input to a displacement vector transformation unit, and transformation, quantization, and encoding can be performed on the displacement value of the normal component.
[0333] Depending on the embodiment, the transformation may be performed by a wavelet transform, a lifting transform, etc.
[0334] In some embodiments, when a lifting transformation is performed, the displacement vector may be transformed through the process of FIG. 30. The number of lifting transformations may be determined based on the mesh subdivision count parameter (subdivision_iteration_count). The lifting transformation process may be performed per mesh subdivision level.
[0335] In the lifting transformation prediction part, the vertex R of the kth subdivision level k When performing displacement vector prediction, t(t <k 또는 t헽)번째 세분화 레벨의 정점 R t Displacement vector prediction of the kth subdivision level can be performed using the displacement vector as a predictor.
[0336] A residual signal can be generated by taking the difference between the displacement vector at the subdivision level and the predicted displacement vector.
[0337] In some embodiments, when performing prediction of a displacement vector, prediction can be performed by averaging or distance-based weighted averaging displacement vectors of m nearby vertices based on connection information among vertices with a lower level of detail than the current vertex.
[0338] According to an embodiment, prediction can be performed based on the displacement vectors of the m vertices used to generate the current vertex in the mesh refinement step.
[0339] Fig. 31 can represent a normal vector-based lifting transformation prediction according to embodiments.
[0340] According to an embodiment, as in Fig. 31, two vertices with a low level of subdivision among the vertices adjacent to the prediction target vertex may be used as predictors, and prediction may be performed as in Equation 12 using weights determined based on the normal vector.
[0341] Equation 12: Signal(v)=Signal(v)-(w 0 * Signal(v0)+w1*Signal(v1))
[0342] In Equation 12, v represents the target vertex to be predicted, v0 and v1 represent vertices used as predictors, signal represents the displacement vector of the corresponding vertex, and w0 and w1 represent normal vector-based prediction weights for the displacement vectors of vertices v0 and v1.
[0343] Normal vector-based prediction weights can be determined through the similarity between the normal vectors of the target vertex to be predicted and the vertex used as a predictor.
[0344] In some embodiments, the normal vector-based prediction weights can be calculated through the ratio of inner product operations between normal vectors, as in Equation 13.
[0345]
[0346] Figure 32 may represent normal vector-based prediction weights according to embodiments.
[0347] In some embodiments, the normal vector-based prediction weights may be calculated using the difference in the inner product operation between the normal vectors as an index into a prediction weight table defined between the encoder and decoder.
[0348] The lifting transformation update unit can perform a process of updating the displacement vector of a vertex used for prediction through a residual signal generated by the lifting transformation prediction unit.
[0349] Depending on the embodiment, updates may be performed based on weights derived from information such as the average or current granularity level, the number of edges connected to a vertex, etc.
[0350] The texture map generation unit of Fig. 30 generates a texture map of the restoration mesh through the relationship between the texture coordinates and connection information of the restoration mesh and the texture map of the original mesh.
[0351] The texture map encoding unit of Fig. 15 stacks the texture map generated through the texture map generation unit in the frame order of the mesh to form a texture map video, and encodes it through a 2D video encoder.
[0352] Depending on the embodiment, if the color space of the texture map is RGB444, encoding may be performed after conversion to a color space such as YUV420 or YUV444.
[0353] Fig. 33 may represent a dynamic mesh decoder according to embodiments.
[0354] Dynamic mesh encoding is a process in which the encoded base mesh bitstream, displacement vector bitstream, and texture map bitstream are transmitted, and the decoder decode each bitstream to restore the mesh. First, the base mesh is decoded using the motion vector or static mesh decoding unit depending on whether it is an inter or intra frame, and then the geometric information is restored along with the decoded displacement vector information through subdivision.
[0355] The decoder of Fig. 33 can follow the reverse process of the operation of the encoder of Fig. 15.
[0356] The decoder of FIG. 33 may be configured as a device including memory and a processor. The processor may be configured to perform the operations of each component illustrated in FIG. 33. Each component of FIG. 33 may correspond to a combination of software, hardware, and / or processor.
[0357] The motion vector decoding unit of Figure 21 can perform motion vector decoding when inter-screen prediction of the current mesh is performed.
[0358] The residual motion vector can be decoded in units of vertices or subblocks through a motion vector bitstream, prediction based on connection information can be performed using a previously decoded motion vector as a predictor, and the motion vector can be decoded by adding it to the residual motion vector.
[0359] The static mesh decoding unit in Figure 21 can restore the connection information, vertex geometry information, vertex texture coordinates, normal information, etc. of the base mesh.
[0360] - In the base mesh restoration section of Figure 21, if the current base mesh is encoded based on a reference mesh, the current base mesh can be restored by adding a motion vector to the reference base mesh and then performing inverse quantization.
[0361] The base mesh restoration unit can restore the current base mesh by performing inverse quantization if the current base mesh has been decoded through the static mesh encoding unit.
[0362] Depending on the embodiment, the dequantization process may be omitted.
[0363] The texture map decoding unit of Fig. 33 may be a process of receiving a texture map bitstream as input and decoding the texture map.
[0364] Texture map decoder types include video decoder, zero run length decoder, and arithmetic decoder.
[0365] Depending on the embodiment, color space conversion of the texture map can be performed.
[0366] In the displacement vector coordinate system inverse transformation section of Fig. 33, if the coordinate system transformation flag (applyLocalCoord) parsed by sequence, GOF (Group Of Frame), frame, and submesh units is 1, the inverse quantized restored displacement vector can be inversely transformed to the (x, y, z) coordinate axes.
[0367] Depending on the embodiment, it is always possible to perform coordinate system inversion without sending flags.
[0368] Figure 34 can represent the normal vector assignment of the restored base mesh according to the embodiments.
[0369] According to an embodiment, a normal vector per vertex may be calculated based on the restored vertex position information of the restored base mesh, and for vertices additionally generated through the subdivision process, the normal vector per vertex of the calculated restored base mesh may be interpolated and assigned as the normal vector of the newly generated vertex.
[0370] At this time, interpolation can be performed by weighting the normal vector information of the base mesh used for subdivision based on average or distance.
[0371] Figure 35 can represent the calculation of normal vectors of restored base meshes according to embodiments.
[0372] In some embodiments, after performing subdivision on a restored base mesh, normal vectors can be calculated for vertices generated through the subdivision section and vertices of the base mesh.
[0373] By calculating the tangential and bi-tangential vectors orthogonal to the normal vector through the calculated normal vector per vertex, the displacement vector coordinate system inverse transformation can be performed using the following formula.
[0374]
[0375] The mesh restoration unit of Fig. 33 can calculate and restore vertex geometric information of the restoration mesh by adding a restoration displacement vector to vertices generated through a subdivision process in the mesh subdivision unit.
[0376] The displacement vector decoding unit of Fig. 33 can perform displacement vector decoding using a 2D video codec such as H.264, HEVC, VVC, or an arithmetic coding method.
[0377] At this time, the decoding method of the displacement vector can be determined by an agreement between the sub-decoder and the decoder, or can be determined by parsing the decoding method determined by the encoder as an index (dispEncType), and an index (porfilesetIdx) for the profile information defined between the sub-decoder and the decoder can be signaled according to the determined decoding method.
[0378] Fig. 36 can represent decoding of a displacement vector based on a video codec according to embodiments.
[0379] In some embodiments, when a displacement vector is decoded through a 2D video codec, displacement vector encoding may be performed through the process of FIG. 36.
[0380] The displacement vector video decoding unit of Fig. 36 can decode a transformation coefficient image from a displacement vector bitstream through a 2D video codec.
[0381] The displacement vector transform coefficient packing part of Fig. 36 can perform reverse packing according to the scanning order defined by the encoder / decoder agreement from the restored transform coefficient image or the scanning order parsed into units such as sequences and frames.
[0382] The displacement vector transform coefficient quantization unit of Fig. 36 can perform inverse quantization on the displacement vector transform coefficient.
[0383] Fig. 37 can represent the decoding of a displacement vector based on an arithmetic coding method according to embodiments.
[0384] In some embodiments, when a displacement vector is decoded using an arithmetic coding method, displacement vector decoding may be performed using the process of Figure 25.
[0385] The displacement vector prediction and restoration unit of Fig. 37 can restore the displacement vector transformation coefficient by adding the transformation coefficient of the reference frame stored in the buffer and the restoration residual transformation coefficient.
[0386] Depending on the embodiment, the order of the displacement vector prediction and restoration unit and the displacement vector transformation coefficient inverse quantization unit may be changed, and the displacement vector prediction and restoration unit may be omitted.
[0387] In the displacement vector inverse transform section of Fig. 37, the inverse transform of the transform performed in the encoder can be performed.
[0388] Depending on the embodiment, the transformation may be performed by a wavelet inverse transform, a lifting inverse transform, etc.
[0389] Fig. 38 can represent the lifting inverse transformation according to embodiments.
[0390] In an embodiment, when lifting inverse transformation is performed, transformation of the displacement vector can be performed through the process of FIG. 38.
[0391] The number of lifting inverse transformations can be determined by the mesh subdivision count parameter (subdivision_iteration_count).
[0392] The lifting inverse transformation process can be performed at each mesh subdivision level.
[0393] In the lifting inverse transformation update section of Fig. 38, a process of updating the displacement vector of the vertex used for prediction through the parsed residual signal can be performed.
[0394] Depending on the embodiment, updates may be performed based on weights derived from information such as the average or current granularity level, the number of edges connected to a vertex, etc.
[0395] The lifting inverse prediction part of Fig. 38 is the vertex R of the kth subdivision level. k When performing displacement vector prediction, t(t <k 또는 t Vertex R of the kth subdivision level t Displacement vector prediction of the kth subdivision level can be performed using the displacement vector as a predictor.
[0396] The displacement vector of the subdivision level vertex can be restored by summing the predicted displacement vector and the parsed residual signal.
[0397] In some embodiments, when performing prediction of a displacement vector, prediction can be performed by averaging or distance-based weighted averaging displacement vectors of m nearby vertices based on connection information among vertices with a lower level of detail than the current vertex.
[0398] In some embodiments, prediction can be performed based on displacement vectors of m vertices used to generate the current vertex in the mesh refinement step.
[0399] According to an embodiment, as in Fig. 31, two vertices with a low level of subdivision among the vertices adjacent to the prediction target vertex may be used as predictors, and prediction may be performed as in Equation 12 using weights determined based on the normal vector.
[0400] In Equation 12, v represents the target vertex to be predicted, v0 and v1 represent the vertices used as predictors, Signal represents the displacement vector of the corresponding vertex, and w0 and w1 represent normal vector-based prediction weights for the displacement vectors of the vertices v0 and v1.
[0401] Normal vector-based prediction weights can be determined through the similarity between the normal vectors of the target vertex to be predicted and the vertex used as a predictor.
[0402] In some embodiments, the normal vector-based prediction weights can be calculated through the ratio of inner product operations between normal vectors (see Equation 13).
[0403] In some embodiments, the normal vector-based prediction weights may be calculated by using the difference in the inner product operation between the normal vectors as an index of a prediction weight table defined between the encoder and decoder (see FIG. 32).
[0404] An encoding method according to embodiments may encode mesh data and generate parameters related to the mesh data. The encoding method may generate a bitstream including the encoded mesh data and parameters. A decoding method according to embodiments may decode mesh data within a bitstream based on parameters within the bitstream. Hereinafter, syntax elements within a bitstream will be described with reference to FIGS. 39 to 41.
[0405] Fig. 39 may represent an atlas sequence parameter set according to embodiments.
[0406] asve_subdivision_iteration_count: Indicates the number of times subdivision is repeated. A value of 0 indicates that subdivision is not performed.
[0407] asve_subdivision_method[i]: This parameter determines the mesh subdivision method in sequence units. If this value is 0, mesh subdivision is performed using the Midpoint subdivision method. If this value is 1, mesh subdivision is performed using the Normal Vector-based subdivision method. If this value is 2, mesh subdivision is performed using the Butterfly subdivision method. If this value is 3, mesh subdivision is performed using the Loop subdivision method. If this value is 4, mesh subdivision is performed using the LS3 method. i represents the subdivision level, and different mesh subdivision methods can be used for each subdivision level.
[0408] The mesh refinement method determined through the parameters described above may be configured to include some of the specified refinement methods and / or additional mesh refinement methods.
[0409] asve_normal_subdivision_refine_weight_numerator[i]: When mesh subdivision is performed using the normal vector-based subdivision method in sequence units, this refers to the numerator of the weight coefficient for the correction vector value.
[0410] asve_normal_subdivision_refine_weight_denominator[i]: When mesh refinement is performed using the normal vector-based refinement method at the sequence level, this value represents the denominator of the weight coefficient for the correction vector value. i represents the refinement level, and different weights can be used for each refinement level.
[0411] The encoding method according to the embodiments can determine the correction vector weight of the normal vector-based segmentation method of the sequence unit by calculating asve_normal_subdivision_refine_weight_numerator[i] / asve_normal_subdivision_refine_weight_denominator[i]. If one of the two parameters is 0, the decoder calculates and derives the weight and determines it. If neither of the two parameters is 0, the weight is calculated and used through the parsed parameter.
[0412] Fig. 40 may represent an atlas frame parameter set according to embodiments.
[0413] afve_overriden_flag: If this value is 1, it means that the parameters afve_subdivision_enable_flag, afve_Quantization_parameters_enable_flag, afve_transform_method_enable_flag and afve_transform_parameters_enable_flag are present in the atlas frame parameter set extension.
[0414] afve_subdivision_enable_flag: If this value is 1, it means that afve_subdivision_method and afve_subdivision_iteration_count are present in the atlas frame parameter set extension. If the value is 0, it can be inferred that afve_subdivision_enable_flag is not present.
[0415] afve_subdivision_iteration_count: Indicates the number of times or count of iterations to perform subdivision. A value of 0 means that subdivision is not performed.
[0416] afve_subdivision_method[i]: Indicates a parameter for determining the mesh subdivision method at the frame level.
[0417] If afve_subdivision_enable_flag is 1, frame unit parameters can be used instead of sequence unit parameters.
[0418] afve_normal_subdivision_refine_weight_numerator[i]: When mesh subdivision is performed using the normal vector-based subdivision method at the frame level, this refers to the numerator of the weight coefficient for the correction vector value.
[0419] afve_normal_subdivision_refine_weight_denominator[i]: When mesh subdivision is performed using the normal vector-based subdivision method in frame units, it refers to the denominator of the weight coefficient for the correction vector value.
[0420] When the afve_subdivision_enable_flag value is 1, frame-level parameters can be used instead of sequence-level parameters. i represents the subdivision level, and different weights can be used for each subdivision level.
[0421] The decoding method according to the embodiments can determine the correction vector weight of the normal vector-based segmentation method in units of frames by calculating afve_normal_subdivision_refine_weight_numerator[i] / afve_normal_subdivision_refine_weight_denominator[i]. If one of the two parameters is 0, the decoder calculates and derives the weight and determines it. If neither of the two parameters is 0, the weight is calculated and used through the parsed parameter.
[0422] Figure 41 may represent a mesh patch data unit according to embodiments.
[0423] mdu_parameters_override_flag: If the value is 1, it may mean that parameters mdu_subdivision_override_flag, mdu_Quantization_override_flag, mdu_transform_method_override_flag and mdu_transform_parameters_override_flag are present in the meshpatch with index patchIdx in the current atlas tile and the tile ID is equal to TileID.
[0424] mdu_subdivision_override_flag[tileID][patchIdx]: If this value is 1, it indicates that mdu_subdivision_method and mdu_subdivision_iteration_count are present in the mesh patch with index patchIdx in the current atlas tile, and the tile ID is the same as the tile ID. If this value is 0, it can be inferred that mdu_subdivision_override_flag[TileID][patchIdx] is absent.
[0425] mdu_subdivision_method[tileID][patchIdx][i]: Indicates a parameter for determining the mesh subdivision method at the Submesh (=meshpatch) level. If mdu_subdivision_override_flag[tileID][patchIdx] is 1, the corresponding Submesh (=meshpatch) unit parameter can be used instead of the frame unit parameter.
[0426] mdu_normal_subdivision_refine_weight_numerator[tileID][patchIdx][i]: When mesh subdivision is performed using the normal vector-based subdivision method in the submesh (=meshpatch) unit, this refers to the numerator of the weight coefficient for the correction vector value.
[0427] mdu_normal_subdivision_refine_weight_denominator[tileID][patchIdx][i]: When mesh subdivision is performed using the normal vector-based subdivision method in the submesh (=meshpatch) unit, it refers to the denominator of the weight coefficient for the correction vector value.
[0428] If the value of mdu_subdivision_override_flag[tileID][patchIdx] is 1, the Submesh(=meshpatch) unit parameter can be used instead of the frame unit parameter. tileID means the ID of the tile to which the submesh(=meshpatch) belongs. patchIdx means the patch index of the submesh(=meshpatch). i indicates the subdivision level, and different weights can be used for each subdivision level.
[0429] The decoding method according to the embodiments calculates mdu_normal_subdivision_refine_weight_numerator[tileID][patchIdx][i] / mdu_normal_subdivision_refine_weight_denominator[tileID][patchIdx][i] to determine the correction vector weight of the normal vector-based subdivision method of the corresponding submesh (=meshpatch) unit. If one of the two parameters is 0, the decoder calculates and derives the weight and determines it. If neither of the two parameters is 0, the weight is calculated and used through the parsed parameter.
[0430] According to embodiments, with respect to the above-described asve_normal_subdivision_refine_weight_numerator[i] and / or asve_normal_subdivision_refine_weight_denominator[i], a subdivision-related weight may be set according to known data between an encoder and a decoder according to embodiments, and in this case, the subdivision-related weight may not be signaled to the decoder.
[0431] Referring to Fig. 15, the encoder (transmitter) generates a base mesh by passing the original mesh to be encoded (transmitted) through a simplification unit and a parameterization unit. The generated base mesh is quantized. In the case of an inter-frame, a motion vector is calculated from a previously referenced restored base mesh and the motion vector is encoded. In the case of an intra-frame, the base mesh is encoded through a static mesh encoding unit. A base mesh bitstream including the encoded base mesh is output and transmitted to a decoder. Then, a displacement vector is calculated between mesh data obtained by subdividing and fitting the simplified mesh through the mesh simplification unit and mesh data restored from the previously encoded base mesh. In order to efficiently encode the calculated displacement vector, the displacement vector coordinate system can be transformed into a local coordinate system, and the displacement vector is transformed and quantized into displacement vector coefficients in the displacement vector conversion unit and encoded into a displacement vector bitstream. The process of subdividing the simplified base mesh and the method of performing the lifting transformation / inverse transformation prediction process based on the normal vector during the encoding and decoding process of the displacement vector are described in more detail below.
[0432] The mesh refinement unit can be performed n times with a simplified base mesh as input, and the number of mesh refinements can be determined by an agreement between the decoder and the decoder, or by units such as sequence, GOF (Group Of Frame), frame, and submesh, and subdivision_iteration_count information can be signaled by refinement level (lod). In addition, the refinement method to be performed at each refinement level can be determined. The subdivision_method information can be signaled as follows depending on the mesh refinement method. For example, it can be signaled as follows: Mid-point refinement method = 0, Normal vector-based refinement method (without planar constraints) = 1, Normal vector-based refinement method (with planar constraints) = 2, Butterfly refinement method = 3, Loop refinement method = 4, LS3 (Least Squares Subdivision Surfaces) method = 5. When the normal vector-based refinement method is determined, refinement can be performed for each vertex using the method in Equation 7. In the case of a method that also applies a planar constraint, refinement can be performed so that the distance between two vertices and adjacent faces becomes the minimum value according to Equation 11. When the refinement of the base mesh is complete, a fitting process is performed to make it similar to the original mesh as mentioned above, and the displacement vector, which is the coordinate difference with the result of refining the base mesh, is calculated. The calculated displacement vector can be encoded using a 2D video codec or an arithmetic coding method. Both methods perform displacement vector transformation, and depending on the embodiment, wavelet transformation, lifting transformation, etc. can be performed. When lifting transformation is performed, the lifting transformation prediction unit can perform prediction based on the normal vector as in Equations 12 and 13.As shown in Fig. 32, normal vector-based prediction weight information can be signaled as normal_subdivision_refine_weight_numerator and normal_subdivision_refine_weight_denominator by calculating the denominator and numerator, respectively. This information can be signaled via ASPS (Atlas Sequence Parameter Set), AFPS (Atlas Frame Parameter Set), and MDU (Meshpatch Data Unit). Alternatively, depending on the embodiments, the weight information may be defined as a preset value between the encoder and decoder and may not be signaled as separate parameter information.
[0433] Afterwards, the lifting transformation prediction unit can perform a process of updating the displacement vector of the vertex used for prediction through the generated residual signal. Once the displacement vector transformation is performed in this way, quantization is performed on the generated displacement vector coefficients, and encoding is performed through a 2D video encoder according to each encoding method or encoded using an arithmetic encoding method to generate a displacement vector sub-bitstream.
[0434] Finally, the texture map generation unit generates a new texture map having color information corresponding to the texture coordinates of the restored mesh, encodes the texture map through a 2D video encoder, and transmits it as a texture bitstream.
[0435] The base mesh bitstream, displacement vector bitstream, and texture bitstream generated through the aforementioned process in the encoder are generated as a single bitstream through a multiplexing unit and transmitted through a transmission unit.
[0436] Referring to FIG. 33, the decoder receives a bitstream transmitted from the encoder and performs a process of decoding a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream through a demultiplexer.
[0437] A decoder can follow the reverse process of the encoder's operation.
[0438] First, the base mesh bitstream is decoded through a motion vector decoding unit for inter-frames and a static mesh decoding unit for intra-frames. The decoded base mesh then passes through a restoration unit, where mesh refinement is applied.
[0439] The displacement vector bitstream is decoded in the reverse order of encoding to decode the displacement vector coefficients, perform inverse quantization and inverse transformation, and then inversely transform to the coordinate system to restore the mesh geometry information along with the base mesh data. The process of the normal vector-based mesh refinement unit of the decoder and the process of inversely transforming the displacement vector coefficients in the displacement vector decoding unit are described in detail below.
[0440] The decrypted base mesh undergoes subdivision through the mesh subdivision unit. Parsing subdivision_iteration_count reveals the level at which subdivision is performed, and parsing subdivision_method for each subdivision level reveals the method used. For example, if this value is 0, the subdivision is performed using the Mid-point subdivision method. If this value is 1, the subdivision is performed using the normal vector-based subdivision method (without planar constraints). If this value is 2, the subdivision is performed using the normal vector-based subdivision method (with planar constraints). If this value is 3, the subdivision is performed using the Butterfly subdivision method. If this value is 4, the subdivision is performed using the Loop subdivision method. If this value is 5, the subdivision is performed using the LS3 (Least Squares Subdivision Surfaces) method. The decoder's subdivision operation, like the encoder, can perform subdivision for each vertex using the method in Equation 7. If the method also applies planar constraints, the subdivision can be performed such that the distance between two vertices and adjacent faces is minimized according to Equation 11.
[0441] Subdivision is performed using a parsed method for each LoD level, and the performed result can be restored as final geometric information together with the final restored displacement vector in the mesh restoration unit.
[0442] The displacement vector decoding unit can determine whether the displacement vector sub-bitstream is encoded using a 2D video codec or an arithmetic coding method by parsing dispEncType, and can determine the profile information defined between the unit and the decoder by parsing profilesetIdx. If dispEncType is a 2D video codec, decoding is performed using the 2D video codec, and depacking and dequantization of the displacement vector transform coefficients are performed. Subsequently, displacement vector inverse transformation is performed. If dispEncType is an arithmetic coding method, arithmetic coding decoding is performed, and displacement vector prediction and restoration and displacement vector inverse quantization are performed. Subsequently, displacement vector inverse transformation is performed similarly.
[0443] During the displacement vector inverse transformation, the lifting inverse transformation is performed the number of times specified by subdivision_iteration_count. The lifting inverse transformation update unit updates the displacement vector of the vertex used for prediction using the parsed residual signal. Afterwards, the lifting inverse transformation prediction unit obtains the prediction result by applying Equations 12 and 13 using the parsed normal_subdivision_refine_weight_numerator and normal_subdivision_refine_weight_denominator information.
[0444] The mesh restoration unit calculates the vertex geometry of the restoration mesh by adding the restoration displacement vector to the vertices generated through the subdivision process in the mesh refinement unit, thereby restoring the final geometry. The received texture map bitstream is decoded through the texture map decoding unit, and the final restoration mesh is generated together with the previously restored geometry.
[0445] Figure 42 may represent an encoding method according to embodiments.
[0446] The encoding method according to the embodiments may include a step of encoding a base mesh of mesh data (S4200); a step of encoding a displacement of mesh data (S4210); and a step of encoding an attribute of mesh data (S4220).
[0447] Referring to FIG. 42, FIG. 17, FIG. 18, and FIG. 22 together, the FIG. 42 encoding method further includes a step of subdividing the base mesh based on normal information (which can be expressed as normal data, normal vectors, normal components, etc.) of vertices of the base mesh in order to encode the base mesh (S4200) and encode displacement (S4210), and the base mesh can be divided based on a count for the subdivision.
[0448] With respect to Equation 7, a new vertex generated based on subdivision can be generated based on the sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh, a weight, and the midpoint of the first and second vertices of the base mesh. The weight can be calculated based on a count for the subdivision. The count can mean the number of times (number) to perform subdivision on the base mesh. The level of detail (LoD) of the base mesh can mean the count. Subdivision of the base mesh can be performed as many times as the level of LoD.
[0449] The sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh can be calculated based on the correction value induced by the normal information of the first vertex and the correction value induced by the normal information of the second vertex.
[0450] Referring also to FIG. 39, with respect to bitstream syntax, the bitstream may include at least one of information indicating how to subdivide the basemesh or information indicating a count regarding the subdivision.
[0451] Referring to Equations 7 and 8, the subdivision step further includes: calculating first normal information of a first vertex of the base mesh and second normal information of a second vertex of the base mesh, adding the first normal information and the second normal information, and adding the product of the weights and the midpoints of the first and second vertices to calculate a new vertex of the base mesh, wherein the weights can be calculated based on counts for the subdivision.
[0452] The embodiments further include a computer-readable storage medium storing a bitstream generated by the method according to FIG. 44.
[0453] Embodiments further include a method comprising: obtaining a bitstream for mesh data, the bitstream being generated based on: encoding a base mesh of the mesh data; encoding a displacement of the mesh data; and encoding an attribute of the mesh data; and transmitting data including the bitstream.
[0454] Figure 43 may represent a decryption method according to embodiments.
[0455] A decoding method according to embodiments may include a step of decoding a base mesh within a bitstream (S4300); a step of decoding a displacement within a bitstream (S4310); and a step of decoding an attribute within a bitstream (S4320).
[0456] Referring to FIG. 43, FIG. 17, FIG. 18, and FIG. 22, the decoding method of FIG. 43 further includes a step of subdividing the base mesh based on normal information (which can be identified by normal data, normal vectors, normal components, etc.) of vertices of the base mesh in order to decode the base mesh (S4500) and decode displacement (S4510), and the base mesh can be divided based on a count for the subdivision.
[0457] With respect to Equation 7, a new vertex generated based on subdivision can be generated based on the sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh, a weight, and the midpoint of the first and second vertices of the base mesh. The weight can be calculated based on a count for the subdivision. The count can mean the number of times (number) to perform subdivision on the base mesh. The level of detail (LoD) of the base mesh can mean the count. Subdivision of the base mesh can be performed as many times as the level of LoD.
[0458] The sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh can be calculated based on the correction value induced by the normal information of the first vertex and the correction value induced by the normal information of the second vertex.
[0459] Referring also to FIG. 39, with respect to bitstream syntax, the bitstream may include at least one of information indicating how to subdivide the basemesh or information indicating a count regarding the subdivision.
[0460] Referring to Equations 7 and 8, the subdivision step further includes: calculating first normal information of a first vertex of the base mesh and second normal information of a second vertex of the base mesh, adding the first normal information and the second normal information, and adding the product of the weights and the midpoints of the first and second vertices to calculate a new vertex of the base mesh, wherein the weights can be calculated based on counts for the subdivision.
[0461] The decoding method of FIG. 43 can be performed by a decoding device (decoder). The decoding device includes a memory; and at least one processor connected to the memory; and the at least one processor can be configured to: decode a base mesh within a bitstream; decode a displacement within the bitstream; and decode an attribute within the bitstream.
[0462] The method and device according to the embodiments provide the following technical effects.
[0463] Current V-DMC technology uses a midpoint segmentation method, which simply segments two points to their midpoints during the segmentation process. While this technique is fast to calculate, it inevitably results in a large displacement vector value due to differences from the actual original mesh. The method and device according to the embodiments include and perform a normal vector-based segmentation method that segments the mesh to resemble the original mesh by utilizing the normal vectors of surrounding vertices rather than simply dividing the mesh by the median value. Furthermore, when performing a lifting transformation of a displacement vector, the method includes a method of predicting based on the normal vector in the transformation prediction, and a syntax signaling method that enables efficient use of this in V-DMC is also included. Performing segmentation based on the normal vector reduces the coordinate difference with the fitting mesh generated in the preprocessing step, and enables prediction to be performed more similarly to the original in the lifting transformation prediction step of the displacement vector. This in turn reduces the displacement vector value, resulting in a bit-saving effect in the displacement vector sub-bitstream.
[0464] The embodiments have been described in terms of methods and / or devices, and the descriptions of methods and devices may be applied complementarily.
[0465] For the convenience of explanation, each drawing has been described separately, but it is also possible to design a new embodiment by combining the embodiments described in each drawing. In addition, designing a computer-readable recording medium having a program recorded thereon for executing the previously described embodiments, as needed by a person skilled in the art, also falls within the scope of the embodiments. The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of the embodiments so that various modifications can be made. Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications can be made by a person skilled in the art to which the present invention pertains without departing from the gist of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[0466] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. According to embodiments, the components according to the embodiments may be implemented by separate chips. According to embodiments, at least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[0467] In this document, “ / ” and “,” are interpreted as “and / or”. For example, “A / B” is interpreted as “A and / or B”, and “A, B” is interpreted as “A and / or B”. Additionally, “A / B / C” means “at least one of A, B, and / or C”. Also, “A, B, C” means “at least one of A, B, and / or C”. Additionally, “or” in this document is interpreted as “and / or”. For example, “A or B” can mean 1) “A” only, 2) “B” only, or 3) “A and B”. In other words, “or” in this document can mean “additionally or alternatively”.
[0468] Terms such as "first," "second," etc. may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted as limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not mean the same user input signals unless the context clearly indicates otherwise.
[0469] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of terms. The expression “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0470] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in the memory.
[0471] Meanwhile, the operations according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting / receiving device may include a transmitting / receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart, and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting / receiving device.
[0472] The processor may be referred to as a controller or the like, and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above-described embodiments may be performed by the processor. Furthermore, the processor may be implemented as an encoder / decoder or the like for the operations of the above-described embodiments.
[0473] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[0474] As described above, the embodiments may be applied in whole or in part to a point cloud data transmission and reception device and system.
[0475] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.
[0476] Embodiments may include modifications / changes, which do not depart from the scope of the claims and their equivalents.
Claims
1. Step of decoding the base mesh in the bitstream; A step of decoding displacement within the bitstream; and A step of decoding an attribute within the bitstream; comprising: How to decrypt.
2. In paragraph 1, the method: Further comprising a step of subdividing the base mesh based on normal information of the vertices of the base mesh, The above base mesh is divided based on the count for the above subdivision, How to decrypt.
3. In paragraph 2, A new vertex generated based on the above subdivision is generated based on the sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh, the weight, and the midpoint of the first vertex and the second vertex of the base mesh. How to decrypt.
4. In paragraph 3, The above weight is calculated based on the count for the above subdivision, How to decrypt.
5. In paragraph 3, The sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh is calculated based on the correction value induced by the normal information of the first vertex and the correction value induced by the normal information of the second vertex. How to decrypt.
6. In paragraph 1, The bitstream includes at least one of information indicating a method of subdividing the base mesh or information indicating a count regarding the subdivision. How to decrypt.
7. In paragraph 2, The steps for subdividing the above are: Calculate the first normal information of the first vertex of the base mesh and the second normal information of the second vertex of the base mesh, It further includes a step of calculating a new vertex of the base mesh by adding the first normal information and the second normal information, multiplying the weights, and adding the midpoints of the first vertex and the second vertex. The above weight is calculated based on the count for the above subdivision, How to decrypt.
8. Memory; and At least one processor connected to the memory; wherein the at least one processor comprises: Decode the basemesh within the bitstream; Decoding displacements within the above bitstream; and configured to decode an attribute within the above bitstream; Decryption device.
9. In paragraph 8, at least one processor: Based on the normal information of the vertices of the above base mesh, the above base mesh is further configured to be subdivided; The above base mesh is divided based on the count for the above subdivision, Decryption device.
10. In paragraph 9, A new vertex generated based on the above subdivision is generated based on the sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh, the weight, and the midpoint of the first vertex and the second vertex of the base mesh. Decryption device.
11. In paragraph 10, The above weight is calculated based on the count for the above subdivision, Decryption device.
12. In paragraph 10, The sum of the normal information of the first vertex of the base mesh and the normal information of the second vertex of the base mesh is calculated based on the correction value induced by the normal information of the first vertex and the correction value induced by the normal information of the second vertex. Decryption device.
13. Step of encoding the base mesh of mesh data; a step of encoding the displacement of the above mesh data; and A step of encoding attributes of the above mesh data; comprising: Encoding method.
14. A computer-readable storage medium storing a bitstream generated by the method according to Article 13.
15. Step of obtaining bitstream for mesh data, The bitstream is generated based on the steps of encoding a base mesh of mesh data; encoding a displacement of the mesh data; and encoding an attribute of the mesh data; and A method comprising the step of transmitting data including the bitstream.
Citation Information
Patent Citations
Three-dimensional data encoding method, three-dimensional data decoding method, three-dimensional data encoding device, and three-dimensional data decoding device
US20230030392A1
Mesh vertex displacements coding
US20230412849A1
Dynamic mesh geometry refinement component adaptive coding
WO2024035762A1
Dynamic mesh compression method and device
WO2024058614A1
3D data transmission device, 3D data transmission method, 3D data reception device, and 3D data reception method
WO2024063544A1