Mesh data encoding device, mesh data encoding method, mesh data decoding device and mesh data decoding method
The V-DMC-based encoder and decoder system addresses the challenges of dynamic mesh data transmission by preprocessing and encoding mesh data into bitstreams, enhancing efficiency and quality for applications like VR, AR, and autonomous driving.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-04-02
AI Technical Summary
The challenge lies in efficiently transmitting and receiving dynamic mesh data due to the large amount of throughput required and the complexity of encoding and decoding processes, particularly in applications like VR, AR, and autonomous driving.
A method and system for encoding and decoding mesh data using a V-DMC-based encoder and decoder, which includes preprocessing to generate a base mesh and displacement vectors, and encoding these components into bitstreams using video codecs for efficient transmission and reception.
This approach enables high-quality mesh services by reducing latency and complexity in mesh data transmission, supporting applications such as autonomous driving and providing immersive 3D content across various platforms.
Smart Images

Figure KR2025015284_02042026_PF_FP_ABST
Abstract
Description
Mesh data encoding device, mesh data encoding method, mesh data decoding device and mesh data decoding method
[0001] The embodiments provide a method for providing Point Cloud content to provide various services to users, such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.
[0002] A point cloud is a collection of points in 3D space. There is a problem in that it is difficult to generate point cloud data because there are many points in 3D space.
[0003] Mesh data refers to a form of data in which connectivity information between the vertices of a mesh is added to point cloud data.
[0004] There is a problem in that a large amount of throughput is required to transmit and receive dynamic mesh data.
[0005] The technical problem according to the embodiments is to provide a point mesh data transmission device, a transmission method, a mesh data reception device, and a reception method for efficiently transmitting and receiving mesh data in order to solve the aforementioned problems, etc.
[0006] The technical problem according to the embodiments is to provide a point mesh transmission device, a transmission method, a mesh data receiving device, and a receiving method for solving latency and encoding / decoding complexity.
[0007] However, the scope of rights of the embodiments is not limited to the technical problems described above, and may be extended to other technical problems that can be inferred by a person skilled in the art based on the entire content of this document.
[0008] To achieve the above-described purpose and other advantages, the decoding method according to the embodiments may include the step of decoding a base mesh in a bitstream; the step of decoding a displacement in a bitstream; and the step of decoding an attribute in a bitstream. The encoding method according to the embodiments may include the step of encoding a base mesh of mesh data; the step of encoding a displacement of mesh data; and the step of encoding an attribute of mesh data.
[0009] A mesh data transmission method, a transmission device, a mesh data reception method, and a reception device according to the embodiments can provide a high-quality mesh service.
[0010] A mesh data transmission method, a transmission device, a mesh data reception method, and a reception device according to the embodiments can achieve various video codec methods.
[0011] The mesh transmission method, transmission device, mesh data reception method, and reception device according to the embodiments can provide general-purpose mesh content such as autonomous driving services.
[0012] Drawings are included to further understand the embodiments, and the drawings illustrate the embodiments along with descriptions related to the embodiments. For a better understanding of the various embodiments described below, one must refer to the description of the embodiments below in relation to the following drawings, which include parts corresponding to similar reference numerals throughout the drawings.
[0013] FIG. 1 shows a V-DMC-based encoder and decoder according to embodiments.
[0014] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.
[0015] FIG. 3 illustrates a V-MESH compression method according to embodiments.
[0016] FIG. 4 shows the pre-processing of V-MESH compression according to the embodiments.
[0017] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.
[0018] FIG. 6 illustrates a displacement generation process according to embodiments.
[0019] FIG. 7 illustrates the V-DMC encoding process according to the embodiments.
[0020] FIG. 8 illustrates a lifting conversion process for displacement according to embodiments.
[0021] FIG. 9 illustrates the process of packing conversion coefficients according to embodiments into a 2D image.
[0022] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.
[0023] FIG. 11 illustrates a V-DMC decoding process according to embodiments.
[0024] FIG. 12 illustrates a V-DMC encoding process according to embodiments.
[0025] FIG. 13 illustrates a V-DMC decoding process according to embodiments.
[0026] FIG. 14 shows a V-DMC encoder according to embodiments.
[0027] FIG. 15 shows a flowchart of a motion vector encoding unit according to embodiments.
[0028] FIG. 16 shows a flowchart of a group mode determination unit according to embodiments.
[0029] FIG. 17 shows a flowchart of a motion vector prediction unit according to embodiments.
[0030] FIG. 18 shows an example of vertex-specific predictor decision parameter signaling according to embodiments.
[0031] FIG. 19 shows a flowchart of a group motion vector encoding unit and a component motion encoding unit according to embodiments.
[0032] FIG. 20 shows a flowchart of a residual motion vector generation unit according to embodiments.
[0033] FIG. 21 shows a flowchart of a V-DMC decoder according to embodiments.
[0034] FIG. 22 shows a flowchart of displacement vector decoding according to embodiments.
[0035] FIG. 23 shows a flowchart of motion vector decoding according to embodiments.
[0036] FIG. 24 shows a flowchart of a group skip determination unit according to embodiments.
[0037] FIG. 25 shows a flowchart of a motion vector restoration unit according to embodiments.
[0038] FIG. 26 shows the base mesh inter-submesh unit syntax within the bitstream according to the embodiments.
[0039] FIG. 27 shows the base mesh inter-submesh data unit syntax within a bitstream according to embodiments.
[0040] FIG. 28 illustrates a mesh data encoding method according to embodiments.
[0041] FIG. 29 illustrates a mesh data decoding method according to embodiments.
[0042] Preferred embodiments of the embodiments are described in detail, and examples thereof are shown in the accompanying drawings. The detailed description below, with reference to the accompanying drawings, is intended to describe preferred embodiments of the embodiments rather than merely embodiments that may be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it is obvious to those skilled in the art that the embodiments can be practiced without these details.
[0043] Most terms used in the embodiments are selected from those commonly used in the field, but some terms are chosen at the applicant's discretion, and their meanings are described in detail in the following description as necessary. Accordingly, the embodiments should be understood based on the intended meaning of the terms, rather than their mere names or meanings.
[0044] FIG. 1 shows a V-DMC-based encoder and decoder according to embodiments.
[0045] The basic structure of the currently ongoing V-DMC (v-mesh) is shown in Fig. 1. The encoder and decoder according to Fig. 1 perform the encoding and decoding processes of media representing a dynamic mesh using V3C technology. The preprocessor converts the input dynamic mesh representation into several V3C components (base mesh, displacement set, 2D representation of attributes, and atlas). The original mesh is simplified into a base mesh. The base mesh can be encoded using any mesh codec. Displacement vectors can be represented by a profile or encoded into V3C geometric video components using any video codec via SEI messages. For example, depending on the profile, displacement vectors (displacement data) can be encoded using arithmetic coding. Attribute data may include additional attributes. For example, texture or material information may be included as additional attributes and can be encoded using any video codec. Atlas data contains information on how to perform inverse reconstruction and is provided to the V3C decoding and / or rendering system. For example, atlas data may include methods for performing subdivision of the base mesh, methods for applying displacement vectors to the vertices of the subdivided mesh, and methods for applying attributes to the reconstructed mesh.
[0046] The encoder may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a preprocessor, an atlas encoder, a basemesh encoder, a displacement vector encoder, a video encoder, and a multiplexer.
[0047] The atlas encoding unit encodes the atlas of the mesh data to generate an atlas bitstream. The basemesh encoding unit encodes the basemesh of the mesh data to generate a basemesh bitstream. The displacement vector encoding unit encodes the displacement vector of the mesh data to generate a displacement vector bitstream. The video encoding unit encodes the attributes of the mesh data to generate an attribute bitstream. The encoder generates parameter information (which may be referred to as signaling information, metadata, etc.) related to each encoding. The encoder can generate a bitstream containing parameter information, the atlas, the basemesh, the displacement vector, and / or attributes.
[0048] The decoder may be composed of a memory and at least one processor connected to the memory. The at least one processor may be configured to perform operations such as a demultiplexer, an atlas decoder, a basemesh decoder, a displacement vector decoder, and a video decoder.
[0049] The atlas decoder decodes the atlas within the bitstream. The basemesh decoder decodes the basemesh within the bitstream. The displacement vector decoder decodes the displacement vector within the bitstream. The video decoder decodes the attributes within the bitstream. The decoder can perform each decoding operation based on parameter information within the bitstream. The decoder can reconstruct dynamic mesh data based on the atlas, displacement vector, attributes, and basemesh.
[0050] Below, the operation of the V-DMC encoder and decoder of FIG. 1 is explained in more detail.
[0051] FIG. 2 shows a system for providing dynamic mesh content according to embodiments.
[0052] The system of FIG. 2 includes a point cloud data transmission device (100) and a point cloud data reception device (110) according to embodiments. The point cloud data transmission device may include a dynamic mesh video acquisition unit (101), a dynamic mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The point cloud data reception device (110) may include a receiver (111), a file / segment decapsulator (112), a dynamic mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the point cloud data transmission device according to embodiments may be interpreted as a term referring to the transmission device (100) or the dynamic mesh video encoder (hereinafter, encoder) (102). The point cloud data receiving device according to the embodiments may be interpreted as a term referring to the receiving device (110) or the dynamic mesh video decoder (hereinafter, decoder) (113).
[0053] The system of Fig. 2 can perform video-based dynamic mesh compression and decompression.
[0054] With advancements in 3D capture, modeling, and rendering, users can access various forms of 3D content, such as AR, XR, the metaverse, and holograms, across multiple platforms and devices. 3D content represents objects more sophisticatedly and realistically to enable users to enjoy immersive experiences, and for this purpose, the creation and use of 3D models require a large amount of data. Among the various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. The embodiments include a series of processing steps in a system that uses such mesh content.
[0055] First, the method for compressing dynamic mesh data originates from the V-PCC (Video-based point cloud compression) standard technology. Point cloud data consists of data containing color information along with vertex coordinates (X, Y, Z). Mesh data refers to data where connectivity information between vertices is added to this vertex data. Content can be created in a mesh data format from the outset when generating content. Point cloud data can be converted into mesh data and used by adding connectivity information.
[0056] Currently, the MPEG standards organization defines dynamic mesh data types as the following two: Category 1: Mesh data containing texture maps as color information. Category 2: Mesh data containing vertex colors as color information.
[0057] Mesh coding standards for Category 1 data are currently being developed, and standardization work for Category 2 data is also scheduled to proceed in the future. As shown in Fig. 1, the entire process for providing mesh content services may include an acquisition process, an encoding process, a transmission process, a decoding process, a rendering process, and / or a feedback process.
[0058] To provide mesh content services, 3D data acquired through multiple cameras or special cameras can be processed into a mesh data type through a series of processes and then generated as a video. The generated mesh video is transmitted after undergoing a series of processes, and at the receiving end, the received data can be processed back into a mesh video and rendered. Through this, the mesh video is provided to the user, and the user can use the mesh content according to their intention through interaction.
[0059] A mesh compression system may include a transmission device and a reception device. The transmission device can encode mesh video to output a bitstream and transmit it to the reception device via a digital storage medium or network in the form of a file or streaming (streaming segment). The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0060] The transmission device may schematically include a mesh video acquisition unit, a mesh video encoder, and a transmission unit. The receiving device may schematically include a receiver, a mesh video decoder, and a renderer. The encoder may be referred to as a mesh video / image / picture / frame encoding device, and the decoder may be referred to as a mesh video / image / picture / frame decoding device. The transmitter may be included in the mesh video encoder. The receiver may be included in the mesh video decoder. The renderer may include a display unit, and the renderer and / or the display unit may be composed of separate devices or external components. The transmission device and the receiving device may further include separate internal or external modules / units / components for a feedback process.
[0061] Mesh data represents the surface of an object using multiple polygons. Each polygon is defined by a vertex in 3D space and connectivity information indicating how those vertices are connected. It may also include vertex attributes such as color and normals. Mapping information, which enables the surface of the mesh to be mapped onto a 2D planar area, can also be included as an attribute of the mesh. Mapping can generally be described by a set of parametric coordinates, referred to as UV coordinates or texture coordinates, associated with the mesh vertices. The mesh contains a 2D attribute map, which can be used to store high-resolution attribute information such as textures, normals, and displacements.
[0062] The mesh video acquisition unit may include processing 3D object data acquired through a camera, etc., into a mesh data type having the attributes described above through a series of processes, and generating a video composed of such mesh data. In a mesh video, the attributes of the mesh, namely vertices, polygons, connectivity information between vertices, color, normals, etc., may change over time. A mesh video having attributes and connectivity information that change over time in this way can be described as a dynamic mesh video.
[0063] A mesh video encoder can encode input mesh video into one or more video streams. A single video may contain multiple frames, and a single frame may correspond to a still image or picture. In this document, the term "mesh video" may include mesh images, frames, or pictures, and the terms mesh video and mesh images, frames, or pictures may be used interchangeably. A mesh video encoder can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. To improve compression and coding efficiency, a mesh video encoder can perform a series of procedures such as prediction, transformation, quantization, and entropy coding. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0064] The file / segment encapsulation module can encapsulate encoded mesh video data and / or mesh video-related metadata into a file or the like. Here, the mesh video-related metadata may be received from a metadata processing module or the like. The metadata processing module may be included in the mesh video encoder or may be configured as a separate component / module. The encapsulation module can encapsulate the data into a file format such as ISOBMFF or process it into other forms such as DASH segments. Depending on the embodiment, the encapsulation module may include mesh video-related metadata in the file format. Mesh video metadata may be included, for example, in various levels of boxes in the ISOBMFF file format or as data within a separate track in the file. Depending on the embodiment, the encapsulation module may encapsulate the mesh video-related metadata itself into a file.
[0065] The transmission processing unit can apply transmission processing to mesh video data encapsulated according to the file format. The transmission processing unit may be included in the transmission unit or may be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. Processing for transmission may include processing for delivery via a broadcast network and processing for delivery via broadband. According to an embodiment, the transmission processing unit may receive mesh video-related metadata from the metadata processing unit in addition to mesh video data, and apply transmission processing to it.
[0066] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can extract the bitstream and transmit it to a decoding device.
[0067] The receiver can receive mesh video data transmitted by a mesh video transmission device. Depending on the transmission channel, the receiver may receive mesh video data via a broadcast network or via broadband. Alternatively, it may receive mesh video data via a digital storage medium.
[0068] The receiving processing unit can perform processing on the received mesh video data according to the transmission protocol. The receiving processing unit may be included in the receiving unit or may be configured as a separate component or module. Corresponding to the processing for transmission performed on the transmitting side, the receiving processing unit may perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit may transmit the acquired mesh video data to the decapsulation processing unit and the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the receiving processing unit may be in the form of a signaling table.
[0069] The file / segment decapsulation module can decapsulate mesh video data in file format received from the receiving module. The decapsulation module can decapsulate files based on ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to a mesh video decoder, and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to a metadata processing module. The mesh video bitstream may also contain metadata (metadata bitstream). The metadata processing module may be included in the mesh video decoder or configured as a separate component / module. The mesh video-related metadata obtained by the decapsulation module may be in the form of boxes or tracks within the file format. If necessary, the decapsulation module may receive metadata required for decapsulation from the metadata processing module. Mesh video-related metadata may be passed to a Mesh video decoder and used in the Mesh video decoding process, or passed to a renderer and used in the Mesh video rendering process.
[0070] The Mesh Video Decoder can decode video by receiving a bitstream as input and performing operations corresponding to those of the Mesh Video Encoder. The decoded Mesh Video can be displayed through a display unit. The user can view all or part of the rendered result through a VR / AR display or a standard display.
[0071] The feedback process may include the process of transmitting various feedback information obtainable during the rendering / display process to the transmitting side or to the decoder of the receiving side. Interactivity in mesh video consumption may be provided through the feedback process. According to an embodiment, head orientation information, viewport information indicating the area the user is currently viewing, etc., may be transmitted during the feedback process. According to an embodiment, the user may interact with elements implemented in a VR / AR / MR / autonomous driving environment, and in this case, information related to such interaction may be transmitted to the transmitting side or the service provider side during the feedback process. According to an embodiment, the feedback process may not be performed.
[0072] Head orientation information can refer to information regarding the user's head position, angle, movement, etc. Based on this information, viewport information—that is, information about the area the user is currently viewing within the mesh video—can be calculated.
[0073] Viewport information may be information about the area currently being viewed by the user within a mesh video. Through this, gaze analysis can be performed to determine how the user consumes the mesh video and to what extent they gaze at specific areas of the mesh video. Gaze analysis may be performed at the receiving end and transmitted to the transmitting end via a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0074] According to an embodiment, the aforementioned feedback information may not only be transmitted to the transmitting side but may also be consumed at the receiving side. That is, decoding and rendering processes at the receiving side may be performed using the aforementioned feedback information. For example, using head orientation information and / or viewport information, only the mesh video for the area currently viewed by the user may be preferentially decoded and rendered.
[0075] This document relates to dynamic mesh video compression as described above. The methods / embodiments disclosed in this document may be applied to the MPEG (Moving Picture Experts Group) Video-based Dynamic Mesh Compression Method (V-Mesh) standard or next-generation video / image coding standards. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.
[0076] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0077] In this document, "picture" or "frame" generally refers to a unit representing a single image of a specific time period.
[0078] A pixel or pel may refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the lumina component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0079] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0080] The encoding process of Fig. 2 is as follows.
[0081] Video-based dynamic mesh compression (V-Mesh) compression methods can provide a method for compressing dynamic mesh video data based on 2D video codecs such as HEVC and VVC. In the V-Mesh compression process, compression is performed by receiving the following data as input.
[0082] Input mesh: It includes the 3D coordinates (geometry) of the vertices constituting the mesh, normal information for each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between the vertices constituting the surface. The surface of the mesh can be represented by triangles or polygons of greater size, and connection information between the vertices constituting each surface is stored according to a defined shape. The input mesh can be saved in the OBJ file format.
[0083] Attribute map: (Hereafter, Texture map is used with the same meaning): It contains information on the attributes (color, normals, displacement, etc.) of a mesh and stores data in the form of mapping the mesh surface onto a 2D image. Mapping which part of the mesh (surface or vertex) corresponds to each data point in this attribute map is based on the mapping information contained in the input mesh. Since the attribute map holds data for each frame of the mesh video, it can also be referred to as an attribute map video (or simply "attribute"). In the V-Mesh compression method, the attribute map primarily contains the mesh's color information and is stored in image file formats (PNG, BMP, etc.).
[0084] Material Library File: Contains material attribute information used in the mesh, specifically information that links the input mesh with its corresponding attribute map. This is saved in the Wavefront Material Template Library (MTL) file format.
[0085] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0086] Base Mesh: By simplifying (decimating) the input mesh through a preprocessing stage, objects from the input mesh are represented using the minimum number of vertices determined according to user criteria.
[0087] Displacement: Displacement information used to represent the input mesh as similarly as possible using the base mesh, expressed in the form of 3D coordinates.
[0088] Atlas information: This is metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be generated and utilized as sub-units (sub-mesh, patch, etc.) that constitute the mesh.
[0089] Referring to FIGS. 3 to 7, a method for encoding mesh position information (vertices) is described, and referring to FIGS. 7-10 and others, a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0090] FIG. 3 illustrates a V-MESH compression method according to embodiments.
[0091] FIG. 3 illustrates the encoding process of FIG. 2, and the encoding process may include pre-processing and encoding processes. The encoder of FIG. 2 may include a pre-processor (200) and an encoder (201) as in FIG. 3. The transmitting device of FIG. 2 may be broadly referred to as an encoder, and the dynamic mesh video encoder of FIG. 2 may be referred to as an encoder. The V-Mesh compression method may include pre-processing (200) and encoding (201) processes as in FIG. 3. The pre-processor of FIG. 3 may be located in front of the encoder of FIG. 3. The pre-processor and encoder of FIG. 3 may be referred to as a single encoder.
[0092] The pre-processor can receive a static dynamic mesh and / or attribute map. The pre-processor can generate a base mesh and / or displacement through preprocessing. The pre-processor can receive feedback information from an encoder and generate a base mesh and / or displacement based on the feedback information.
[0093] The encoder can receive a base mesh, displacement, static and dynamic meshes, and / or attribute maps. The encoder can encode mesh-related data to generate a compressed bitstream.
[0094] FIG. 4 shows the pre-processing of V-MESH compression according to the embodiments.
[0095] Figure 4 shows the configuration and operation of the pre-processor of Figure 3.
[0096] FIG. 3 illustrates a process of performing preprocessing on an input mesh. The preprocessing process (200) may include four main steps: 1) generation of a Group of Frame (GoF), 2) mesh decimation, 3) UV parameterization, and 4) fitting subdivision surface (300). The pre-processor (200) may receive the input mesh, generate displacement and / or a base mesh, and transmit it to the encoder (201). The pre-processor (200) may transmit GoF information associated with GoF generation to the encoder (201).
[0097] Below, each step of Fig. 4 is explained.
[0098] GoF Generation: This is the process of creating a reference structure for mesh data. If the number of vertices, texture coordinates, vertex connectivity, and texture coordinate connectivity of the mesh in the previous frame and the current mesh are all identical, the previous frame can be set as the reference frame. In other words, if only the vertex coordinate values differ between the current input mesh and the reference input mesh, inter-frame encoding can be performed. Otherwise, the frame undergoes intra-frame encoding.
[0099] Mesh Decimation: This is the process of simplifying the input mesh to generate a simplified mesh, or base mesh. After selecting vertices to remove from the original mesh based on user-defined criteria, the selected vertices and the triangles connected to them can be removed.
[0100] In the process of performing mesh decimation, information regarding the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) is passed as input, and a decimated mesh can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0101] UV Parameterization: This is the process of mapping 3D surfaces to a texture domain for a decimated mesh. Parameterization can be performed using a UV Atlas tool. Through this process, mapping information is generated regarding where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process.
[0102] OrthoAtlas technology is a technique that generates texture coordinates using orthographic projection. In orthoAtlas technology, the processes of patch generation and patch packing are performed sequentially. First, Connected Components (CCs) are generated by splitting adjacent triangles, and then the optimal CCs are merged using a cost function to generate a patch. The cost function can measure the cost based on the degree of distortion that occurs when orthogonally projecting the patch in each direction. Finally, texture coordinates can be calculated by packing the patch that minimizes the cost function into the texture domain. In the case of orthoAtlas technology, texture coordinates can be derived in the base mesh decoder without compressing texture coordinate and texture connection information during the base mesh encoding process.
[0103] Fitting subdivision surface: This is the process of performing subdivision on a decimated mesh. User-defined methods, such as the mid-edge method, can be applied as the subdivision method. A fitting process is performed to make the input mesh and the subdivision mesh similar to each other.
[0104] This is a process of performing fitting so that the mesh obtained by subdividing the base mesh becomes similar to the surface of the input mesh. As for the subdivision method, a user-defined method such as the mid-edge method (Fig. 5), loop method, or LS3 method may be applied.
[0105] FIG. 5 illustrates a mid-edge subdivision method according to embodiments.
[0106] Figure 5 illustrates the mid-edge method of the fitting subdivision surface described in Figure 4. Referring to Figure 5, an original mesh containing 4 vertices is subdivided to create a sub-mesh. A sub-mesh can be created by creating a new vertex at the midpoint of the edge between vertices.
[0107] When a fitted subdivided mesh (hereinafter referred to as the fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The positional difference between this result and the fitted subdivided mesh for each vertex becomes the displacement for each vertex. Since displacement represents a positional difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of the Cartesian coordinate system. Depending on user input parameters, (x, y, z) coordinate values can be converted into (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0108] FIG. 6 illustrates a displacement generation process according to embodiments.
[0109] FIG. 6 illustrates in detail the displacement calculation method of a fitting subdivision surface (300) as described in FIG. 5.
[0110] An encoder and / or pre-processor according to the embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may receive a restored base mesh and generate a subdivided restored base mesh. The local coordinate system calculation unit may receive a fitted subdivided mesh and a subdivided restored base mesh and convert a coordinate system relating to the mesh to a local coordinate system. The local coordinate system calculation operation may be optional. The displacement calculation unit calculates the positional difference between the fitted subdivision mesh and the subdivided restored base mesh. For example, it may generate a positional difference value between the vertices of the two input meshes. The vertex positional difference value becomes the displacement.
[0111] The method and apparatus for transmitting point cloud data according to the embodiments can encode the point cloud as follows. The point cloud data according to the embodiments (which may be referred to simply as point cloud) may refer to data including vertex coordinates and color information. Point cloud is a term that includes mesh data, and in this document, point cloud and mesh data may be used interchangeably.
[0112] The V-Mesh compression (restoration) method according to the embodiments may include intra-frame encoding (Fig. 6) and inter-frame encoding (Fig. 7).
[0113] Intra-frame encoding or inter-frame encoding is performed based on the results of the aforementioned GoF generation. In the case of intra-frame encoding, the data to be compressed may include a base mesh, displacement, and attribute map. In the case of inter-frame encoding, the data to be compressed may include displacement, attribute map, and the motion field between the reference base mesh and the current base mesh.
[0114] FIG. 7 illustrates the V-DMC encoding process according to the embodiments.
[0115] The encoding process of Fig. 7 illustrates the encoding of Figs. 1 and 2 in detail.
[0116] A pre-processor receives an input mesh and can perform the aforementioned preprocessing. Through preprocessing, a base mesh and / or a fitted subdivided mesh can be generated. A quantizer can quantize the base mesh and / or the fitted subdivided mesh. A static mesh encoder can encode a static mesh. The static mesh encoder can generate a bitstream containing the encoded base mesh. A motion encoder can encode motion vectors for the base mesh based on inter-frame motion estimation and motion compensation for inter-prediction. An atlas encoder can encode an atlas for the vertices of the base mesh. The encoded base mesh can be reconstructed and inversely quantized through an inverse quantizer. A displacement computer receives the reconstructed mesh and, based on the fitted subdivided mesh, can generate displacement, which is the position difference. A lifting transform can receive the displacement and generate lifting coefficients. The quantizer can quantize the lifting coefficients. Depending on the encoding method, the image packing unit can pack the image based on the quantized lifting coefficients. The video encoder can encode the packed image. Depending on the encoding method, it can apply interprediction to the quantized lifting coefficients and encode the predicted lifting coefficients according to an arithmetic encoding method. The mesh restoration unit restores the deformed mesh through the restored displacement and the restored base mesh. The displacement data is restored, and the deformed mesh is restored based on the restored displacement data and the restored base mesh and provided to the attribute transfer. The attribute transfer receives the input mesh and / or input attribute map and generates an attribute map based on the restored deformed mesh. Push-pull padding can pad data into the attribute map based on a push-pull method. The color space converter can convert the space of the color component that is an attribute. The video encoder can encode the attributes.A multiplexer can generate a bitstream by multiplexing a compressed base mesh, compressed displacement, and compressed attributes.
[0117] Base Mesh Encoding: Base mesh compression methods can be divided into INTRA, INTER, and SKIP types depending on the base mesh type, and encoding can be performed in different ways for each. If the base mesh is of the INTRA type, it can be encoded using a static mesh encoding method. If the base mesh is of the INTER type, the motion field between the reference base mesh and the current base mesh can be encoded. If the current base mesh is of the SKIP type, the reference base mesh can be derived into the current base mesh.
[0118] After being encoded in the encoder, the base mesh can be subdivided into a subdivided mesh through a subdivision process. Subdivision algorithms such as mid-point subdivision and loop subdivision can be used.
[0119] Static Basemesh Encoding (Intra Basemesh Encoding): When performing intra encoding on the current basemesh, the base mesh generated during the preprocessing stage can be encoded using static mesh compression technology after undergoing a quantization process. Static mesh compression utilizes MPEG EdgeBreaker (MEB) technology, and the base mesh's vertex position information, mapping information (texture coordinates), vertex connectivity information, and normals are subject to compression.
[0120] The technology for compressing connection information can be encoded based on the edgebreaker algorithm. The edgebreaker algorithm is a technique that sequentially traverses triangles according to rules, maps symbols based on the characteristics of each triangle, and then encodes the corresponding symbols.
[0121] Techniques for compressing vertex location information can calculate predicted values based on prediction techniques such as multiple parallelogram prediction, and then encode the residual value, which is the difference between the current vertex and the predicted value.
[0122] A technique for compressing mapping information (texture coordinates) can calculate a predicted value based on a prediction technique such as stretching, and then encode the residual value, which is the difference between the current mapping information (texture coordinates) and the predicted value.
[0123] Normal compression techniques can obtain predicted values based on prediction techniques such as delta coding, multiple parallelogram prediction, and cross product-based prediction, and then encode the residual value, which is the difference between the current normal and the predicted value.
[0124] Motion Field Encoding (Inter Basemesh Encoding): Inter basemesh encoding can be performed when a one-to-one correspondence exists between the reference mesh and the current input mesh, differing only in vertex position information. When performing inter encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh—that is, the motion field—is calculated and this information is encoded. The reference base mesh is the result of quantizing already decoded base mesh data and is determined by the reference frame index.
[0125] The motion field can be encoded as is, or the predicted motion field can be calculated by averaging the motion fields of the restored vertices among those connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the current vertex's motion field value, can be encoded. This value can be encoded using entropy coding.
[0126] Displacement Encoding: After encoding the base mesh, reconstruction and inverse quantization are performed to reconstruct it. Once the base mesh is generated and subdivision is performed on it, the displacement between the result and the fitted subdivided mesh can be calculated. For effective encoding, data transform processes such as wavelet transform can be applied to the displacement information, and Figure 7 shows the process of transforming displacement information using the lifting transform in V-Mesh. The transform coefficients generated through the transformation process are quantized, and the quantized transform coefficients can be compressed through a video codec or through arithmetic coding, depending on the compression method.
[0127] When compressed through a video codec, the data is packed into a 2D image as shown in Figure 8. Transform coefficients are organized into one block for every N^2 (N*N) units, and each block can be packed in z-scan order. The number of horizontal blocks is fixed at N, while the number of vertical blocks can be determined by the number of vertices in the subdivided base mesh. Within a single block, transform coefficients can be packed by aligning them using Morton code. The packed images generate a displacement video for every GoF unit, and this displacement video can be encoded using an existing video compression codec.
[0128] When compressed via arithmetic coding, cross-frame prediction can be performed on the quantized displacement vector transformation coefficients. When cross-frame prediction is performed on the current quantized displacement vector transformation coefficients, the residual value, which is the difference between the current displacement vector transformation coefficient and the reference displacement vector transformation coefficient, can be encoded, and information about the reference target can be encoded. Depending on the displacement vector type, the quantized displacement vector transformation coefficients can be arithmetic encoded if it is of the INTRA type, and the residual value if it is of the INTER type. Arithmetic coding can be performed based on Context Adaptive Binary Arithmetic Coding (CABAC). The CABAC process can first binarize the displacement vector data and map it to a bin string. The bin string can be an output binarized into 0s and 1s, where each 0 or 1 can be a bin. Each bin can be arithmetic encoded using context information selected from the context model, and a process of updating probabilities can be performed.
[0129]
[0130] FIG. 8 illustrates a lifting conversion process for displacement according to embodiments.
[0131] FIG. 9 illustrates the process of packing conversion coefficients according to embodiments into a 2D image.
[0132] FIGS. 8-9 respectively illustrate the process of converting the displacement of the encoding process of FIG. 7 and the process of packing the conversion coefficients.
[0133] The encoding method according to the embodiments includes displacement encoding.
[0134] After encoding the base mesh through base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through reconstruction and inverse quantization. Subdivision is then performed on this reconstructed base mesh, and the displacement between the result and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated. For effective encoding, a data transform process such as a wavelet transform can be applied to the displacement information.
[0135] FIG. 8 illustrates the process of transforming displacement information using a lifting transform in V-Mesh. The transformation coefficients generated through the transformation process are quantized and then packed into a 2D image as shown in FIG. 9. The transformation coefficients are organized into one block for every 256 (=16 x 16) units, and each block can be packed in z-scan order. The number of horizontal blocks is fixed at 16, while the number of vertical blocks can be determined by the number of vertices in the subdivided base mesh. Within a single block, the transformation coefficients can be packed by aligning them using a Morton code. The packed images generate a displacement video for every GoF unit, and this displacement video can be encoded using a conventional video compression codec.
[0136] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. A first subdivision mesh generated by dividing the base mesh includes vertices generated by further dividing the edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes subdivided vertices and vertices of the base mesh (LoD0). A second subdivision mesh may be generated by dividing the first subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 containing vertices additionally generated from LoD0, and vertices additionally divided from LoD1. LoD is a Level of Detail indicating the degree of detail; as the index of the level increases, the distance between vertices becomes closer, and the level of detail increases. LoD N includes the vertices included in the previous LoDN-1 as they are. When vertices are further subdivided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1 and v2 and the subdivided vertex v. Instead of encoding the information for the current LoD N as is, the size of the bitstream can be reduced by generating residuals between the previous LoD N-1 and encoding the mesh using these residuals. The prediction process refers to the operation of predicting the current vertex v based on the previous vertices v1 and v2. Since adjacent subdivided meshes possess similar data, efficient encoding can be achieved by utilizing this property. The current vertex position information is predicted using the residuals from the previous vertex position information, and the previous vertex position information is updated using these residuals.
[0137] Referring to FIG. 9, the vertices have coefficients generated through a lifting transformation. The coefficients of the vertices related to the lifting transformation can be encoded by packing them into an image.
[0138] FIG. 10 illustrates the attribute transfer process of the V-MESH compression method according to the embodiments.
[0139] Figure 10 shows the detailed operation of the attribute transfer of the encoding of Figure 7.
[0140] The encoding according to the embodiments includes attribute map encoding.
[0141] Information regarding the input mesh is compressed through base mesh encoding, motion field encoding, and displacement encoding. The input mesh compressed during the encoding process is restored through base mesh decoding (intra frame), motion field encoding (inter frame), and displacement video decoding. The resulting restored deformed mesh (hereinafter referred to as Recon. deformed mesh) is used to compress the input attribute map, as shown in FIGS. 6 and 7. The Recon. deformed mesh possesses vertex position information, texture coordinates, and corresponding connection information, but lacks color information corresponding to the texture coordinates. Accordingly, as shown in FIG. 10, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is generated through the attribute transfer process in the V-Mesh compression method.
[0142] Attribute transfer first checks for every point P(u, v) in the 2D texture domain whether the point belongs to a texture triangle of the reconstructed deformed mesh, and if it does, determines the barycentric coordinates (α, of P(u, v) according to that triangle T). Calculate , γ). Then, the 3D vertex positions of triangle T and (α, Calculate the 3D coordinates M(x, y, z) of P(u, v) using , γ). Find the vertex coordinates M'(x', y', z') corresponding to the location most similar to the calculated M(x, y, z) in the input mesh domain, and find triangle T' containing this point. Then, in triangle T', find the coordinates of the centroid of M'(x', y', z') (α', Calculate ', γ'). The texture coordinates corresponding to the three vertices of Triangle T' and (α', Texture coordinates (u', v') are calculated using ', γ'), and color information corresponding to these coordinates is found in the input attribute map. The found color information is then assigned to the (u, v) pixel location in the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm such as the push-pull algorithm.
[0143] The new attribute map generated through attribute transfer is bundled in GoF units to form an attribute map video, which is then compressed using a video codec.
[0144] Atlas Encoding: Atlas information may be transmitted during the aforementioned process. The Atlas consists of information required during the mesh decoding and / or rendering process, and may include information required during the process of performing subdivision, displacement decoding, base mesh decoding, etc., as well as tile information, patch information, etc. Atlas data may be encoded using Exp-Golomb coding, etc.
[0145] Referring to FIG. 10, the reference relationships between the input mesh, input attribute map, restored mesh, and generated attribute map can be seen.
[0146] The decoding process of Fig. 1 can perform the reverse process of the corresponding process of the encoding process of Fig. 1. The specific decoding process is as follows.
[0147] FIG. 11 illustrates a VV-DMC decoding process according to embodiments.
[0148] Figure 11 shows the configuration and operation of the decoder of the receiving device of Figure 1.
[0149] The input bitstream can be separated into a Basemesh sub-stream, a Displacement sub-stream, an Attribute map sub-stream, and an Atlas sub-stream.
[0150] The Atlas sub-stream can be decoded through Exp-Golomb coding, and as a result, information necessary for performing decoding, tile information, patch information, etc. can be obtained.
[0151] Depending on the basemesh type, if the basemesh sub-stream is of the INTRA type, it can be decoded through a static mesh decoder based on MEB (MPEG EdgeBreaker) technology, and as a result, the connectivity information, vertex geometry information, and vertex mapping information (texture coordinates) of the base mesh can be restored.
[0152] When the texture parameterization method in the encoder is orthoAtlas, the decoder can derive mapping information (texture coordinates) and attribute information (texture) connection information using vertex coordinates. The process of deriving mapping information (texture coordinates) and connection information can generate mapping information (texture coordinates) and attribute information (texture) connection information by calculating the homography transform of each face and then projecting the vertex based on it.
[0153] If the Basemesh type is INTER type, motion information can be decoded through entropy decoding and inverse prediction processes. The restored motion information is combined with the reference Basemesh that has already been restored and stored in the buffer to generate a Reconstructed quantized basemesh for the current frame. An inverse quantization process can be performed on the restored Basemesh.
[0154] If the displacement sub-stream is compressed through a video codec according to the compression method used in encoding, it is decoded into a displacement video through the video compression codec's decoder, and then the image unpacking process is performed.
[0155] When compressed through arithmetic coding, the displacement vector bitstream can decode binarized syntax elements through arithmetic decoding, and a Contextual Probability Model (CPM) can be adaptively determined according to each bin of the syntax elements, and arithmetic decoding can be performed by predicting the probability of occurrence of the bin through the CPM. The binarized syntax elements can be decoded through inverse binarization. Quantized displacement vector transformation coefficients can be derived from the decoded syntax elements. Depending on the displacement information type, if it is INTER (where inter prediction is performed), an inverse inter prediction process is performed using reference information for the quantized displacement coefficients.
[0156] The quantized displacement coefficient is restored as displacement information for each vertex through inverse quantization, inverse transform, and coordinate system transformation processes.
[0157] The restored Base mesh and the restored Displacement information are combined to generate the final Decoded mesh. The Attribute map sub-stream is decoded through the decoder of the video compression codec used in Encoding, and then restored to the final Attribute map through processes such as color format conversion.
[0158] The restored Decoded mesh and Decoded attribute map can be utilized at the receiving end as final mesh data available to the user.
[0159] The atlas decoder decodes atlas data within the bitstream.
[0160] When mesh data within the bitstream is encoded based on inter-prediction, the motion decoder derives the motion field of the current frame's basemesh through motion estimation and compensation, based on the basemesh within the reference frame. When mesh data within the bitstream is encoded based on intra-prediction, the sectic decoder decodes the basemesh. Depending on the encoding method, it decodes displacement data by applying either arithmetic coding or video decoding. The video decoder decodes attribute data within the bitstream.
[0161] The decoding method of FIG. 11 can follow the reverse process of the encoding method according to the embodiments.
[0162] FIG. 12 illustrates a V-DMC encoding process according to embodiments.
[0163] FIG. 12 illustrates the configuration and operation of an encoder of a transmitting device such as FIG. 1 or FIG. 2. Each component of FIG. 12 corresponds to hardware, software, a processor, and / or a combination thereof.
[0164] Figure 12 shows the encoding process of V-Mesh technology.
[0165] The mesh preprocessing unit receives the original mesh as input and generates a decimated mesh. Decimation can be performed based on the number of target vertices or polygons constituting the mesh. For the decimated mesh, parameterization can be performed to generate mapping information (texture coordinates) and attribute information (texture) connection information per vertex. Additionally, floating-point mesh information can be quantized into fixed-point form. This result can be encoded as a base mesh through the static mesh encoding unit. The mesh preprocessing unit can generate additional vertices by performing mesh subdivision on the base mesh. Depending on the subdivision method, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information can be generated. The subdivided mesh can be fitted by adjusting vertex positions to resemble the original mesh, thereby generating a fitted subdivided mesh.
[0166] The base mesh generated through the mesh preprocessing unit can perform intra-encoding or inter-encoding depending on the base mesh type. If the base mesh frame undergoes intra-encoding, it can be compressed through the static mesh encoding unit. In this case, encoding can be performed on the base mesh's connectivity information, vertex geometry information, vertex texture information, normal information, etc. If the base mesh frame undergoes inter-encoding, the motion vector encoding unit is executed; using the base mesh and the reference-restored base mesh as inputs, the motion vector between the two meshes is calculated and its value encoded. The motion vector encoding unit performs connectivity-based prediction using previously encoded / decoded motion vectors as predictors and can encode the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The base mesh bitstream generated through the base mesh encoding unit is transmitted to the multiplexer.
[0167] The encoded base mesh bitstream can generate a restored base mesh through the base mesh restoration unit.
[0168] The displacement vector calculation unit can perform mesh subdivision on the reconstructed base mesh. The displacement vector can be calculated as the difference in vertex positions between the subdivided reconstructed base mesh and the fitted subdivision mesh generated by the preprocessing unit. As a result, displacement vectors can be calculated for each vertex of the subdivision mesh. The displacement vector calculation unit can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0169] The displacement vector processing unit can transform the displacement vector for effective encoding. Depending on the embodiment, the transformation may be performed using a lifting transformation, a wavelet transformation, etc. Additionally, quantization can be performed on the transformed displacement vector values, i.e., the transformation coefficients. Different quantization parameters can be applied to each axis of the transformation coefficients, and the quantization parameters can be derived by the agreement between the encoder and decoder. The quantized displacement vector transformation coefficients calculated by the displacement vector processing unit can be encoded through the displacement vector video encoding unit or the displacement vector arithmetic encoding unit, depending on the compression method.
[0170] The displacement vector video encoding unit can pack displacement vector information that has undergone transformation and quantization into 2D images. A displacement vector video can be generated by bundling the packed 2D images for each frame, and the displacement vector video can be generated for each GoF (Group of Frame) unit of the input mesh. The generated displacement vector video can be encoded using a video compression codec. The generated displacement vector video bitstream is transmitted to the multiplexer.
[0171] In the displacement vector arithmetic encoding unit, for the quantized displacement vector transformation coefficients, if the displacement vector type is INTER type, inter-frame prediction can be performed. The inter-frame prediction process may be a process of calculating the residual value, which is the difference between the current transformation coefficient and the reference transformation coefficient. The displacement vector transformation coefficient or the residual value can be encoded through the arithmetic encoding process.
[0172] The displacement vector restored through the displacement vector restoration unit and the base mesh restored and subdivided through the base mesh restoration unit are restored through the mesh restoration unit, and the restored mesh contains restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0173] The attribute information (texture map) of the original mesh can be regenerated into the attribute information (texture map) for the restored mesh through the attribute information (texture map) video generation unit. Vertex-specific color information contained in the original mesh's texture map can be assigned to the texture coordinates of the restored mesh. The texture maps regenerated for each frame can be grouped by GoF unit to generate a texture map video.
[0174] The generated texture map video can be encoded using a video compression codec through the texture map video encoding unit. The texture map video bitstream generated through encoding is transmitted to the multiplexer.
[0175] The atlas encoding unit can encode the atlas, which is additional information required for the mesh decoding and rendering processes. The generated atlas bitstream is transmitted to the multiplexer.
[0176] The generated base mesh bitstream, displacement vector bitstream, texture map bitstream, and atlas bitstream can be multiplexed into a single bitstream and transmitted to the receiver via the transmitter. Alternatively, the generated base mesh bitstream, displacement vector bitstream, texture map bitstream, and atlas bitstream can be generated into a file with one or more track data or encapsulated into segments and transmitted to the receiver (decoder) via the transmitter.
[0177] The data input unit can receive the original mesh and / or original texture map ('attribute'). The mesh preprocessing unit can simplify the original mesh to generate a base mesh and fit it to generate a refined mesh. If the mesh encoding method is inter-prediction, the motion vector encoding unit can generate motion vectors (motion fields) by referencing the restored base mesh within a previously processed reference frame and encode them based on motion estimation and compensation methods. If the mesh encoding method is intra-prediction, the static mesh encoding unit can encode the base mesh within the frame. The displacement vector calculation unit can calculate displacement vectors for vertices from the fitted refined mesh based on the restored base mesh. The displacement vector processing unit can process the displacement vectors into a form suitable for encoding. Depending on the encoding method for the displacement vectors, the displacement vectors can be encoded based on a video method or an arithmetic encoding method. The displacement vectors are restored and can be provided to the mesh restoration unit along with the restored base mesh. Based on the restored mesh, the attribute (texture map) video generation unit can generate a video for encoding the texture map using the original mesh and the texture map for the original mesh. The attributes are encoded based on the video method. The atlas is encoded by the atlas encoding unit.
[0178] FIG. 13 illustrates a V-DMC decoding process according to embodiments.
[0179] FIG. 13 corresponds to a decoder such as FIG. 1 to FIG. 3. Each component of FIG. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0180] The received Mesh bitstream is demultiplexed into a compressed base mesh bitstream, displacement vector bitstream, attribute information (texture map) bitstream, and atlas bitstream after file / segment decapsulation.
[0181] If the current mesh has inter-frame encoding applied based on the frame header information, decoding can be performed on the base mesh bitstream in the motion vector decoder. The final motion vector can be restored by using the previously decoded motion vector as a predictor and adding it to the residual motion vector decoded from the bitstream. The current base mesh can be restored by adding the decoded motion vector to the reference base mesh.
[0182] If the current mesh has in-frame encoding applied based on the frame header information, the base mesh bitstream can restore the base mesh's connectivity information, vertex geometry, texture coordinates, normal information, etc., through the static mesh decoder.
[0183] In the base mesh restoration unit, a restored base mesh can be generated by performing inverse quantization on the decoded base mesh.
[0184] Depending on the encoding codec type, if the displacement vector bitstream is encoded through a video codec, it can be decoded using the video codec and then subjected to a reverse packing process. If encoded through arithmetic coding, arithmetic decoding can be performed through the displacement vector arithmetic decoding unit, and if inter-frame prediction is performed, the current displacement vector transformation coefficient can be generated by adding the residual value to the reference displacement vector transformation coefficient through inter-frame prediction.
[0185] The displacement vector restoration unit restores the displacement vector by applying the decoded displacement vector transformation coefficients to inverse quantization and inverse transformation processes. If the restored displacement vector is a value in the local coordinate system, an inverse transformation to the Cartesian coordinate system can be performed.
[0186] The mesh restoration unit can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and connection information of the texture coordinates can be generated. The subdivided restored base mesh can be combined with the restored displacement vectors to generate the final restored mesh.
[0187] The texture map bitstream can be decoded as a video bitstream using a video codec in the texture map video decoder. The restored texture map contains color information for each vertex contained in the restored mesh, and the color value of the corresponding vertex can be retrieved from the texture map using the texture coordinates of each vertex.
[0188] The atlas bitstream can be decoded through the atlas decoding unit.
[0189] The restored mesh and texture map are displayed to the user through a rendering process using a mesh data renderer, etc.
[0190] The decoder receives the encoded bitstream and decodes the base mesh, displacement vector, attributes, and atlas within the bitstream based on the parameter information (which may be referred to as signaling information, metadata, etc.) contained within the bitstream. The decoding process may follow the inverse of the encoding process. Based on the decoded atlas, a mesh is reconstructed from the reconstructed base mesh and the reconstructed displacement mesh. Based on the reconstructed mesh and the reconstructed attributes, the mesh can be rendered.
[0191] A point cloud data encoding device and method according to the embodiments can encode mesh data and transmit a bitstream containing the encoded mesh data. A point cloud data decoding device and method according to the embodiments can receive a bitstream containing mesh data and decode the mesh data. A point cloud data encoding / decoding method / device according to the embodiments may be referred to simply as a method / device according to the embodiments. A point cloud data encoding / decoding method / device according to the embodiments may also be referred to as a mesh data encoding / decoding method / device according to the embodiments. Additionally, it may be used in this document simply as an encoding / decoding method / device.
[0192] The encoding method according to the embodiments may include an encoder of FIG. 1, a transmission device of FIG. 1, a preprocessor of the encoder of FIG. 3 to 4 and FIG. 6, an encoder of FIG. 7, a transmission device of FIG. 12, encoding of FIG. 14 to 20, syntax generation of FIG. 26 to 27, an encoding method of FIG. 28, etc.
[0193] The decoding method according to the embodiments may include a decoder of FIG. 1, a receiving device of FIG. 1, a decoder of FIG. 11, a receiving device of FIG. 13, decoding of FIG. 21 to 25, syntax acquisition of FIG. 26 to 27, a decoding method of FIG. 29, etc.
[0194] A method according to the embodiments may include a method for determining motion vector predictors based on encoding mode for V-DMC.
[0195] The method according to the embodiments is performed based on V-Mesh (Video-based dynamic mesh compression), a method for compressing 3D dynamic mesh data using a 2D video codec. In V-DMC, which is currently in the process of standardization, inter-mode occurs under specific conditions. In this inter-mode, a modified base mesh is used by applying motion information to the base mesh of the previous frame. Conventional motion vector prediction is performed by averaging the already encoded motion vectors around the current vector, and in this process, the skipped motion vector is used when calculating the average value.
[0196] The embodiments include a method for performing motion vector prediction based on V-DMC, which is a method for compressing 3D dynamic mesh data using an existing 2D video codec.
[0197] Conventional motion vector prediction is performed by averaging the already encoded motion vectors surrounding the current vector, and in this process, skipped motion vectors are used to calculate the average. The embodiments improve prediction performance per vertex by performing motion vector prediction using adjacent motion vectors that have not been skipped.
[0198] During the motion vector encoding / decoding process, when performing motion vector prediction, adjacent motion vectors that were not skipped are used to generate the predicted motion vector value.
[0199] As the predicted value becomes similar to the original motion vector value, the residual value of the motion vector decreases, improving encoding performance.
[0200] As used in this document, the term V-PCC (Video-based Point Cloud Compression) may be used interchangeably with V3C (Visual Volumetric Video-based Coding), and the two terms may be used interchangeably. Therefore, in this document, the term V-PCC may be interpreted as V3C.
[0201] FIG. 14 shows a V-DMC encoder according to embodiments.
[0202] The encoder of FIG. 14 can correspond to the encoder of FIG. 1, the transmission device of FIG. 1, the preprocessor of the encoders of FIG. 3 to 4 and FIG. 6, the encoder of FIG. 7, the transmission device of FIG. 12, the encoding of FIG. 15 to 20, the syntax generation of FIG. 26 to 27, the encoding method of FIG. 28, etc.
[0203] The encoder of FIG. 14 may be composed of a memory and at least one processor connected to the memory. Each of the at least one processors can perform the operations of the block diagrams of the encoder of FIG. 14.
[0204] The mesh simplification unit takes the original mesh as input and generates a simplified base mesh.
[0205] The input mesh can be simplified to the target number of vertices or the target number of faces.
[0206] Parameterization is performed to generate texture coordinates (UV coordinates) and texture connectivity information per vertex of the input mesh.
[0207] The mesh subdivision unit can generate additional vertices by performing subdivision on the base mesh. In this process, depending on the subdivision method, it can generate them by implicitly deriving geometric connectivity, texture coordinate connectivity, and texture coordinates.
[0208] Mesh subdivision can be performed n times by user parameters or by an encoder / decoder agreement, and according to the embodiment, the vertices of the base mesh are R0, the vertices newly generated by performing subdivision 1 time are R1, and the vertices generated by performing subdivision n times are R n When defined as LoD n It can be defined as follows.
[0209]
[0210] The mesh fitting section performs vertex position adjustments so that the subdivided mesh becomes similar to the original mesh.
[0211] According to the embodiment, the mesh simplification unit, mesh parameterization unit, mesh subdivision unit, and mesh fitting unit may be omitted, and if the processes are omitted, the original mesh may be applied as an input to the mesh quantization unit.
[0212] At this time, coordinate information of the original mesh may be applied as input to the displacement vector calculation unit, and depending on the embodiment, the displacement vector encoding process (displacement vector calculation unit, displacement vector coordinate system conversion unit, displacement vector encoding unit) may be omitted.
[0213] The mesh quantization unit can quantize geometric information (x, y, z) in floating-point form or / and texture coordinates (u, v) or / and normal information (nx, ny, nz), etc. into fixed-point form.
[0214] Depending on the embodiment, quantization for specific components may be omitted.
[0215] Static Mesh Encoding Unit: Performs encoding for the base mesh's connectivity information, vertex geometry information, vertex texture coordinates, normal information, etc.
[0216] Motion vector encoding can be performed by calculating motion vectors using the reference restored base mesh and the current base mesh as inputs.
[0217] The motion vector encoding unit can perform connection-based prediction using previously encoded / decoded motion vectors as predictors, and perform entropy encoding on the residual motion vector obtained by subtracting the predicted motion vector from the current motion vector.
[0218] Motion vector encoding can be performed at the vertex level or subgroup level depending on the embodiment.
[0219] The displacement vector calculation unit calculates the displacement vector between the mesh that performed subdivision on the fitted subdivided mesh and the restored current base mesh.
[0220] The displacement vector of the number of vertices of the subdivided mesh can be calculated through the displacement vector calculation unit.
[0221] Displacement vector coordinate system transformation unit: A transformation of a vertex displacement vector calculated in (x, y, z) space into a (normal, tangential, bi-tangential) coordinate system based on the normal vector of each vertex can be performed.
[0222] According to the embodiment, only the normal component of the (normal, tangential, bi-tangential) coordinate system may be encoded, and according to the encoder / decoder agreement, when a coordinate system transformation is applied, only the normal component may always be encoded, or the encoder may decide to signal a 1-bit flag (onlyNormFlag).
[0223] In this case, the normal vector can be calculated for each subdivided vertex based on the geometric information of neighboring vertices and / or connectivity information.
[0224] Whether to perform a displacement vector coordinate system transformation can be determined by the agreement between the decoding and encoding departments, or by transmitting a coordinate system transformation flag (applyLocalCoord) in units such as sequences, GOF (Group of frames), frames, and sub-meshes to determine whether to perform a coordinate system transformation.
[0225] The base mesh restoration unit performs base mesh restoration based on the encoding type of the current mesh (inter-frame encoding or intra-frame encoding).
[0226] When cross-frame encoding is performed, the current base mesh can be generated by adding the restoration motion vector to the reference restoration base mesh.
[0227] In this case, if the motion vector is not quantized, the motion vector restoration process is omitted, and the current base mesh can be restored using the motion vector calculated in the motion vector encoding unit.
[0228] When in-screen encoding is performed, the current base mesh can be restored by performing inverse quantization on the quantized base mesh through the mesh quantization unit.
[0229] Base mesh inverse quantization unit: The mesh quantization unit performs inverse quantization with inputs such as the reconstructed geometric information (x, y, z) or / and texture coordinates (u, v) or / and normal information (nx, ny, nz) of the reconstructed base mesh.
[0230] Depending on the embodiment, inverse quantization for a specific component may be omitted.
[0231] Displacement vector restoration unit: Decodes a bitstream that is packed into a 2D image / video and encoded through a 2D video encoder using a 2D video decoder, and performs reverse packing.
[0232] The quantized transform coefficients with inverse packing are subjected to nutrient magnetization, and an inverse transform is performed to calculate the restored displacement vector.
[0233] Mesh Restoration Unit: Subdivision is performed on the restored base mesh, which is restored by indequantization in the base mesh indequantization unit, to generate subdivided vertex position information, texture coordinates, and connection information.
[0234] Restored vertex location information is generated by adding the restored displacement vector to the detailed vertex location information.
[0235] The texture map generation unit generates the texture map of the restored mesh through the relationship between the texture coordinates and connection information of the restored mesh, the original mesh, and the texture map of the original mesh.
[0236] Texture map encoding unit: Texture maps generated through the texture map generation unit are stacked in the order of mesh frames to form a texture map video, and encoding is performed through a 2D video encoder.
[0237] According to the embodiment, if the color space of the texture map is RGB444, encoding can be performed after converting to a color space such as YUV420 or YUV444.
[0238] Referring to FIG. 14, initial 3D object data, such as the original mesh, is provided as input to the encoder. The mesh simplification unit reduces the complexity of the original mesh and converts it into a simplified form. The mesh parameterization unit assigns parameters to the simplified mesh. The mesh quantization unit quantizes the parameterized mesh to reduce the data size. The reference restoration base mesh is a base mesh that serves as the standard for restoration. The motion vector encoding unit encodes the motion vectors of the mesh. The static mesh encoding unit encodes the unchanging static mesh data. The base mesh encoding unit encodes the base mesh. The base mesh bitstream is a mesh bitstream containing the encoded base. The mesh subdivision unit finely divides the mesh to subdivide the processing units. The mesh fitting unit fits the subdivided mesh to the original data. The displacement vector calculation unit calculates the displacement vector for each part. The displacement vector coordinate system transformation unit converts the calculated displacement vector to a different coordinate system. The base mesh dequantization unit dequantizes the encoded base mesh to restore it. The mesh restoration unit reconstructs the entire mesh. The displacement vector encoder encodes the displacement vector. The displacement vector bitstream is a bitstream containing the encoded displacement vector. The texture map generator generates a texture map corresponding to the mesh. The texture map encoder encodes the generated texture map. The texture map bitstream is a bitstream containing the encoded texture map. The displacement vector restorer restores the encoded displacement vector.
[0239] Referring to FIG. 14, the effect of the base mesh decoding unit can be provided due to the motion vector encoding of the encoding process according to the embodiments. The motion vector encoding unit will be described in more detail below.
[0240] FIG. 15 shows a flowchart of a motion vector encoding unit according to embodiments.
[0241] FIG. 15 illustrates in detail the motion vector encoding operation of the encoder in FIG. 14.
[0242] The motion vector encoding unit can be performed when the reference base mesh and the current base mesh are in a 1-to-1 mapping relationship, and can perform prediction based on already encoded motion vectors around each vertex, generate residual MV by differentiating with the original MV, and transmit the residual MV by entropy coding.
[0243] According to the embodiments, as in the embodiment of FIG. 14, motion vectors undergo a process of identifying neighboring vertices for each vertex and removing duplicate vertices by a method defined according to the encoder convention, and then the motion vectors to be encoded are divided into groups, and the prediction mode is determined and the encoding process is performed on a group basis.
[0244] According to the embodiments, for group partitioning, the vertex order can be changed to a sequence such as a Morton code or a Hilbert curve based on the vertex position information of the base mesh, and then group partitioning can be performed based on the changed vertices.
[0245] The adjacent vertex configuration stores adjacent vertices for each vertex and is used to determine the predictor during the process of predicting each vertex.
[0246] According to the example, up to k adjacent vertices can be determined.
[0247] At most k can be determined by the encoder / decoder agreement or through explicit signaling.
[0248] In this case, the k adjacent vertices selected may refer to the k encoded vertices among the vertices connected to the current vertex by an edge, determined according to the encoder / decoder agreement, or the k encoded vertices closest to the current vertex.
[0249] According to the embodiment, when k vertices are already configured, a new vertex may be added instead of the last vertex in the encoding / decoding order or instead of the vertex furthest from the current vertex.
[0250] According to the embodiments, all encoded adjacent vertices can be determined as predictors.
[0251] According to the embodiments, adjacent vertex information by vertex already configured in the reference frame can be used identically in the current frame.
[0252] The duplicate vertex encoding unit can check for vertex redundancy and encode the motion vectors of duplicate vertices identically without performing encoding.
[0253] Vertex redundancy can be determined based on the vertex coordinates and motion vector values according to the encoder / decoder agreement.
[0254] According to the embodiments, whether to check for redundancy can be signaled via a flag, and if redundancy is checked, instead of encoding the motion vector for each vertex, a 1-bit flag (duplicate_flag) can be signaled to perform encoding identically to the motion vector of the duplicated vertex.
[0255] Referring to FIG. 15, the adjacent vertex configuration unit configures adjacent vertices of the mesh data. The duplicate vertex encoding unit removes or encodes duplicate vertices. The group mode determination unit determines which group mode to process the vertices into. The MV prediction unit predicts motion vectors (MV). The residual MV generation unit generates residual MV by calculating the difference between the predicted MV and the actual MV. The residual MV entropy encoding unit entropizes the generated residual MV. It determines whether the current group of mesh data is the last group. For example, it determines whether the current group is the last. When all processing is completed, the process is terminated.
[0256] Referring to FIG. 15, the motion vector encoding unit according to the embodiments may further include a Molton code alignment unit. Although the Molton code alignment unit is not illustrated in FIG. 15, referring to the encoding process of FIG. 14, the Molton code alignment unit may be further included after the mesh quantization unit. Additionally, referring to the decoding process of FIG. 21, the Molton code alignment unit may be included before or within the base mesh restoration.
[0257] FIG. 16 shows a flowchart of a group mode determination unit according to embodiments.
[0258] FIG. 16 illustrates in detail the operation of the group mode determination unit of the encoder of FIG. 14.
[0259] According to embodiments, motion vectors may be divided into groups by a method defined according to a sub / decoder agreement, and a prediction mode determination and encoding process may be performed on a group basis.
[0260] The definition of a group can vary (1 <= group_size <= N, N = number of vertices in the base mesh), and the group size can be determined by the decoding / coding agreement or through explicit signaling.
[0261] The group mode determination unit can determine a motion vector prediction method according to the sub / decoder agreement for each group, as shown in the embodiment of FIG. 16.
[0262] According to the embodiment, different prediction methods can be determined for each component (x, y, z axes) of the motion vector.
[0263] According to the embodiment, the group skip decision unit of FIG. 16 can determine whether to derive all motion vectors within a group into a representative vector on a group basis and can signal the execution status with a 1-bit flag (group_skip_flag).
[0264] According to the embodiments, the representative vector may be determined as the zero vector, or determined by using the average of the motion vectors of the previous group in the encoding / decoding order, or determined as the i-th motion vector of the previous group in the encoding / decoding order.
[0265] According to the embodiment, the component skip determination unit of FIG. 16 can determine whether to derive a motion value as a representative value for each component (x, y, z axis) of each motion vector when group_skip_flag is 0, and can signal whether to perform using a 1-bit flag (group_comp_skip_flag).
[0266] In this case, if group_skip_flag==0 and group_comp_skip_flag[0]==1 and group_comp_skip_flag[1]=1, the skip flag group_comp_skip_flag[2] of the 3rd component can be implicitly induced to 0.
[0267] According to the embodiments, the representative value may be determined as a value of 0, or determined using the component-wise average of the motion vectors of the previous group in the encoding / decoding order, or determined as the component of the i-th motion vector of the previous group in the encoding / decoding order.
[0268] Referring to FIG. 16, the group skip decision unit determines whether to omit the current group during the encoding process. The component skip decision unit determines whether to skip each component within the group. The MV prediction mode decision unit determines which mode to apply for the motion vector (MV) prediction method.
[0269] FIG. 17 shows a flowchart of a motion vector prediction unit according to embodiments.
[0270] FIG. 17 illustrates in detail the operation of the motion vector prediction unit of the encoder in FIG. 14.
[0271] The MV prediction mode determination unit can encode residual motion in units of vertex motion vectors or / and component motions and determine the prediction method through signaling.
[0272] Prediction methods may include prediction modes such as zero motion and / or average motion with neighboring vertices and / or representative values of neighboring vertices.
[0273] In this case, if k=0, prediction may not be performed using the prediction method determined at the group level.
[0274] Predictor determination unit (1 / 2): The predictor determination unit determines m vertices among k adjacent vertices as predictors.
[0275] In this case, the value of m represents the number of adjacent vertices used as predictors and can be determined according to the decoding agreement or through explicit signaling.
[0276] According to the embodiments, when determining according to the encoded order, m vertices are determined first according to the encoded order, or adjacent vertices determined last may be determined according to the last encoded order.
[0277] According to the embodiments, when determining based on vertex distance, m vertices may be determined in order of smallest difference in distance between the current vertex and adjacent vertices, or m vertices may be determined based on the distance difference between adjacent vertices.
[0278] According to the embodiments, when determining based on the motion values of adjacent vertices, the motion values may be determined as m non-zero values or m absolute values of the motion values that fall within a range.
[0279] In this case, the range is determined by a value agreed upon between the encoder and decoder, or the range can be determined by the encoder through signaling.
[0280] Predictor determination unit (2 / 2): According to the embodiment, when determining based on the motion vector encoding mode of adjacent vertices, m can be determined by whether adjacent vertices are skipped or m can be determined by whether adjacent vertices are skipped according to the component of adjacent vertices.
[0281] According to the embodiments, when determining whether to skip for each vertex, a vertex may be determined to be skipped if the vertex is encoded by a group skip method or if one or more components of the vertex (x, y, z axes) are encoded as skip.
[0282] At this time, 1 bit of memory can be used per vertex or / and 1 bit of memory per group to store whether a vertex is skipped, which can be used in the process of determining the predictor.
[0283] According to the embodiments, when determining whether to skip for each vertex component, a vertex encoded by a component skip method can be determined to be skipped.
[0284] At this time, 3 bits of memory can be used per vertex or / and 3 bits of memory per group to store whether a vertex is skipped, which can be used in the process of determining the predictor.
[0285] Referring to FIG. 17, k is derived as the number of decoded neighbor vertices (neighbor_decoded_vertex_count), where neighbor_decoded_vertex_count represents the number of adjacent decoded vertices, denoted by the variable k. It is determined whether k is 0. It is determined whether the number of adjacent vertices is 0. If k is 0, the process is terminated because it is unpredictable. The predictor determination unit selects a predictor to use. The prediction method determination unit determines a specific prediction method based on the selected predictor. The prediction processing process is completed and terminated.
[0286] Referring to FIG. 17, in the process of predicting a motion vector, the motion vector can be predicted based on the average value of already encoded motion vectors around the current vector. In this case, the average value may be calculated including vertices labeled as skips. If the motion vector is predicted by including skipped vertices, the prediction result may be inaccurate. The method according to the embodiments can predict the motion vector using vertices that have not been skipped. Vertices that have not been skipped may have valid values necessary for prediction. Since the motion vector is predicted based on adjacent vertices that have not been skipped, the prediction accuracy and prediction performance are increased.
[0287] For example, when predicting using the average value of the motion vectors of adjacent vertices, there is a technical effect that can effectively resolve the technical problem caused by the phenomenon where even vertices labeled as skips are included.
[0288] FIG. 18 shows an example of vertex-specific predictor decision parameter signaling according to embodiments.
[0289] FIG. 18 shows the syntax of parameters related to the predictor determination of FIG. 17. The parameters according to the embodiments are included in the bitstream.
[0290] According to an embodiment, the predictor determination method can determine and signal predictor determination parameters (predictor_selelct_method) in units such as sequence, GOF (Group Of Frames), frame, and submesh.
[0291] Referring to FIG. 18, the definition according to the value of the predictor_select_method is explained.
[0292] This defines various methods for selecting predictors along with code values (0, 1, 2, 3, ...). The selection method based on each number is as follows:
[0293] 0: Encoded / Decoded Order-Based Method, selects a predictor based on the encoding / decoding order of adjacent vertices. 1: Vertex Distance-Based Method, selects the closest vertex as the predictor based on the distance (proximity) between vertices. 2: Vertex Encoding Mode-Based Method, determines a predictor based on the encoding mode of the vertex (e.g., intra / inter, etc.). 3: Vertex Motion Value-Based Method, selects a predictor using the motion vector (motion value) of the vertex. Additional methods can be defined through additional values.
[0294] The prediction method determination unit determines the method for generating predicted values through the determined predictors.
[0295] According to an embodiment, when calculating the average value of m MVs for the predicted value, or / and Depending on whether rounding is used when predicting the average value, two predicted values may be used, and the two modes may be configured as separate prediction modes according to the embodiment.
[0296] According to an embodiment, when a predicted value is determined as a representative value of m MVs, the representative value may be determined as a minimum / middle / maximum value or a mode value among the values of the predictors, and may be determined as the motion value of the i-th predictor.
[0297] FIG. 19 shows a flowchart of a group motion vector encoding unit and a component motion encoding unit according to embodiments.
[0298] FIG. 19 shows the operation of the group motion vector encoding unit and the component motion encoding unit of FIG. 14 encoder in more detail.
[0299] MV prediction unit (1 / 2): The MV prediction unit can perform motion prediction in groups or vertices through the prediction unit and prediction method determined by the group mode determination unit.
[0300] If the prediction method is determined in the group skip decision unit or the component skip decision unit, motion vectors can be encoded through a process such as that shown in Fig. 20.
[0301] The representative vector determination unit of FIG. 19 determines one representative vector of the motion group.
[0302] According to the embodiments, the representative vector may be determined as the zero vector, or the average vector of the motion vectors of the previous group in order may be selected according to the encoder / decoder agreement, or it may be determined as the i-th motion vector within the previous group.
[0303] The group motion vector encoding unit of FIG. 19 encodes the motion vectors of the motion group into a determined representative vector.
[0304] The representative value determination unit of FIG. 19 determines a representative value for each component of the motion group.
[0305] According to the embodiments, the representative value may be determined as 0, or the average value of the component of the previous group in order may be selected according to the decoder agreement, or it may be determined as the i-th motion value within the previous group.
[0306] The component motion encoding section of FIG. 19 is encoded with a representative value determined for each component of the motion group.
[0307] Referring to FIG. 19, regarding the group skip method, a method for determining whether to skip at the group level is established. The group skip status is determined. For example, it is determined whether to skip the corresponding group. The representative vector determination unit determines a representative motion vector when the group is not skipped. The group motion vector encoding unit encodes the determined group representative motion vector. The component skip status is determined. The skip status is determined for each component within the group. The representative value determination unit determines the representative value of the component that is not skipped. The component motion encoding unit encodes the motion value of the determined component.
[0308] FIG. 20 shows a flowchart of a residual motion vector generation unit according to embodiments.
[0309] FIG. 20 illustrates in detail the residual motion vector generation operation following FIG. 19.
[0310] MV prediction unit (2 / 2): When the prediction method is determined by the prediction method determination unit, residual MV can be generated through a process such as that shown in FIG. 21.
[0311] The vertex prediction value determination unit of FIG. 21 determines the prediction value for each vertex of the motion group.
[0312] For each vertex, m vertices can be selected from among the adjacent vertices to determine the predicted value.
[0313] According to an embodiment, the predicted value can be determined as the average value of m vertices.
[0314] According to the embodiment, the predicted value can be determined as a representative value of adjacent vertices.
[0315] The prediction representative value determination unit of FIG. 21 determines the representative prediction value of the motion group.
[0316] According to the embodiments, a representative predicted value may be selected as 0 or / and selected as the average value of the motion vector of the previous group in the encoding / decoding order or / and selected as the i-th motion vector of the previous group.
[0317] Referring to FIG. 20, regarding the group prediction method, a method for predicting motion vectors (MV) at the group level is established. It is determined whether to predict MV. It is determined whether to predict the MV of the corresponding group. If prediction is required, the vertex prediction value determination unit determines the prediction value of an individual vertex. The prediction representative value determination unit determines the prediction value representing the entire group. The residual MV generation unit generates a residual MV by calculating the difference between the actual MV and the predicted MV.
[0318] Residual MV generation unit: Residual MV can be generated by differentiating the MV of the current group or vertex with the predicted MV signal generated by the MV prediction unit.
[0319] Residual MV Entropy Encoding Unit: Entropy encoding for the motion of each component of the residual MV can be performed through isZero, ..., isK(1<=K) flags and mv_sign, mv_rem, etc. In this case, the isZero flag indicates whether the absolute value of the current motion is 0, and isK may indicate whether the absolute value of the current motion is equal to K.
[0320] In the case of the motion vector sign (mv_sign), if the isZero flag is 0, that is, if the absolute value of the current motion is greater than 0, it can be entropy encoded.
[0321] For the motion vector remainder (mv_rem), if the isZero, ..., and isK flags are all 0, (abs(v)-K) can be encoded using an entropy encoding method such as exponential-Golomb encoding. (abs(v) = current motion absolute value)
[0322] In this case, according to the embodiment, for each flag and mv_rem, encoding may be performed using context-information-based entropy encoding or a bypass mode using a fixed probability.
[0323] FIG. 21 shows a flowchart of a V-DMC decoder according to embodiments.
[0324] FIG. 21 can correspond to FIG. 1 decoder, FIG. 1 receiving device, FIG. 11 decoder, FIG. 13 receiving device, FIG. 22 to 25 decoding, FIG. 26 to 27 syntax acquisition, FIG. 29 decoding method, etc.
[0325] Static Mesh Decoding Unit: The static mesh decoding unit can decode the connectivity information, vertex geometry information, vertex texture coordinates, normal information, etc. of the current base mesh.
[0326] Motion Vector Decoding Unit: The motion vector decoding unit can decode the current motion vector.
[0327] According to the embodiments, a connection information-based prediction can be performed using a previously decoded motion vector as a predictor, and a current motion vector can be decoded by adding the decoded residual motion vector to the predicted motion vector.
[0328] According to the embodiments, decoding of motion vectors may be performed or omitted at the vertex level or subgroup level.
[0329] Base Mesh Restoration Unit: The base mesh restoration unit restores the current base mesh based on the encoding type (cross-frame encoding or intra-frame encoding) of the current base mesh.
[0330] When in-screen encoding is performed, a reconstructed base mesh can be generated by performing inverse quantization on the connectivity information, vertex geometry information, vertex texture coordinates, normal information, etc. of the current base mesh decoded through the static mesh decoding unit.
[0331] When cross-frame encoding is performed, the motion vector decoded through the motion vector decoding unit is added to the reference restoration base mesh to decode the connectivity information, vertex geometry information, vertex texture coordinates, normal information, etc. of the current base mesh, and perform inverse quantization to generate the restoration base mesh.
[0332] In this case, if quantization for a specific component is not performed, inverse quantization for that component may be omitted.
[0333] Mesh Subdivision Unit: The mesh subdivision unit can generate additional vertices by performing subdivision on the base mesh. Depending on the subdivision method, this can be generated by implicitly deriving geometric connectivity, texture coordinate connectivity, and texture coordinates.
[0334] The mesh subdivision unit can perform subdivision through methods such as mid-edge, Loop, and Catmul & Clark depending on the embodiment.
[0335] Mesh subdivision can be performed n times by user parameters or by an encoder / decoder agreement, and according to the embodiment, the vertices of the base mesh are R0, the vertices newly generated by performing subdivision 1 time are R1, and the vertices generated by performing subdivision n times are R n When defined as LoD n It can be defined as follows.
[0336]
[0337] Referring to FIG. 21, the reference restored base mesh refers to the base mesh decoded prior to the current restored base mesh. The base mesh bitstream refers to the compressed base mesh data. The static mesh decoding unit restores the static mesh bitstream. The motion vector decoding unit decodes the motion vector data and utilizes it for mesh correction. The base mesh restoration unit reconstructs the base mesh using the decoded information. The mesh subdivision unit subdivides the base mesh into smaller units to increase precision. The mesh restoration unit combines the subdivided mesh with correction information to restore the final mesh. The restored mesh refers to the finally restored 3D mesh data. The displacement vector bitstream refers to the bitstream containing the displacement vector data. The displacement vector decoding unit restores the displacement vector from the bitstream. The displacement vector coordinate system inverse transformation unit inversely transforms the restored displacement vector to fit the coordinate system. The texture map bitstream refers to the bitstream containing the texture map data. The texture map decoding unit decodes the texture map data from the bitstream and applies it to the mesh.
[0338] FIG. 22 shows a flowchart of displacement vector decoding according to embodiments.
[0339] FIG. 22 illustrates the displacement vector decoding operation of the decoder in FIG. 21 in more detail.
[0340] Displacement vector coordinate system inverse transformation unit: The displacement vector coordinate system inverse transformation unit can perform an inverse transformation of the inversely quantized restored displacement vector to the (x, y, z) coordinate axes when the coordinate system transformation flag (applyLocalCoord) parsed in units of sequence, GoF (group of frame), frame, or sub-mesh is 1.
[0341] At this time, a normal vector per vertex is calculated based on the reconstructed vertex position information of the reconstructed base mesh, and the normal value of the newly generated vertex can be assigned by interpolating with the vertex normal vector of the reconstructed base mesh calculated for the additional vertices generated through the subdivision process. (Fig. 22)
[0342] In this case, for interpolation, the normal information of the base mesh used for segmentation can be averaged or distance-based weighted summed to perform interpolation.
[0343] According to the embodiment, after performing subdivision on the restored base mesh, normal vectors can be calculated for the vertices generated through the subdivision unit and the vertices of the base mesh. (Fig. 22)
[0344] By calculating tangential and bi-tangential vectors perpendicular to the normal vectors using the calculated normal vectors per vertex, the inverse transformation of the displacement vector coordinate system can be performed using the following formula.
[0345]
[0346] According to the embodiments, the inverse coordinate system transformation can always be performed without flag transmission.
[0347] Mesh Restoration Unit: The mesh restoration unit calculates the vertex position information of the restored mesh by adding restoration displacement vectors to the vertices generated through the subdivision process in the mesh subdivision unit.
[0348] Texture map decoding unit: The texture map decoding unit may be a process that takes a texture map bitstream as input and decodes the texture map.
[0349] Texture map decoder types can be decoded using video decoders, zero-run length decoders, arithmetic decoders, etc.
[0350] Color space conversion of the texture map can be performed according to the embodiments.
[0351] Referring to FIG. 22, the reconstructed base mesh normal calculation unit calculates the face and vertex normals of the reconstructed base mesh. The reconstructed base mesh subdivision unit subdivides the reconstructed base mesh into a finer mesh. The reconstructed base mesh subdivision unit subdivides the reconstructed base mesh and provides it to the normal calculation path. The subdivided vertex normal calculation unit calculates the normals of the vertices newly created by subdivision. The subdivided vertex normal interpolation unit interpolates the subdivided vertex normals using the base mesh normals and surrounding vertex information.
[0352] FIG. 23 shows a flowchart of motion vector decoding according to embodiments.
[0353] FIG. 23 illustrates the operation of motion vector decoding of the decoder in FIG. 21 in more detail.
[0354] FIG. 23 is an example of a motion vector decoding unit, and depending on the examples, some steps may be added, omitted, or the order changed.
[0355] According to an embodiment, for group partitioning, the vertex order can be changed to a sequence such as a Morton code or a Hilbert curve based on the vertex position information of the base mesh, and then group partitioning can be performed based on the changed vertices.
[0356] Adjacent Vertex Configuration: The adjacent vertex configuration stores adjacent vertices for each vertex and is used to determine the predictor during the process of predicting each vertex.
[0357] According to the embodiments, the predictors are determined as adjacent vertices equal to the number of adjacent vertices explicitly or implicitly derived.
[0358] In this case, the k adjacent vertices selected may refer to the maximum k decoded vertices among the vertices connected to the current vertex by an edge, determined according to the decoder agreement, or the maximum k decoded vertices closest to the current vertex.
[0359] According to the embodiment, the method of adding a new vertex when K are already configured may be added instead of the last vertex in the encoding / decoding order, or added instead of the vertex furthest from the current vertex, or determined according to the motion vector encoding mode.
[0360] According to the embodiments, all decoded adjacent vertices can be determined as predictors.
[0361] According to the embodiments, adjacent vertex information by vertex already configured in the reference frame can be used identically in the current frame.
[0362] Duplicate Vertex Encoding Unit: The duplicate vertex encoding unit checks for vertex redundancy and can decode the motion vectors of duplicate vertices identically without performing decoding.
[0363] Vertex redundancy can be determined based on the vertex coordinates and motion vector values according to the encoder / decoder agreement.
[0364] According to the embodiments, whether to check for duplication can be signaled through a flag, and if duplication is checked, instead of decoding the motion vector for each vertex, a 1-bit flag (duplicate_flag) can be signaled to perform decoding identical to the motion vector of the duplicated vertex.
[0365] Referring to FIG. 23, the adjacent vertex configuration unit finds adjacent vertices and creates groups. The duplicate vertex encoding unit compresses and encodes duplicate vertex information within the group. The group mode determination unit determines which processing mode to apply to each group. The residual MV entropy decoding unit entropies the residual data of the motion vector. The MV prediction unit predicts the motion vector based on surrounding information and the group mode. The MV restoration unit combines the predicted motion vector with the residual information to restore the final motion vector. It determines whether the current group is the last group. When all group processing is completed, the process ends.
[0366] FIG. 24 shows a flowchart of a group skip determination unit according to embodiments.
[0367] FIG. 24 shows the operation of the group skip decision unit of the decoder FIG. 22.
[0368] According to embodiments, motion vectors may be divided into groups by a method defined according to the encoder / decoder agreement, and the prediction mode may be determined and the decoding process performed on a group basis.
[0369] The definition of a group can vary (1 <= group_size <= N, N = number of vertices in the base mesh), and the group size can be determined by the decoding / coding agreement or through explicit signaling.
[0370] Group mode determination unit: As shown in the embodiment of FIG. 24, the group mode determination unit can determine the motion vector prediction method on a group basis according to the sub / decoder agreement for each group.
[0371] According to the embodiment, different prediction methods can be determined for each component (x, y, z axes) of the motion vector.
[0372] According to the embodiment, the group skip decision unit of FIG. 24 may determine whether to derive all motion vectors within a group into a representative vector on a group basis, and the decision on whether to perform may be determined by parsing a 1-bit flag (group_skip_flag).
[0373] According to the embodiments, the representative vector may be determined as the zero vector, or determined using the average of the motion vectors of the previous group in the encoding / decoding order, or determined as the i-th motion vector of the previous group in the encoding / decoding order.
[0374] According to the embodiment, the component skip determination unit of FIG. 24 may determine whether to derive a motion value as a representative value for each component (x, y, z axis) of each motion vector when group_skip_flag is 0, and the execution status may be determined by parsing a 1-bit flag (group_comp_skip_flag).
[0375] In this case, if group_skip_flag==0 and group_comp_skip_flag[0]==1 and group_comp_skip_flag[1]=1, the skip flag group_comp_skip_flag[2] of the 3rd component can be implicitly induced to 0.
[0376] According to the embodiments, the representative value may be determined as a value of 0, or determined using the component-wise average of the motion vectors of the previous group in the encoding / decoding order, or determined as the component of the i-th motion vector of the previous group in the encoding / decoding order.
[0377] Referring to FIG. 24, the group skip determination unit determines whether the entire specific group can be skipped during the encoding / decoding process. The component skip determination unit determines whether each motion vector (MV) component within the group can be skipped. The MV prediction mode determination unit finally selects the prediction mode of the motion vector based on whether to skip and adjacent information.
[0378] Referring to FIG. 17, the MV prediction mode determination unit decodes the residual motion in units of the vertex motion vector and / or component motion, parses the prediction method, and performs the prediction. Prediction modes such as zero motion and / or the average motion with neighboring vertices and / or representative values of neighboring vertices may be used as prediction methods. In this case, when k=0, the prediction may not be performed using the prediction method determined in units of groups.
[0379] The predictor decision unit determines m vertices out of k adjacent vertices as predictors.
[0380] In this case, the value of m represents the number of adjacent vertices used as predictors and can be determined according to the decoding agreement or through explicit signaling.
[0381] According to the embodiments, when determining according to the decoded order, m vertices are determined first according to the decoded order, or adjacent vertices determined last may be determined according to the last decoded order.
[0382] According to the embodiments, when determining based on vertex distance, m vertices may be determined in order of smallest difference in distance between the current vertex and adjacent vertices, or m vertices may be determined based on the distance difference between adjacent vertices.
[0383] According to the embodiments, when determining based on the motion values of adjacent vertices, the motion values may be determined as m non-zero values or m absolute values of the motion values that fall within a range.
[0384] In this case, the range can be determined by a value agreed upon between the encoder and decoder, or by signaling the range determined by the encoder.
[0385] According to the embodiments, when determining based on the motion vector decoding mode of adjacent vertices, m can be determined by whether adjacent vertices are skipped or by whether adjacent vertices are skipped per component.
[0386] According to the embodiments, when determining whether to skip for each vertex, a vertex may be determined to be skipped if the vertex is decoded by the group skip method or if one or more components of the vertex (x, y, z axes) are decoded as skip.
[0387] At this time, 1 bit of memory can be used per vertex or / and 1 bit of memory per group to store whether a vertex is skipped, which can be used in the process of determining the predictor.
[0388] According to the embodiments, when determining whether to skip for each vertex component, the vertex component can be determined to be skipped by decoding the vertex using the component skip method.
[0389] At this time, 3 bits of memory can be used per vertex or / and 3 bits of memory per group to store whether a vertex is skipped, which can be used in the process of determining the predictor.
[0390] Referring to FIG. 19, according to embodiments, a predictor determination method can be determined by receiving a predictor determination parameter (predictor_selelct_method) in units such as a sequence, GOF (Group Of Frames), frame, or submesh.
[0391] The prediction method determination unit determines the method for generating predicted values through the determined predictors.
[0392] According to an embodiment, when calculating the average value of m MVs for the predicted value, or / and Depending on whether rounding is used when predicting the average value, two predicted values may be used, and the two modes may be configured as separate prediction modes according to the embodiment.
[0393] According to an embodiment, when a predicted value is determined as a representative value of m MVs, the representative value may be determined as a minimum / middle / maximum value or a mode value among the values of the predictors, and may be determined as the motion value of the i-th predictor.
[0394] The MV prediction unit can perform motion prediction in groups or vertices through the prediction unit and prediction method parsed from the group mode determination unit.
[0395] When the prediction method is parsed in the group skip decision unit or evaded in the component skip decision unit, motion vectors can be decoded through a process such as that shown in Fig. 15.
[0396] The representative vector determination unit of FIG. 20 determines one representative vector of the motion group.
[0397] According to the embodiments, the representative vector may be determined as the zero vector, or the average vector of the motion vectors of the previous group in order may be selected according to the encoder / decoder agreement, or it may be determined as the i-th motion vector within the previous group.
[0398] The group motion vector decoding unit of FIG. 19 decodes the motion vectors of the motion group into a determined representative vector.
[0399] The representative value determination unit of FIG. 19 determines a representative value for each component of the motion group.
[0400] According to the embodiments, the representative value may be determined as 0, or the average value of the component of the previous group in order may be selected according to the decoder agreement, or it may be determined as the i-th motion value within the previous group.
[0401] The component motion decoding unit of FIG. 19 is decoded using a representative value determined for each component of the motion group.
[0402] FIG. 25 shows a flowchart of a motion vector restoration unit according to embodiments.
[0403] FIG. 25 is a flowchart of the motion vector restoration operation of the decoder of FIG. 21.
[0404] MV prediction unit: When the prediction method is parsed in the prediction method determination unit, MV restoration can be performed through a process such as that shown in Fig. 25.
[0405] The vertex prediction value determination unit of FIG. 25 determines the prediction value for each vertex of the motion group.
[0406] For each vertex, m vertices can be selected from among the adjacent vertices to determine the predicted value.
[0407] According to an embodiment, the predicted value can be determined as the average value of m vertices.
[0408] According to the embodiment, the predicted value can be determined as a representative value of adjacent vertices.
[0409] The prediction representative value determination unit of FIG. 25 determines the representative prediction value of the motion group.
[0410] According to the embodiment, a representative predicted value may be selected as 0 or / and selected as the average value of the motion vector of the previous group in the encoding / decoding order or / and selected as the i-th motion vector of the previous group.
[0411] Referring to FIG. 26, the group prediction method determines how to predict motion vectors (MV) at the group level. It determines whether to perform MV prediction. For example, it determines whether to apply motion vector prediction to the group. The vertex prediction value determination unit determines the motion vector prediction value for each vertex when the prediction is applied. The prediction representative value determination unit selects a motion vector prediction value that represents the entire group. The MV restoration unit restores the final motion vector based on the determined prediction value and residual information.
[0412] The encoding method according to the embodiments (encoder of FIG. 1, transmission device of FIG. 1, preprocessor of encoder of FIG. 3 to 4 and 6, encoder of FIG. 7, transmission device of FIG. 12, encoding of FIG. 14 to 20, encoding method of FIG. 28, etc.) can generate syntax of FIG. 26 to 27, and generate and transmit a bitstream including encoded mesh data and syntax information.
[0413] A decoding method according to embodiments (decoder of FIG. 1, receiving device of FIG. 1, decoder of FIG. 11, receiving device of FIG. 13, decoding of FIG. 21 to 2526, decoding method of FIG. 29, etc.) can receive a bitstream, obtain syntax of FIG. 26 to 27 from the bitstream, and decode mesh data.
[0414] The syntax and semantics of parameter information within a bitstream are explained below.
[0415] FIG. 26 shows the base mesh inter-submesh unit syntax within the bitstream according to the embodiments.
[0416] FIG. 27 shows the base mesh inter-submesh data unit syntax within a bitstream according to embodiments.
[0417] Basemesh Morton mode reordering flag (bm _Morton_code_reordering_flag): bm_Morton_code_reordering_flag is a flag that determines whether to change the decoding order of vertices of a submesh based on Morton code when decoding motion vectors at the submesh level. If 0, the decoding order of vertices is determined according to the decoding order of connectivity information of the reference mesh; if 1, the decoding order of vertices is determined according to the sorted order after converting the position information of the vertices of the reference mesh into Morton code.
[0418] Basemesh motion vector predictor count (bm_mv_predictor_count): bm_mv_predictor_count is a flag that determines the number of adjacent vertices to be used as predictors for each vertex when encoding / decoding motion vectors at the submesh level; if it is 0, it determines all possible adjacent vertices as predictors.
[0419] Basemesh Inter-Submesh Data Unit Motion Vector Predictor Selection Flag (bmidu_mv_predictor_select_flag): bmidu_mv_predictor_select_flag is a flag that determines how to select a predictor when performing vertex prediction during motion vector encoding / decoding at the submesh level. If 0, the predictor is determined based on the encoding / decoding order of the vertices; if 1, the predictor is determined based on the distance between vertices; if 2, the predictor is determined based on the encoding mode of adjacent vertices; and if 3, the predictor is determined based on the motion value of adjacent vertices. The name bmidu_mv_predictor_select_flag may be referred to as bmidu_mv_predictor_select_method.
[0420] FIG. 28 illustrates a mesh data encoding method according to embodiments.
[0421] The encoding method according to the embodiments may include the step of encoding a base mesh of mesh data (S2800); the step of encoding a displacement of mesh data (S2810); and / or the step of encoding attributes of mesh data (S2820), etc.
[0422] The steps of encoding the base mesh of the mesh data (S2800), encoding the displacement of the mesh data (S2810), and encoding the attributes of the mesh data (S2820) can correspond to the base mesh encoding, displacement encoding, attribute encoding operations, etc. described in the encoder of FIG. 1, the transmission device of FIG. 1, the preprocessor of the encoder of FIG. 3 to 4 and FIG. 6, the encoder of FIG. 7, the transmission device of FIG. 12, and the encoding of FIG. 14 to 20, etc.
[0423] Based on the step of encoding the base mesh of the mesh data (S2800), the step of encoding the displacement of the mesh data (S2810), and the step of encoding the attributes of the mesh data (S2820), the syntax of FIGS. 26 and FIGS. 27, etc., can be generated and a bitstream can be generated.
[0424] The encoding method of FIG. 28 can be further configured as follows by referring to the encoding of FIG. 14.
[0425] The step of encoding the base mesh may include: a step of quantizing the base mesh; a step of encoding the geometry of the vertices of the base mesh, the attributes of the vertices, and the connectivity information of the vertices; or a step of encoding motion vectors for the base mesh based on a reference base mesh of the base mesh.
[0426] The encoding method of Fig. 28 can be further configured as follows by referring to the motion vector encoding of Fig. 15.
[0427] The step of encoding motion vectors includes deriving adjacent vertices for a vertex of a base mesh and deriving groups for the vertices, and based on the groups, motion vectors for the base mesh can be encoded.
[0428] The encoding method of FIG. 28 can be further configured as follows by referring to FIG. 17 together.
[0429] The step of encoding motion vectors includes deriving predictors from adjacent vertices for a vertex of the base mesh, and the predictors may include vertices excluding those skipped during the prediction process within the group.
[0430] The encoding method of FIG. 28 can be performed by an encoder. The encoder includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: encode the base mesh of the mesh data; encode the displacement of the mesh data; and encode the attributes of the mesh data.
[0431] The embodiments further include a computer-readable storage medium for storing a bitstream generated by the method according to FIG. 28.
[0432] The embodiments may include the step of obtaining a bitstream for mesh data, the bitstream being generated based on the step of encoding a base mesh of the mesh data; the step of encoding a displacement of the mesh data; and the step of encoding attributes of the mesh data; and the step of transmitting data including the bitstream.
[0433] FIG. 29 illustrates a mesh data decoding method according to embodiments.
[0434] The decoding method according to the embodiments may include the step of decoding a base mesh in a bitstream (S2900); the step of decoding a displacement in a bitstream (S2910); and / or the step of decoding an attribute in a bitstream (S2920), etc.
[0435] The steps of decoding the base mesh (S2900), decoding the displacement (S2910), and decoding the attribute (S2920) can correspond to the base mesh encoding, displacement encoding, attribute encoding operations, etc. described in the decoder of FIG. 1, the receiving device of FIG. 1, the decoder of FIG. 11, the receiving device of FIG. 13, and the decoding of FIG. 21 to 25, respectively.
[0436] The steps of decoding the base mesh (S2900), decoding the displacement (S2910), and decoding the attribute (S2920) can each perform a decoding operation based on the syntax of FIGS. 26 to 29.
[0437] Methods in Fig. 28 and Fig. 29 can correspond to each other as inverse processes.
[0438] The mesh decoding operation regarding the decoding method of FIG. 29 can be described as follows with reference to FIG. 21 and FIG. 22 together.
[0439] The step of decoding the base mesh (S2900) includes: decoding the geometry of the vertices of the base mesh, the attributes of the vertices, and the connectivity information of the vertices; or decoding the motion vectors for the base mesh based on the reference base mesh of the base mesh; and the step of decoding the displacement (S2910) may include: inversely transforming the coordinate system for the displacement based on the normal vector of the base mesh.
[0440] Looking at the mesh decoding operation regarding the decoding method of FIG. 29, it can be further described as follows by referring together with the motion vector decoding of FIG. 24.
[0441] The step of decoding motion vectors includes deriving adjacent vertices for a vertex of a base mesh and deriving groups for the vertices, and based on the groups, motion vectors for the base mesh can be decoded.
[0442] Looking at the mesh decoding operation regarding the decoding method of FIG. 29, the group skip determination unit of FIG. 24, group_skip_flag, and group_comp_skip_flag can be described as follows by referring together.
[0443] Representative vectors for motion vectors within a group can be derived, or representative values for components of motion vectors within a group can be derived.
[0444] Looking at the mesh decoding operation regarding the decoding method of FIG. 29, it can be further described as follows by referring to motion vector prediction using adjacent motion vectors that were not skipped in FIG. 17.
[0445] The step of decoding motion vectors includes deriving predictors from adjacent vertices for a vertex of the base mesh, and the predictors may include vertices excluding vertices skipped in the prediction process within the group.
[0446] Looking at the mesh decoding operation regarding the decoding method of FIG. 29, it can be further described as follows by referring together to the group skip and component skip in FIG. 19.
[0447] Based on skipping for a group, a representative vector for the group is derived, or based on skipping for a component of a vertex within a group, a representative value for a vertex within a group is derived, and a motion vector can be predicted based on vertices excluding the skipped vertices.
[0448] Regarding the mesh decoding operation for the decoding method of FIG. 29, it can be further described as follows by referring together with FIG. 26 bm_Morton_code_reordering_flag.
[0449] The bitstream includes syntax regarding the base mesh, and the syntax regarding the base mesh may include a flag indicating whether a motion vector for the base mesh is derived based on a Molton code regarding the vertices of the submesh for the base mesh.
[0450] The decoding method of FIG. 29 can be performed by a decoder. The decoder includes a memory; and at least one processor connected to the memory; and the at least one processor may be configured to: decode a base mesh in a bitstream; decode a displacement in a bitstream; and decode an attribute in a bitstream.
[0451] The encoding / decoding method and apparatus according to the embodiments provide the following technical effects.
[0452] For each vertex, motion vectors of adjacent vertices that were not skipped are determined as predictors to generate predicted values similar to the original motion vectors, thereby increasing prediction performance, reducing the bit size for motion vectors, and improving encoding performance.
[0453] The embodiments have been described in terms of methods and / or devices, and the description of the methods and the description of the devices may be applied complementarily.
[0454] Although the drawings have been described separately for the convenience of explanation, it is also possible to design a new embodiment by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by a person skilled in the art, falls within the scope of the claims of the embodiments. The apparatus and method according to the embodiments are not limited to the configuration and method of the embodiments described above; rather, the embodiments may be configured by selectively combining all or part of each embodiment to allow for various modifications. Although preferred embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above. It is not only possible for a person skilled in the art to make various modifications without departing from the essence of the embodiments claimed in the claims, but such modifications should not be understood individually from the technical concept or perspective of the embodiments.
[0455] Various components of the device of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented as a single chip, for example, a single hardware circuit. Depending on the embodiments, the components according to the embodiments may each be implemented as separate chips. Depending on the embodiments, at least one of the components of the device according to the embodiments may be composed of one or more processors capable of executing one or more programs, and one or more programs may include instructions for performing or executing any one or more of the operations / methods according to the embodiments. Executable instructions for performing the methods / operations of the device according to the embodiments may be stored in non-transient CRMs or other computer program products configured to be executed by one or more processors, or may be stored in transient CRMs or other computer program products configured to be executed by one or more processors. Additionally, memory according to the embodiments may be used as a concept that includes not only volatile memory (e.g., RAM, etc.) but also non-volatile memory, flash memory, PROM, etc. In addition, it may also include implementation in the form of carrier waves, such as transmission over the Internet. Furthermore, processor-readable recording media are distributed across networked computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[0456] In this document, “ / ” and “,” are interpreted as “and / or.” For example, “A / B” is interpreted as “A and / or B,” and “A, B” is interpreted as “A and / or B.” Additionally, “A / B / C” means “at least one of A, B and / or C.” Also, “A, B, C” means “at least one of A, B and / or C.” Additionally, in this document, “or” is interpreted as “and / or.” For example, “A or B” may mean 1) “A” alone, 2) “B” alone, or 3) “A and B.” In other words, “or” in this document may mean “additionally or alternatively.”
[0457] Terms such as "first," "second," etc., may be used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, the first user input signal may be referred to as the second user input signal. Similarly, the second user input signal may be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not mean the same user input signals unless clearly indicated in the context.
[0458] The terms used to describe the embodiments are intended for the purpose of describing specific embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless explicitly indicated in the context. Expressions of and / or are used to mean including all possible combinations between the terms. Expressions of include describe the presence of features, numbers, steps, elements, and / or components and do not imply the exclusion of additional features, numbers, steps, elements, and / or components. Conditional expressions such as "if" or "when" used to describe the embodiments are not limited to being optional. It is intended to be interpreted as "when a specific condition is satisfied," "when a related action is performed in response to a specific condition," or "when a related definition is interpreted."
[0459] Additionally, operations according to the embodiments described herein may be performed by a transmitting and receiving device including memory and / or a processor, depending on the embodiments. The memory may store programs for processing / controlling operations according to the embodiments, and the processor may control various operations described in this document. The processor may be referred to as a controller, etc. Operations in the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in the processor or in memory.
[0460] Meanwhile, the operation according to the embodiments described above may be performed by a transmitting device and / or a receiving device according to the embodiments. The transmitting and receiving device may include a transmitting and receiving unit for transmitting and receiving media data, a memory for storing instructions (program code, algorithm, flowchart and / or data) for a process according to the embodiments, and a processor for controlling the operations of the transmitting and receiving devices.
[0461] The processor may be referred to as a controller, etc., and may correspond, for example, to hardware, software, and / or a combination thereof. The operation according to the embodiments described above may be performed by the processor. Additionally, the processor may be implemented as an encoder / decoder, etc., for the operation of the embodiments described above.
[0462] As described above, the relevant details have been explained in the best mode for carrying out the embodiments.
[0463] As described above, the embodiments may be applied wholly or partially to point cloud data transmission and reception devices and systems.
[0464] Those skilled in the art may make various changes or modifications to the embodiments within the scope of the embodiments.
[0465] The embodiments may include modifications / variations, and such modifications / variations do not exceed the scope of the claims and their equivalents.
Claims
1. Step of decoding the basemesh within the bitstream; A step of decoding displacement within the bitstream; and A step of decoding attributes within the bitstream; comprising Decryption method.
2. In Paragraph 1, The step of decoding the above base mesh is: The method comprises: a step of decoding the geometry of the vertices of the base mesh, the attributes of the vertices, and the connectivity information of the vertices; or a step of decoding a motion vector relating to the base mesh based on a reference base mesh of the base mesh. The above displacement decoding step is: A step comprising inversely transforming the coordinate system regarding the displacement based on the normal vector of the base mesh. Decryption method.
3. In Paragraph 2, The step of decoding the above motion vector is: Deriving adjacent vertices for a vertex of the above base mesh, and Includes inducing a group for the above vertex, Motion vectors for the base mesh are decoded based on the above group, Decryption method.
4. In Paragraph 3, A representative vector for motion vectors within the above group is derived, or a representative value for a component of a motion vector within the above group is derived, Decryption method.
5. In Paragraph 3, The step of decoding the above motion vector is: It includes deriving a predictor from adjacent vertices for a vertex of the above base mesh, and The above predictor includes vertices excluding those skipped in the prediction process within the above group, Decryption method.
6. In Paragraph 5, Based on the skip for the above group, a representative vector for the above group is derived, or Based on the skip for the component of the vertex within the group, a representative value for the vertex within the group is derived, and The above motion vector is predicted based on vertices excluding the skipped vertices, Decryption method.
7. In Paragraph 1, The above bitstream includes syntax regarding the basemesh, and The syntax regarding the above base mesh is A flag indicating whether a motion vector for the base mesh is induced based on a Molton code regarding the vertices of the submesh for the base mesh, Decryption method.
8. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Decoding the basemesh within the bitstream; Decoding displacement within the above bitstream; and Configured to decode attributes within the above bitstream, Decoding device.
9. Step of encoding the base mesh of the mesh data; A step of encoding the displacement of the above mesh data; and A step of encoding attributes of the above mesh data; comprising Encoding method.
10. In Paragraph 9, The step of encoding the above basemesh is: A step of quantizing the above base mesh; A step of encoding the geometry of the vertices of the base mesh, the attributes of the vertices, and the connectivity information of the vertices; or a step of encoding a motion vector relating to the base mesh based on a reference base mesh of the base mesh; comprising Encoding method.
11. In Paragraph 10, The step of encoding the above motion vector is: Deriving adjacent vertices for a vertex of the above base mesh, and Includes inducing a group for the above vertex, Motion vectors for the base mesh are encoded based on the above group, Encoding method.
12. In Paragraph 11, The step of encoding the above motion vector is: It includes deriving a predictor from adjacent vertices for a vertex of the above base mesh, and The above predictor includes vertices excluding those skipped in the prediction process within the above group, Encoding method.
13. Memory; and At least one processor connected to the memory; comprising, wherein the at least one processor: Encode the base mesh of the mesh data; Encoding the displacement of the above mesh data; and Configured to encode the attributes of the above mesh data, Encoding device.
14. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 9.
15. Step of acquiring a bitstream for mesh data, The bitstream is generated based on the steps of: encoding the base mesh of the mesh data; encoding the displacement of the mesh data; and encoding the attributes of the mesh data; and A method comprising the step of transmitting data including the bitstream above.
Citation Information
Patent Citations
Submesh coding for dynamic mesh coding
US20240244232A1
Integrating duplicated vertices and vertices grouping in mesh motion vector coding
US20240244260A1
3D data transmission device, 3D data transmission method, 3D data reception device, and 3D data reception method
WO2024049197A1
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
WO2024085653A1
Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method
WO2024191192A1