Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method

By encoding mesh data with displacement and texture maps into one frame using scalable video codecs, the method addresses the challenges of high processing requirements and latency in 3D data transmission, enabling efficient and flexible 3D service delivery.

WO2025211851A1PCT designated stage Publication Date: 2025-10-09LG ELECTRONICS INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/004568
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-05
Filing Date
2025-04-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

The sheer number of points in 3D space makes it difficult to generate and process point cloud or mesh data, leading to high processing requirements for transmission and reception of 3D data such as point cloud or mesh data, along with issues of latency and encoding/decoding complexity.

Method used

A method and device for efficiently transmitting and receiving mesh data by encoding displacement information and a texture map into one frame using scalable video codecs, allowing for scalable coding and minimizing synchronization problems through packing geometry and texture data into one frame.

Benefits of technology

This approach enables efficient encoding and decoding of mesh data, providing quality 3D services and general-purpose applications like autonomous driving, while optimizing coding efficiency and handling hardware constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025004568_09102025_PF_FP_ABST
    Figure KR2025004568_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A decoding method according to embodiments may comprise: a base mesh processing step for restoring a base mesh from a base mesh bitstream included in a bitstream; a packed video data processing step for restoring displacement information and attribute information by separating same from a packed video data bitstream included in the bitstream on the basis of signaling information; and a restoration step for restoring a mesh on the basis of the base mesh and the displacement information.
Need to check novelty before this filing date? Find Prior Art

Description

Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method

[0001] The embodiments provide a method for providing 3D content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.

[0002] Among 3D content, point cloud data and mesh data are collections of points in 3D space. However, the sheer number of points in 3D space makes it difficult to generate point cloud or mesh data.

[0003] That is, there is a problem that a lot of processing is required to transmit and receive 3D data with a large amount of points, such as point cloud data or mesh data.

[0004] The technical problem according to the embodiments is to provide a device and method for efficiently transmitting and receiving mesh data in order to solve the problems described above.

[0005] The technical problem according to the embodiments is to provide a device and method for resolving latency and encoding / decoding complexity of mesh data.

[0006] A technical problem according to embodiments is to provide an apparatus and method for efficiently performing encoding and decoding of packed video data in which displacement information and a texture map are packed into one frame.

[0007] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.

[0008] To achieve the above-described purpose and other advantages, a decoding method according to embodiments may include a step of receiving a bitstream including mesh data and a step of decoding the mesh data.

[0009] According to embodiments, the step of decoding the mesh data may include a base mesh processing step of restoring a base mesh from a base mesh bitstream included in the bitstream; a packed video data processing step of separating and restoring displacement information and attribute information from a packed video data bitstream included in the bitstream based on signaling information; and a restoration step of restoring a mesh based on the base mesh and the displacement information.

[0010] According to embodiments, the packed video data processing step may include: scalable decoding packed video data from the packed video data bitstream based on layers including a base layer and one or more enhancement layers; and separating and restoring displacement information and attribute information from the packed video data of the decoded target packed video frame based on the signaling information.

[0011] According to embodiments, the signaling information includes packing area information for a packed area for each layer, and the packing area information may include information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

[0012] According to embodiments, the signaling information includes packing area information for a packed area for the base layer, and the packing area information may include upper left position information of the packed area, width information of the packed area, and height information of the packed area.

[0013] According to embodiments, the packing area information for the enhancement layer is derived from the packing area information of the base layer based on scale factor information, wherein the scale factor information may be included in the signaling information or may be derived directly from a decoding device.

[0014] According to embodiments, a decoding method includes a memory and at least one processor connected to the memory, wherein the at least one processor is configured to receive a bitstream including mesh data and decode the mesh data.

[0015] According to embodiments, the at least one processor may include a base mesh processing unit that restores a base mesh from a base mesh bitstream included in the bitstream; a packed video data processing unit that separates and restores displacement information and attribute information from a packed video data bitstream included in the bitstream based on signaling information; and a mesh restoration unit that restores a mesh based on the base mesh and the displacement information.

[0016] According to embodiments, the packed video data processing unit may scalably decode packed video data from the packed video data bitstream based on a layer including a base layer and one or more enhancement layers, and may separate and restore displacement information and attribute information from the packed video data of the decoded target packed video frame based on the signaling information.

[0017] According to embodiments, the signaling information includes packing area information for a packed area for each layer, and the packing area information may include information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

[0018] According to embodiments, the signaling information includes packing area information for a packed area for the base layer, and the packing area information may include upper left position information of the packed area, width information of the packed area, and height information of the packed area.

[0019] According to embodiments, the packing area information for the enhancement layer is derived from the packing area information of the base layer based on scale factor information, wherein the scale factor information may be included in the signaling information or may be derived directly from a decoding device.

[0020] According to embodiments, the encoding step may include encoding mesh data, and transmitting a bitstream including the encoded mesh data.

[0021] According to embodiments, a computer-readable storage medium can store a bitstream generated by the encoding method described above.

[0022] According to embodiments, a transmission method may include the steps of obtaining a bitstream for image information, wherein the bitstream is generated based on the steps of encoding mesh data and transmitting a bitstream including the encoded mesh data; and the step of transmitting data including the bitstream.

[0023] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide a quality 3D service.

[0024] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can achieve various video codec methods.

[0025] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide general-purpose 3D content such as autonomous driving services.

[0026] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments pack a texture map and displacement information into one frame and then encode it using a video codec, thereby combining a video codec used for displacement information compression and a video codec used for texture map compression into one, thereby obtaining the effect of increasing coding efficiency.

[0027] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can efficiently provide a scalable coding service by minimizing synchronization problems between data by packing geometry data and texture data into one frame at a time.

[0028] The mesh data receiving method and the mesh data receiving device according to the embodiments can effectively decode and render scalable coded data in a manner such as in whole or selectively, depending on hardware constraints such as resources and displays of the receiving device and / or the intention and profile definition of the content creator, user, or receiver itself.

[0029] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.

[0030] FIG. 1 illustrates a system for providing dynamic mesh content according to embodiments.

[0031] Figure 2 illustrates a V-MESH compression method according to embodiments.

[0032] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.

[0033] Figure 4 illustrates a mid-edge subdivision method according to embodiments.

[0034] Figure 5 illustrates a displacement generation process according to embodiments.

[0035] Figure 6 illustrates an intra-frame encoding process of V-MESH data according to embodiments.

[0036] Figure 7 illustrates an inter-frame encoding process of V-MESH data according to embodiments.

[0037] Figure 8 illustrates a lifting conversion process for displacement according to embodiments.

[0038] Figure 9 illustrates a process of packing transformation coefficients into a 2D image according to embodiments.

[0039] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.

[0040] Figure 11 illustrates an intra-frame decoding process of V-MESH data according to embodiments.

[0041] Figure 12 shows an inter-frame decoding processor of V-MESH data.

[0042] Fig. 13 is a drawing showing an example of a transmitting device according to embodiments.

[0043] Fig. 14 is a drawing showing an example of a receiving device according to embodiments.

[0044] FIG. 15 is a diagram showing an example of a dynamic mesh bitstream structure encoded and transmitted in a transmitting device of the present disclosure.

[0045] FIG. 16 is a diagram showing an example of a syntax structure of a V3C unit payload (V3C_unit_payload) according to embodiments.

[0046] FIG. 17 is a diagram showing an example of codec group profile components according to embodiments.

[0047] FIG. 18a and FIG. 18b are diagrams showing an example of a syntax structure of packing information (packing_information()) according to embodiments.

[0048] FIG. 19 is a diagram showing another example of the syntax structure of packing information (packing_information()) according to embodiments.

[0049] FIG. 20 is a diagram showing another example of the syntax structure of packing information (packing_information()) according to embodiments.

[0050] FIG. 21 is a diagram showing an example of packing information of a texture map and a displacement vector when there are two layers according to embodiments.

[0051] FIG. 22a and FIG. 22b are diagrams showing an example of the syntax structure of an SEI message according to embodiments.

[0052] Fig. 23 is a block diagram showing another example of a receiving device according to embodiments.

[0053] Fig. 24 is a block diagram showing another example of a receiving device according to embodiments.

[0054] Figure 25 is a flowchart showing an example of a transmission method according to embodiments.

[0055] Fig. 26 is a flowchart showing an example of a receiving method according to embodiments.

[0056] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.

[0057] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.

[0058] With the recent development of 3D data modeling and rendering technology, research on creating and processing 3D data is being conducted in various fields such as Virtual Reality (VR), Augmented Reality (AR), autonomous driving, Computer-Aided Design (CAD) / Computer-Aided Manufacturing (CAM), and Geographic Information Systems (GIS). 3D data can be represented as point clouds, meshes, etc., depending on the format of expression. Among these, a mesh is composed of geometric information expressing the coordinate values ​​of each vertex (or point), connection information indicating the connection relationship between vertices, a texture map expressing the color information of the mesh surface as 2D image data, and texture coordinates indicating mapping information between the surface of the mesh and the texture map. In the present disclosure, a mesh is defined as a dynamic mesh if one or more of the elements constituting the mesh change over time, and a static mesh if they do not change. In other words, dynamic mesh data may refer to mesh data having an object or movement.

[0059] Because dynamic mesh data has a large amount of data for elements that constitute the mesh compared to two-dimensional image data, technologies have been developed to efficiently compress this large amount of mesh data to store and transmit it.

[0060] FIG. 1 illustrates a system for providing dynamic mesh content according to embodiments.

[0061] The system of FIG. 1 includes a transmitting device (100) and a receiving device (110) according to embodiments. The transmitting device (100) may include a mesh video acquisition unit (101), a mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The receiving device (110) may include a receiving unit (111), a file / segment decapsulator (112), a mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the mesh data transmitting device according to embodiments may be interpreted as a term referring to a 3D data transmitting device or transmitting device (100), or a mesh video encoder (hereinafter, referred to as an encoder) (102). The mesh data receiving device according to the embodiments may be interpreted as a term referring to a 3D data receiving device or receiving device (110), or a mesh video decoder (hereinafter, decoder) (113).

[0062] The system of FIG. 1 can perform video-based dynamic mesh compression and decompression.

[0063] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.

[0064] First, the method of compressing dynamic mesh data starts from the V-PCC (Video-based point cloud compression) standard technology for point cloud data. Point cloud data is data that has color information at the coordinates (X, Y, Z) of a vertex (or point). In the present disclosure, the coordinates (i.e., position information) of a vertex are referred to as geometry information, the color information of a vertex is referred to as attribute information, and the geometry information and attribute information are referred to as vertex information or point cloud data. The vertex information to which connectivity information between vertices is added is referred to as mesh data. When creating content, it can be created in the form of mesh data from the beginning. Alternatively, it can be used by converting it into mesh data by adding connectivity information to point cloud data.

[0065] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.

[0066] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.

[0067] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.

[0068] A mesh compression system may include a transmitting device (100) and a receiving device (110) as shown in FIG. 1. The transmitting device (100) may encode mesh video to output a bitstream, and transmit the bitstream to the receiving device (110) in the form of a file or streaming (streaming segment) via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0069] In the above transmitting device (100), the encoder may be called a mesh video / video / picture / frame encoding device, and in the receiving device (110), the decoder may be called a mesh video / video / picture / frame decoding device. The transmitter may be included in a mesh video encoder. The receiver may be included in a mesh video decoder. The renderer (114) may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device (100) and the receiving device (110) may further include separate internal or external modules / units / components for a feedback process.

[0070] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by vertices in 3D space and connection information that describes how the vertices are connected. It can also contain vertex attributes such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included in the attributes of the mesh. The mapping can be described as a set of parameter coordinates, generally called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement. Here, displacement can be used interchangeably with displacement, displacement information, or displacement vectors (i.e., displacement vectors).

[0071] The mesh video acquisition unit (101) may include processing 3D object data acquired through a camera, etc. into a mesh data type having the attributes described above through a series of processes and generating a video composed of such mesh data. The mesh video may have attributes of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video having attributes and connection information that change over time in this way may be expressed as a dynamic mesh video.

[0072] A mesh video encoder (102) can encode an input mesh video into one or more video streams. One video can include multiple frames, and one frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and a mesh video can be used interchangeably with a mesh image / frame / picture. The mesh video encoder (102) can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. The mesh video encoder (102) can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.

[0073] The file / segment encapsulator (103) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from a metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder (102) or may be configured as a separate component / module. The file / segment encapsulator (103) can encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The file / segment encapsulator (103) may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. Depending on the embodiment, the file / segment encapsulator (103) may encapsulate the mesh video related metadata itself into a file.

[0074] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit (104) or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process it for transmission.

[0075] The transmission unit (104) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (111) of the reception device (110) via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (104) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (111) can extract the bitstream and transmit it to a decoding device.

[0076] The receiving unit (111) can receive mesh video data transmitted by a mesh data transmission device. Depending on the channel through which it is transmitted, the receiving unit (111) can receive mesh video data through a broadcast network, through a broadband, or through a digital storage medium.

[0077] The receiving processing unit can perform processing according to the transmission protocol on the received mesh video data. The receiving processing unit can be included in the receiving unit (111) or can be configured as a separate component / module. In order to correspond to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit can transfer the acquired mesh video data to the file / segment decapsulator (112) and transfer the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.

[0078] The file / segment decapsulator (112) can decapsulate mesh video data in the form of a file received from a receiving processing unit. The file / segment decapsulator (112) can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to the mesh video decoder (113), and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to the metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder (113) or may be configured as a separate component / module. The mesh video-related metadata obtained by the file / segment decapsulator (112) may be in the form of a box or track within a file format. The file / segment decapsulator (112) may receive metadata required for decapsulation from the metadata processing unit, if necessary. The mesh video related metadata may be passed to the mesh video decoder (113) and used in the mesh video decoding procedure, or may be passed to the renderer (114) and used in the mesh video rendering procedure.

[0079] The mesh video decoder (113) can receive a bitstream and perform a reverse operation corresponding to the operation of the mesh video encoder (102) to decode the video / image. The decoded mesh video / image can be displayed through the display unit of the renderer (114). The user can view all or part of the rendered result through a VR / AR display or a general display.

[0080] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.

[0081] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.

[0082] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of ​​the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.

[0083] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.

[0084] This document relates to embodiments of dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.

[0085] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.

[0086] In this document, picture / frame can generally mean a unit representing one video of a specific time period.

[0087] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.

[0088] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.

[0089] As described above, the encoding process of Fig. 1 is as follows.

[0090] That is, the video-based dynamic mesh compression (V-Mesh) compression method can provide a method of compressing dynamic mesh video data based on 2D video codecs such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding). The V-Mesh compression process receives the following data as input and performs compression.

[0091] Input mesh: Contains the 3D coordinates of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.

[0092] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the attributes of the mesh (color, normal, displacement, etc.), and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video. The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).

[0093] Material Library File: Contains information about the material attributes used in a mesh, and in particular, information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.

[0094] In the V-Mesh compression method, the following data and information can be generated through the compression process.

[0095] Base mesh: The input mesh is simplified (decimated) through a pre-processing process, thereby expressing the objects of the input mesh using the minimum number of vertices determined by the user's standards.

[0096] Displacement: Displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.

[0097] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.

[0098] Referring to FIGS. 2 to 7, a method for encoding mesh position information (or vertex position information) is described, and referring to FIGS. 6 to 10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.

[0099] Figure 2 illustrates a V-MESH compression method according to embodiments.

[0100] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The mesh video encoder (102) of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. In addition, the transmitting device of Fig. 1 may be broadly referred to as an encoder, and the mesh video encoder (102) of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing process (Pre-processing, 200) and an encoding process (Encoding, 201) as in Fig. 2. The pre-processor (200) of Fig. 2 may be located in front of the encoder (201) of Fig. 2. The pre-processor (200) and the encoder (201) of Fig. 2 may be referred to as a single encoder.

[0101] The pre-processor (200) can receive a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). The pre-processor (200) can generate a base mesh (m(i)) and / or a displacement (or displacement) (d(i)) through pre-processing. The pre-processor (200) can receive feedback information from the encoder (201) and generate the base mesh and / or the displacement based on the feedback information.

[0102] The encoder (201) can receive a base mesh (m(i)), a displacement (d(i)), a static of a dynamic mesh (M(i)), and / or an attribute map (A(i)). In the present disclosure, at least one of the base mesh (m(i)), the displacement (d(i)), the static of a dynamic mesh (M(i)), and / or the attribute map (A(i)) can be referred to as mesh-related data. The encoder (201) can encode the mesh-related data to generate a compressed bitstream.

[0103] Figure 3 illustrates a pre-processing process of V-MESH compression according to embodiments.

[0104] Fig. 3 illustrates the configuration and operation of the preprocessor of Fig. 2. In Fig. 3, the input mesh may include a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). In addition, the input mesh may include three-dimensional coordinates of vertices constituting the mesh, normal information of each vertex, mapping information for mapping the mesh surface to a 2D plane, connection information between vertices constituting the surface, etc.

[0105] Fig. 3 shows a process of performing pre-processing on an input mesh. The pre-processing process (200) may largely include four steps: 1) GoF (Group of Frame) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). According to embodiments, GoF generation may be referred to as a GoF generation process or a GoF generation unit, mesh simplification may be referred to as a mesh simplification process or a mesh simplification unit, UV parameterization may be referred to as a UV parameterization process or a UV parameterization unit, and the fitting subdivision surface may be referred to as a fitting subdivision surface process or a fitting subdivision surface unit. The pre-processor (200) can generate displacement and / or base meshes from the received input mesh and transmit them to the encoder (201). The pre-processor (200) can transmit GoF information associated with GoF generation to the encoder (201).

[0106] Below, each step of Fig. 3 is described.

[0107] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, the number of texture coordinates, the vertex connection information, and the texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. That is, if only the vertex coordinate values ​​are different between the current input mesh and the reference input mesh, the encoder (201) can perform inter frame encoding. Otherwise, intra frame encoding is performed for the corresponding frame.

[0108] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.

[0109] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.

[0110] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.

[0111] Fitting subdivision surface (300): This is a process of performing subdivision on a decimated mesh (i.e., a simplified mesh having texture coordinates). The displacement and base mesh generated through this process are output to the encoder (201). A user-defined method, such as a mid-edge method, may be applied as the subdivision method. A fitting process is performed so that the input mesh and the mesh on which the subdivision has been performed become similar to each other. In the present disclosure, the mesh on which the fitting process has been performed is referred to as a fitted subdivision mesh (or fitted subdivision mesh).

[0112] Figure 4 illustrates a mid-edge subdivision method according to embodiments.

[0113] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to generate a sub-mesh. A sub-mesh can be generated by creating a new vertex in the middle of the edge between the vertices. Then, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, thereby generating a fitted sub-division mesh.

[0114] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents the position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values ​​can be converted to (normal, tangential, bi-tangential) coordinate values ​​of the local coordinate system.

[0115] Fig. 5 illustrates a displacement generation process according to embodiments. The displacement generation process of Fig. 5 may be performed in a pre-processor (200) or in an encoder (201).

[0116] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.

[0117] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may perform subdivision on a reconstructed base mesh to generate a subdivided reconstructed base mesh. Here, the restoration of the base mesh may be performed in the pre-processor (200) or in the encoder (201). The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and may convert a coordinate system of the mesh into a local coordinate system based on the fitted subdivision mesh and the subdivided reconstructed base mesh. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.

[0118] The mesh data transmission method and device according to the embodiments can encode mesh data as follows. Mesh data is a term including point cloud data. Point cloud data (which may be referred to as point cloud for short) according to the embodiments can refer to data including vertex coordinates (or geometry information) and color information (or attribute information). In addition, geometry images, attribute images, occupancy maps, and additional information (or patch information) generated through patch generation and packing based on vertex coordinates and color information are also referred to as point cloud data. Therefore, point cloud data including connection information can be referred to as mesh data. In this document, point cloud and mesh data can be used interchangeably.

[0119] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).

[0120] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.

[0121] Fig. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments. Each component for the intra-frame encoding process of Fig. 6 corresponds to hardware, software, a processor, and / or a combination thereof.

[0122] The encoding process of FIG. 6 details the encoding of the mesh video encoder (102) of FIG. 1. That is, it shows the configuration of the mesh video encoder (102) when the encoding of FIG. 1 is an intra-frame method. The encoder of FIG. 6 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of FIG. 6 may correspond to the pre-processor (200) and encoder (201) of FIG. 3.

[0123] The preprocessor (200) can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivision mesh.

[0124] The quantizer (411) of the encoder (201) can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder (412) can encode the static mesh (i.e., the quantized base mesh) and generate a bitstream (i.e., a compressed base mesh bitstream) including the encoded base mesh. The static mesh decoder (413) can decode the encoded static mesh (i.e., the encoded base mesh). The inverse quantizer (414) can inversely quantize the quantized static mesh (i.e., the base mesh) to output a reconstructed (or restored) base mesh. The displacement calculation unit (415) can generate displacements (or displacements) based on the reconstructed static mesh (i.e., the base mesh) and the fitted subdivided mesh. According to embodiments, the displacement calculation unit (415) calculates displacement, which is the position difference between each vertex of the subdivided base mesh and the fitted subdivided mesh after subdividing (or refining) the restored base mesh. In other words, the displacement is a displacement vector, which is the position difference between the vertices of the two meshes so that the fitted subdivided (or refining) mesh becomes similar to the original mesh. The forward linear lifting unit (416) can perform lifting transformation on the input displacement to generate lifting coefficients (or transform coefficients). The quantizer (417) can quantize the lifting coefficients. The image packing unit (418) can pack an image based on the quantized lifting coefficients. The video encoder (419) can encode the packed image. That is, the quantized lifting coefficients are packed into one frame as a 2D image by the image packing unit (418), compressed through the video encoder (419), and output as a displacement bitstream (i.e., compressed displacement bitstream).

[0125] A video decoder (420) decodes a compressed displacement bitstream. An image unpacking unit (421) can perform unpacking on the decoded displacement frame to output quantized lifting coefficients. A dequantizer (422) can dequantize the quantized lifting coefficients. An inverse linear lifting unit (423) applies inverse lifting to the inverse quantized lifting coefficients to generate restored displacement. A mesh restoration unit (424) reconstructs and deforms a mesh using the restored displacement output from the inverse linear lifting unit (423) and the restored base mesh (or subdivided restored base mesh) output from the inverse quantization unit (414). The present disclosure refers to the reconstructed and deformed mesh as a restored deformed mesh.

[0126] The attribute transfer (425) receives an input mesh and / or an input attribute map, and regenerates an attribute map based on the restored deformed mesh. The attribute map refers to a texture map corresponding to attribute information among mesh data components, and in the present disclosure, the attribute map and the texture map may be used interchangeably. The push-pull padding (426) may pad data in the attribute map based on the push-pull method. The color space conversion unit (427) may convert the space of the color component of the attribute map. For example, the attribute map may be converted from an RGB color space to a YUV color space. The video encoder (428) may encode the attribute map and output it as a compressed attribute bitstream.

[0127] A multiplexer (430) can generate a compressed bitstream by multiplexing a compressed base mesh bitstream, a compressed displacement bitstream, and a compressed attribute bitstream.

[0128] In Fig. 6, the displacement calculation unit (415) may be included in the pre-processor (200). In addition, at least one of the quantizer (411), the static mesh encoder (412), the static mesh decoder (413), and the inverse quantizer (414) may be included in the pre-processor (200).

[0129] As described in FIG. 6, the intra-frame encoding method includes base mesh encoding (also called static mesh encoding). That is, when performing intra-frame encoding on the current input mesh frame, the base mesh generated in the pre-processing process of the pre-processor (200) can be encoded using a static mesh compression technology in a static mesh encoder (412) after undergoing a quantization process in a quantizer (411). In the V-Mesh compression method, for example, Draco technology is applied to base mesh encoding, and vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh become compression targets.

[0130] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.

[0131] Fig. 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments. Each component for the inter-frame encoding process of Fig. 7 corresponds to hardware, software, a processor, and / or a combination thereof.

[0132] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of Fig. 7 may correspond to the pre-processor (200) and encoder (201) of Fig. 3.

[0133] For a description of the components corresponding to the encoding operation of FIG. 6 among the encoding operations of FIG. 7, refer to the description of FIG. 6. That is, the operation of the quantizer (511), displacement calculation unit (515), wavelet transformer (516), quantizer (517), image packing unit (518), video encoder (519), video decoder (520), image unpacking unit (521), inverse quantizer (522), inverse wavelet transformer (523), mesh restoration unit (524), attribute transfer (525), push-pull padding (526), ​​color space conversion unit (527), video encoder (528), and multiplexer (530) of FIG. 7 is similar to that of the quantizer (411), static mesh encoder (412), static mesh decoder (413), inverse quantizer (414), displacement calculation unit (415), forward linear lifting unit (416), quantizer (417), image Since the operations described in the packing unit (418), video encoder (419), video decoder (420), image unpacking unit (421), inverse quantizer (422), inverse linear lifting unit (423), mesh restoration unit (424), attribute transfer (425), push-pull padding (426), color space conversion unit (427), video encoder (428), and multiplexer (430) are the same or similar, a detailed description thereof is omitted in FIG. 7 to avoid redundant description.

[0134] In Fig. 7, for inter-frame based encoding, the motion encoder (512) can obtain a motion vector between the two base meshes based on the restored quantized reference base mesh and the quantized current base mesh, and then encode the motion vector to output a compressed motion bitstream. The motion encoder (512) can be referred to as a motion vector encoder. The base mesh restoration unit (513) can restore the base mesh based on the restored quantized reference base mesh and the encoded motion vector. The restored base mesh is dequantized in the dequantizer (514) and then output to the displacement calculation unit (515).

[0135] In Fig. 7, the displacement calculation unit (515) may be included in the pre-processor (200). In addition, at least one of the quantizer (511), the motion encoder (512), the base mesh restoration unit (513), and the inverse quantizer (514) may be included in the pre-processor (200).

[0136] As described in Fig. 7, the inter-frame encoding method may include motion field encoding (also called motion vector encoding). Inter-frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter-frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field (also called motion vector), may be calculated and encoded to encode this information. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may also be encoded as a value. Alternatively, the predicted motion field can be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the motion field value of the current vertex, can be encoded. This residual motion field value can be encoded using entropy coding.The process of encoding displacement and attribute maps, excluding the motion field encoding process of inter frame encoding, is the same as the structure of the intra frame encoding method except for the base mesh encoding.

[0137] Figure 8 illustrates a lifting conversion process for displacement according to embodiments.

[0138] Figure 9 illustrates a process of packing transformation coefficients (or lifting coefficients) according to embodiments into a 2D image.

[0139] Figures 8 and 9 illustrate the process of transforming displacement and packing transform coefficients of the encoding process of Figures 6 and 7, respectively.

[0140] The encoding method according to the embodiments includes displacement encoding.

[0141] After base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated (see 415 in FIG. 6 or 515 in FIG. 7). For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information (see 416 in FIG. 6 or 516 in FIG. 7).

[0142] FIG. 8 shows a process of transforming displacement information using a lifting transform in the forward linear lifting unit (416) of FIG. 6 or the wavelet transformer (516) of FIG. 7. For example, a linear wavelet-based lifting transform may be performed. The transform coefficients generated through the transform process are quantized in a quantizer (417 or 517) and then packed into a 2D image through an image packing unit (418 or 518) as in FIG. 9. The transform coefficients are configured as one block for every 256 (= 16×16) units, and each block can be packed in a z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. Transform coefficients can be packed by aligning them with Morton codes within a single block. The packed images generate displacement videos for each GoF unit, and these displacement videos can be encoded using a conventional video compression codec in a video encoder (419 or 519).

[0143] Referring to FIG. 8, a base mesh (original) may include vertices and edges for LoD (Level of Detail) 0. A first subdivision mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) edges of the base mesh. The first subdivision mesh includes vertices for LoD 0 and vertices for LoD 1. LoD 1 includes the subdivided vertices and the vertices of the base mesh (LoD 0). The first subdivision mesh may be further divided (or subdivided) to generate a second subdivision mesh. The second subdivision mesh includes LoD 2. LoD 2 includes base mesh vertices (LoD 0), LoD 1 including vertices further divided (or subdivided) from LoD 0, and vertices further divided (or subdivided) from LoD 1. LoD is a level of detail that indicates the degree of detail of mesh data content. As the level index increases, the distance between vertices becomes closer and the level of detail increases. In other words, the smaller the LoD value, the lower the detail of the mesh data content, and the larger the LoD value, the higher the detail of the mesh data content. LoD N contains the vertices included in the previous LoDN-1 as is. When a mesh (or vertex) is further divided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1, v2, and the subdivided vertex v. Instead of directly encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated and the mesh can be encoded using the residual value to reduce the bitstream size. The prediction process refers to the operation of predicting the current vertex v using the previous vertices v1 and v2. Since adjacent subdivision meshes have similar data, this property can be utilized for efficient encoding.Current vertex position information is predicted as a residual of previous vertex position information, and the previous vertex position information is updated through the residual. In the present disclosure, vertex, apex, and point may be used with the same meaning. In addition, LoDs may be defined during the subdivision process of the base mesh. According to embodiments, the subdivision process of the base mesh may be performed in the pre-processor (200) or in a separate component / module.

[0144] Referring to FIG. 9, a vertex has a transform coefficient (also called a lifting coefficient) generated through a lifting transformation. The transform coefficient of a vertex related to a lifting transformation can be packed into an image by an image packing unit (418 or 518) and then encoded by a video encoder (419 or 519).

[0145] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.

[0146] According to the embodiments, FIG. 10 shows the detailed operation of the attribute transfer (425 or 525) of the encoding of FIG. 6, FIG. 7, etc.

[0147] Encoding according to embodiments includes attribute map encoding. According to embodiments, attribute map encoding may be performed in the video encoder (428) of FIG. 6 or the video encoder (528) of FIG. 7.

[0148] According to embodiments, in the present disclosure, the encoder compresses information about the input mesh through base mesh encoding (i.e., intra encoding), motion field encoding (i.e., inter encoding), and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding processes, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has position information of vertices, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of attribute transfer (425 or 525).

[0149] According to embodiments, attribute transfer (425 or 525) first checks whether each point P(u, v) of a 2D texture domain belongs to a texture triangle of a reconstructed deformed mesh, and if it exists in a texture triangle T, the barycentric coordinate of P(u, v) according to the triangle T ( , , ) is calculated. And the 3D vertex positions of triangle T and ( , , ) is used to compute the 3D coordinates M(x, y, z) of P(u, v). Find the vertex coordinates M'(x', y', z') and the triangle T' containing this vertex that corresponds to the most similar position to the calculated M(x, y, z) in the input mesh domain. Then, the center of mass coordinates of M'(x', y', z') in this triangle T' ( ', ', ') is calculated. The texture coordinates corresponding to the three vertices of Triangle T' and ( ', ', ') is used to calculate the texture coordinates (u', v'), and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm, such as the push-pull algorithm of push-pull padding (426 or 526).

[0150] The new attribute map generated through attribute transfer (425 or 525) is grouped into GoF units to form an attribute map video, which is compressed using the video codec of the video encoder (428 or 528).

[0151] Referring to Figure 10, the reference relationship between the input mesh, the input attribute map, the reconstructed deformed mesh, and the regenerated attribute map can be seen.

[0152] The decoding process of Fig. 1 can perform the reverse process of the corresponding process of the encoding process of Fig. 1. The specific decoding process is as follows.

[0153] FIG. 11 illustrates an intra-frame decoding (or intra-decoding) process of V-Mesh technology according to embodiments.

[0154] Fig. 11 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1. In addition, Fig. 11 can restore mesh data by performing the reverse process of the intra-frame encoding process of Fig. 6. Each component for the intra-frame decoding process of Fig. 11 corresponds to hardware, software, and / or a combination thereof.

[0155] First, the bitstream (i.e., compressed bitstream) received and input to the demultiplexer (611) of the intra frame decoding unit (610) can be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information of the mesh, such as V-PCC / V3C. The term V-PCC (Video-based Point Cloud Compression) used in this document can be used with the same meaning as V3C (Visual Volumetric Video-based Coding), and the two terms can be used interchangeably. Therefore, the term V-PCC in this document can be interpreted as the term V3C.

[0156] According to embodiments, the mesh sub-stream may be input to a static mesh decoder (612) and decoded, the displacement sub-stream may be input to a video decoder (613) and decoded, and the attribute map sub-stream may be input to a video decoder (617) and decoded.

[0157] According to embodiments, the mesh sub-stream is decoded through a decoder (612) of a static mesh codec used in encoding, such as Google Draco, and as a result, a reconstructed quantized base mesh, for example, connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be reconstructed.

[0158] According to embodiments, the displacement sub-stream is decoded into displacement video through a decoder (613) of a video compression codec used in encoding, and is restored as displacement information for each vertex (i.e., Recon. displacements) through an image unpacking process of an image unpacking unit (614), an inverse quantization process of an inverse quantizer (615), and an inverse transform process of an inverse linear lifting unit (616).

[0159] According to embodiments, the base mesh restored by the static mesh decoder (612) is inverse quantized by the inverse quantizer (620) and then output to the mesh restoration unit (630). The mesh restoration unit (630) reconstructs and restores the deformed mesh (i.e., decoded mesh) through the restored displacement output from the inverse linear lifting unit (616) and the restored base mesh output from the inverse quantizer (620). That is, the inverse quantized restored base mesh is combined with the restored displacement information to generate the final decoded mesh. In the present disclosure, the final decoded mesh is referred to as a reconstructed deformed mesh.

[0160] According to embodiments, an attribute map sub-stream is decoded through a decoder (617) corresponding to a video compression codec used in encoding, and then restored to a final attribute map (i.e., decoded attribute map) through a color conversion unit (640) through processes such as color format conversion and color space conversion.

[0161] According to embodiments, the restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.

[0162] Referring to FIG. 11, the received compressed bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. A substream is interpreted as a term referring to a part of a bitstream included in a bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).

[0163] As described above, the decoder of FIG. 11 performs the following intra-frame decoding operations. The static mesh decoder (612) decodes the mesh sub-stream to generate a reconstructed quantized base mesh, and the inverse quantizer (620) applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder (613) decodes the displacement sub-stream, the image unpacking unit (614) unpacks the images of the decoded displacement video, and the inverse quantizer (615) inversely quantizes the quantized images. The inverse linear lifting unit (616) applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit (630) generates a reconstructed deformed mesh based on the reconstructed base mesh and the reconstructed displacement. The video decoder (617) decodes the attribute map sub-stream, and the color conversion unit (640) converts the color format and / or space of the decoded attribute map to generate a decoded attribute map.

[0164] Figure 12 illustrates the inter-frame decoding (or inter-decoding) process of V-Mesh technology.

[0165] Fig. 12 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1. In addition, Fig. 12 can restore mesh data by performing the reverse process of the inter-frame encoding process of Fig. 7. Each component for the inter-frame decoding process of Fig. 12 corresponds to hardware, software, and / or a combination thereof.

[0166] First, the bitstream received and input to the demultiplexer (711) of the intra frame decoding unit (710) can be separated into a motion sub-stream (also called a motion sub-stream or motion vector sub-stream), a displacement sub-stream, an attribute map sub-stream, and a sub-stream including patch information of a mesh such as V3C / V-PCC.

[0167] According to embodiments, a motion sub-stream may be input to a motion decoder (712) and decoded, a displacement sub-stream may be input to a video decoder (713) and decoded, and an attribute map sub-stream may be input to a video decoder (717) and decoded.

[0168] According to embodiments, a motion sub-stream is decoded through entropy decoding and inverse prediction processes in a motion decoder (712) and restored into motion information (or motion vector information). A base mesh restoration unit (718) combines the restored motion information with a reference base mesh that has already been restored and stored to generate a reconstructed quantized base mesh for the current frame. An inverse quantizer (720) applies inverse quantization to the restored quantized base mesh to generate a reconstructed base mesh. A video decoder (713) decodes a displacement sub-stream, an image unpacking unit (714) unpacks an image of the decoded displacement video, and an inverse quantizer (715) inversely quantizes a quantized image. The reverse linear lifting unit (716) applies a lifting transformation in the reverse process of the encoder to generate a restored displacement. The mesh restoration unit (730) generates a reconstructed deformed mesh, i.e., a final decoded mesh, based on the restored base mesh and the restored displacement.

[0169] According to embodiments, the video decoder (717) decodes the attribute map sub-stream in the same manner as intra decoding, and the color conversion unit (740) converts the color format and / or space of the decoded attribute map to generate a decoded attribute map. The decoded mesh and the decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.

[0170] Referring to Fig. 12, the bitstream includes motion information (also called motion vectors), displacement, and an attribute map. Since Fig. 12 performs inter-frame decoding, it further includes a process of decoding inter-frame motion information. The motion information is decoded, and a restored quantized base mesh for the motion information is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operation of Fig. 12, which is identical to that of Fig. 11, refer to the description of Fig. 11.

[0171] Fig. 13 illustrates a mesh data transmission device according to embodiments.

[0172] FIG. 13 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 1, the encoder (preprocessor and encoder) of FIG. 2, FIG. 6, or FIG. 7, and / or a transmitting encoding device corresponding thereto. Each component of FIG. 13 corresponds to hardware, software, a processor, and / or a combination thereof.

[0173] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13. The transmitter of Fig. 13 may perform an intra-frame encoding (or intra-encoding or intra-screen encoding) process and / or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.

[0174] The pre-processor (811) receives the original mesh as input and generates a simplified mesh (decimated mesh) (or base mesh) and a fitted decimated mesh (or subdivision). Simplification can be performed based on the target number of vertices or target number of polygons that constitute the mesh. Parameterization, which generates texture coordinates and texture connection information per vertex, can be performed on the simplified mesh. For example, parameterization is a process of mapping a 3D surface to a texture domain for the decimated mesh. If parameterization is performed using the UVAtlas tool, mapping information is generated that can identify where each vertex of the decimated mesh can be mapped on a 2D image. The mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process. In addition, the work of quantizing the mesh information in floating-point form into fixed-point form can be performed. This result can be output as a base mesh to a motion vector encoder (813) or a static mesh encoder (814) through a switching unit (812). The pre-processor (811) can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The pre-processor (811) can generate a fitted subdivided mesh by adjusting the vertex positions so that the subdivided mesh becomes similar to the original mesh.

[0175] According to embodiments, the base mesh is output to a motion vector encoder (813) via a switching unit (812) when performing inter-encoding for the corresponding mesh frame, and is output to a static mesh encoder (814) via a switching unit (812) when performing intra-encoding for the corresponding mesh frame. The motion vector encoder (813) may be referred to as a motion encoder.

[0176] For example, when performing intra-encoding (or intra-frame encoding) on ​​the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (814). In this case, encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to a multiplexer (823).

[0177] As another example, when performing inter-encoding (or inter-frame encoding) for the corresponding mesh frame, the motion vector encoder (813) can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as input, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (813) can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexer (823).

[0178] The base mesh restoration unit (815) can receive the base mesh encoded by the static mesh encoder (814) or the motion vector encoded by the motion vector encoder (813) and generate a reconstructed base mesh. For example, the base mesh restoration unit (815) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (814) to restore the base mesh. At this time, quantization can be applied before the static mesh decoding, and inverse quantization can be applied after the static mesh decoding. As another example, the base mesh restoration unit (815) can restore the base mesh based on the reconstructed quantized reference base mesh and the motion vector encoded by the motion vector encoder (813). The reconstructed base mesh is output to the displacement calculation unit (816) and the mesh restoration unit (820).

[0179] The displacement calculation unit (816) can perform mesh refinement on the restored base mesh. The displacement calculation unit (816) can calculate a displacement vector, which is a difference value between the vertex positions of the restored base mesh and the fitted subdivision (or refined) mesh generated by the pre-processor (811). At this time, the displacement vector can be calculated as many times as the number of vertices of the refined mesh. The displacement calculation unit (816) can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.

[0180] The displacement vector video generation unit (817) may include a linear lifting unit, a quantizer, and an image packing unit. That is, in the displacement vector video generation unit (817), the linear lifting unit may transform the displacement vector for effective encoding. The transformation may be performed by a lifting transformation, a wavelet transformation, etc., according to embodiments. In addition, quantization may be performed on the transformed displacement vector value, i.e., the transform coefficient, in a quantizer. At this time, a different quantization parameter may be applied to each axis of the transform coefficient, and the quantization parameter may be derived according to an encoder / decoder agreement. The transformed and quantized displacement vector information may be packed into a 2D image in the image packing unit. The displacement vector video generation unit (817) may generate a displacement vector video by bundling packed 2D images for each frame, and the displacement vector video may be generated for each GoF (Group of Frame) unit of the input mesh.

[0181] The displacement vector video encoder (818) can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer (823).

[0182] The displacement vector restoration unit (819) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (819) performs decoding on an encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (820). The mesh restoration unit (820) restores a deformed mesh based on the base mesh restored by the base mesh restoration unit (815) and the displacement vector restored by the displacement vector restoration unit (819). The restored mesh (or referred to as a restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.

[0183] The texture map video generation unit (821) can regenerate a texture map based on the texture map (or attribute map) of the original mesh and the restored deformed mesh output from the mesh restoration unit (820). According to embodiments, the texture map video generation unit (821) can assign color information per vertex of the texture map of the original mesh to the texture coordinates of the restored deformed mesh. According to embodiments, the texture map video generation unit (821) can generate a texture map video by grouping the regenerated texture maps by GoF unit for each frame.

[0184] The generated texture map video can be encoded using a video compression codec of a texture map video encoder (822). The texture map video bitstream generated through encoding is transmitted to a multiplexer (823).

[0185] A multiplexer (823) multiplexes a motion vector bitstream (e.g., in case of inter encoding), a base mesh bitstream (e.g., in case of intra encoding), a displacement vector bitstream, and a texture map bitstream into a single bitstream. The single bitstream can be transmitted to a receiver via a transmitter (824). Alternatively, the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream can be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via the transmitter (824).

[0186] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map). A transmitting device according to inter-encoding can generate a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map). The texture map obtained from the data input unit is generated and encoded based on the restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the divided (or subdivided or subdivided) mesh. More specifically, the displacement is the difference in position between the fitted subdivided mesh and the subdivided restored base mesh, i.e., the difference in vertex positions between the two meshes. In addition, the base mesh is generated by simplifying and encoding the original mesh through pre-processing. Motion is generated as motion vectors for the mesh of the current frame based on the reference base mesh of the previous frame.

[0187] Fig. 14 illustrates a mesh data receiving device according to embodiments.

[0188] Fig. 14 corresponds to the receiving device (110) or mesh video decoder (113) of Fig. 1, the decoder of Fig. 11 or Fig. 12, and / or the receiving decoding device corresponding thereto. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.

[0189] The bitstream of the mesh data received by the receiver (910) is demultiplexed into a compressed motion vector bitstream (e.g., inter decoding) or a base mesh bitstream (e.g., intra decoding), a displacement vector bitstream, and a texture map bitstream after file / segment decapsulation in the demultiplexer (911). For example, if the current mesh has inter-screen encoding (i.e., inter encoding) applied, the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder (913) via the switching unit (912). As another example, if the current mesh has intra-screen encoding (i.e., intra encoding) applied, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (914) via the switching unit (912). Here, the motion vector decoder (913) may be referred to as a motion decoder.

[0190] According to embodiments, if the current mesh has inter-screen encoding applied according to frame header information, the motion vector decoder (913) can perform decoding on the motion vector bitstream. According to embodiments, the motion vector decoder (913) can reconstruct the final motion vector by adding the previously decoded motion vector as a predictor to the residual motion vector decoded from the bitstream.

[0191] According to embodiments, if the current mesh has been encoded within the screen according to the frame header information, the static mesh decoder (914) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.

[0192] According to embodiments, the base mesh restoration unit (915) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh has intra-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by performing inverse quantization on the base mesh decoded through the static mesh decoder (914).

[0193] According to embodiments, the displacement vector video decoder (917) can decode the displacement vector bitstream as a video bitstream using a video codec.

[0194] According to embodiments, the displacement vector restoration unit (918) extracts displacement vector transform coefficients from the decoded displacement vector video, and restores the displacement vector by applying inverse quantization and inverse transformation processes to the extracted displacement vector transform coefficients. To this end, the displacement vector restoration unit (918) may include an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. If the restored displacement vector is a value in a local coordinate system, a process of inversely transforming it into a Cartesian coordinate system may be performed.

[0195] The mesh restoration unit (916) can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information can be generated. At this time, the mesh restoration unit (916) can generate a final restored mesh (or a restored deformed mesh) by combining the subdivided restored base mesh with the restored displacement vector.

[0196] According to embodiments, the texture map video decoder (919) can decode the texture map bitstream as a video bitstream using a video codec to restore the texture map. The restored texture map has color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of each vertex.

[0197] According to embodiments, the mesh restored by the mesh restoration unit (916) and the texture map restored by the texture map video decoder (919) are shown to the user through a rendering process in the mesh data renderer (920).

[0198] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map. A receiving device according to inter-decoding can receive a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map.

[0199] A mesh data transmission device and method according to embodiments may pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream including the encoded mesh data. A point mesh data reception device and method according to embodiments may receive a bitstream including mesh data and decode the mesh data. The mesh data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The mesh data transmission and reception method / device according to embodiments may also be referred to as 3D data transmission and reception method / device or point cloud data transmission and reception method / device.

[0200] As described above, the V-Mesh method can encode displacement information generated during the encoding process using a video codec-based encoder. In this case, the displacement information can be packed into a 2D image (or frame) and then encoded using a 2D video codec (i.e., a video compression codec). In the present disclosure, the displacement information can be referred to as a displacement vector, a displacement vector transform coefficient, or displacement data.

[0201] Meanwhile, the present disclosure can pack a texture map and a displacement information frame into a single frame and then encode them using a video codec (i.e., a video compression codec). This allows the video codec used for conventional displacement information compression and the video codec used for texture map compression to be combined into one, thereby increasing coding efficiency. For example, assuming that a video codec (e.g., HEVC or SHVC or VVC) is used for displacement information representing geometry information, and another video codec (e.g., HEVC or SHVC or VVC) is used for a texture map, packing the texture map and displacement information into a single frame allows one video codec to be used instead of two. Here, HEVC supports encoding / decoding of a single layer, and SHVC, as an extension of HEVC, supports encoding / decoding of multiple layers (i.e., one base layer and one or more enhancement layers). Additionally, VVC is a successor standard to SHVC and supports encoding / decoding of multiple layers (i.e., multilayers).

[0202] In the present disclosure, data in which a texture map and displacement information are packed into a single frame is referred to as PVD (Packed Video Data) or PVD bitstream. In addition, a frame in which a texture map and a displacement information frame are packed together is referred to as a 'packed video frame', a 'texture map and displacement information frame', or a 'texture map and displacement vector frame'. In addition, in the present disclosure, a texture map may be referred to as texture map data, texture data, attribute data, attribute information, or an attribute map. In addition, displacement information may be referred to as geometry data.

[0203] The present disclosure may apply scalable coding to texture maps and / or displacement information. If scalable coding is applied to both the texture map and displacement information, the resolution of the texture map and the resolution of the displacement information may be the same or different. Furthermore, when packing the texture map and displacement information with scalable coding into a single frame, the resolutions of the texture map and displacement information packed into a single frame may be the same or different.

[0204] The present disclosure proposes signaling (e.g., High Level Syntax, hereinafter referred to as HLS) for supporting spatial resolution scalability of dynamic mesh data based on the PVD type. In addition, the present disclosure proposes signaling (HLS) related to a method of packing displacement data and texture data (or texture map data) into a single frame (packed video) and encoding / decoding them through a scalable codec, among methods for supporting spatial resolution scalability of dynamic meshes. In this way, the present disclosure proposes a scalable coding method for dynamic mesh data, and in particular, a method for defining a scalable coding HLS for an integrated packing method of displacement data (or referred to as geometry data) and texture data (or referred to as texture map data or attribute data).

[0205] That is, a dynamic mesh content receiver may be required to decode / render dynamic mesh content in various configurations depending on resources such as memory, decoder performance, display, or the intention of the producer, and the content producer may encode / decode dynamic mesh content by applying scalable coding (i.e., scalability) according to pre-defined and / or newly defined criteria such as Level of Detail (hereinafter referred to as LoD) taking these factors into consideration.

[0206] At this time, since the dynamic mesh content receiver must decode each data included in the dynamic mesh content to know whether scalable coding is applied or not, a dynamic mesh content receiver with restrictions related to scalable coding may experience unnecessary or abnormal operation and inefficient resource management.

[0207] To solve this, the present disclosure defines a syntax capable of signaling scalable encoded information by integrating and packing (packed video) geometry data and texture data among the data constituting dynamic mesh content into a single frame. In the present disclosure, data of the PVD (Packed Video Data) type among the data constituting the dynamic mesh means data configured by packing a displacement frame and a texture map, which are among the geometry data (or referred to as geometry information), into a single frame, and the present disclosure proposes HLS to support PVD-based scalability functions.

[0208] As described above, the present disclosure can integrate displacement information and attribute information (texture maps) and pack them into a single video frame. For example, by packing displacement information into a video frame and attribute information into the remaining area, two types of data can be jointly packed into a single frame. The decoder of the receiving device according to the embodiments has the effect of unpacking the data packed into a single frame to simultaneously decode displacement information and / or attribute information of related areas.

[0209] Hereinafter, the present disclosure will describe in detail the scalable coded PVD data type, scalable coding codec ID signaling for scalable coding, scalable coding packing information signaling, and packing information signaling for scalable coding independent decoding.

[0210] The V-DMC referred to in this disclosure may also be referred to as V-Mesh, and the terms are used with the same meaning. Dynamic mesh data refers to a type of mesh data, which is a form of point cloud data, and refers to mesh data in which objects and / or people corresponding to objects change over time, i.e., have movement.

[0211] FIG. 15 is a diagram showing an example of a dynamic mesh bitstream structure encoded and transmitted by a transmitting device of the present disclosure. That is, dynamic mesh content can be encoded with a bitstream structure as in FIG. 15 and transmitted to a receiving device. In particular, the present disclosure can use a sample stream data unit used when encoding V3C content of the V3C codec specification (ISO / IEC 23090-5) as in FIG. 15. The present disclosure relates to a method for applying scalable coding to geometry (i.e., displacement data) and texture data packed in one frame when the V3C content format is PVD.

[0212] That is, a bitstream (called a V-DMC bitstream or a dynamic mesh bitstream) transmitted from a transmitting device to a receiving device of the present disclosure may be composed of a sample stream DMC header and a plurality of sample stream DMC units. In the present disclosure, the sample stream DMC header may be referred to as a sample stream header, and the sample stream DMC unit may be referred to as a sample stream data unit.

[0213] At this time, if the sample stream DMC unit follows the V3C codec specification (ISO / IEC 23090-5), the sample stream DMC unit can be composed of V3C sample stream size information and a V3C unit. The V3C unit is again composed of a V3C unit header (V3C_unit_header) and a V3C unit payload (V3C_unit_payload).

[0214] The above V3C sample stream size information specifies the size of the subsequent V3C unit in bytes.

[0215] The above V3C unit header includes type information (vuh_unit_type) indicating the type of data carried by the corresponding V3C unit payload. The above V3C unit payload can carry one of a V3C / V-DMC parameter set (VPC), Atlas Data (AD), Base Mesh Data (BMD), Displacement Data / Geometry Video Data (DD / GVD), Attribute Video Data (AVD), and Packing Video Data (PVD) according to the type information (vuh_unit_type).

[0216] Here, VPS may include parameter set information such as decoder configuration information related to mesh encoding / decoding and sequence header. Atlas data (AD) may include additional information such as 2D mapping or texture mapping for 3D objects. Base mesh data (BMD) is compressed base mesh data for mesh encoding / decoding. In addition, DD / GVD represents displacement data (or displacement information), where DD represents displacement data that is arithmetic coded and GVD represents displacement data that is encoded using a video codec. Attribute video data (AVD) is attribute or texture data (or texture map information) compressed using a video codec. Packed video data (PVD) is packed texture map and displacement information compressed using a video codec.

[0217] Fig. 16 is a diagram showing an example of the syntax structure of a V3C unit payload (V3C_unit_payload) according to embodiments. In Fig. 16, numBytesInV3CPayload indicates the size of the corresponding V3C unit, which can be specified by the V3C sample stream size information.

[0218] The V3C unit payload of FIG. 16 may include one of a V3C parameter set (v3c_parameter_set()), an atlas sub-bitstream (atlas_sub_bitstream()), and a video sub-bitstream (video_sub_bitstream()) depending on the value of the vuh_unit_type field of the corresponding V3C unit header.

[0219] For example, if the vuh_unit_type field indicates a V3C parameter set (V3C_VPS), the V3C unit payload includes a V3C parameter set (v3c_parameter_set()) that contains overall encoding information of the bitstream, and if it indicates atlas data (V3C_AD) or common atlas data (V3C_CAD), it includes an atlas sub-bitstream (atlas_sub_bitstream()) that carries the atlas data or common atlas data. And, if the vuh_unit_type field indicates accumulation video data (V3C_OVD), the V3C unit payload includes an accumulation video sub-bitstream (video_sub_bitstream()) carrying accumulation video data, if it indicates geometry video data (V3C_GVD), it includes a geometry video sub-bitstream (video_sub_bitstream()) carrying geometry video data, and if it indicates attribute video data (V3C_AVD), it includes an attribute video sub-bitstream (video_sub_bitstream()) carrying attribute video data. In addition, if the type information (vuh_unit_type) of the V3C unit header indicates PVD, the V3C unit payload includes a packed video sub-bitstream (video_sub_bitstream()) carrying packed video data. That is, in the present disclosure, packet video data (i.e., texture and displacement data packed into one frame) is transmitted to a receiving device through a V3C unit corresponding to V3C_PVD (vuh_unit_type == V3C_PVD).

[0220] According to embodiments, an atlas sub-bitstream is also referred to as an atlas substream, an occupancy video sub-bitstream is also referred to as an occupancy video substream, a geometry video sub-bitstream is also referred to as a geometry video substream, an attribute video sub-bitstream is also referred to as an attribute video substream, and a packed video sub-bitstream is also referred to as a packed video substream.

[0221] The V3C unit payload according to the embodiments has a NAL unit structure coded in HEVC or VVC (SHVC or multi-layer VVC).

[0222] The following describes a method for supporting spatial resolution scalability for data in which the V3C content format is PVD, i.e., displacement vector frames and texture maps are packed into one frame for video encoding.

[0223] As described above, scalable coded packed video data (i.e., data in which a texture map and a displacement vector frame are packed into one frame is scalably coded) is transmitted through a V3C unit payload of a V3C unit whose vuh_unit_type corresponds to V3C_PVD (vuh_unit_type == V3C_PVD). In addition, the V3C unit payload of the scalable coded packed video data uses a video sub-bitstream format, which has a NAL unit structure coded with HEVC or VVC (SHVC or multi-layer VVC).

[0224] Setting of V3C parameter set to signal scalable V-DMC using packed video

[0225] According to embodiments, SHVC and / or multi-layer VVC may be signaled in the V3C parameter set and / or values ​​may be set in the packing information.

[0226] The following describes examples of signaling SHVC and / or multi-layer VVC in the V3C parameter set.

[0227] According to embodiments, the V3C parameter set includes profile_tier_level(). The present disclosure can signal the codec used for scalable V-DMC through the ptl_profile_codec_group_idc field of profile_tier_level(). In the present disclosure, the field can be referred to as a syntax element.

[0228] FIG. 17 is a table showing examples of codec group profile components defined in profile_tier_level() according to embodiments.

[0229] Figure 17 shows an example of newly assigning the ptl_profile_codec_group_idc value for HEVC Scalable Main10 and VVC Multilayer Main10, which are codecs used in scalable V-DMC.

[0230] That is, in FIG. 17, HEVC Scalable Main10 and VVC Multilayer Main10 are defined as new CodecGroups. In addition, ptl_profile_codec_group_idc of HEVC Scalable Main10 is defined as 5, 4CC code as 'shv1', and ptl_profile_codec_group_idc of VVC Multilayer Main10 is defined as 6, 4CC code as 'mvv1'. This is an example, and the ptl_profile_codec_group_idc value and 4ccc code value defined for HEVC Scalable Main10, and the ptl_profile_codec_group_idc value and 4ccc code value defined for VVC Multilayer Main10 may vary depending on the designer. In addition, it is natural that the varied values ​​also belong to the present disclosure.

[0231] The following describes an example of setting values ​​for packing information within a V3C parameter set.

[0232] According to embodiments, packed video frames may be divided into one or more rectangular regions (rectangular frames), and each region is mapped to exactly one atlas tile. Furthermore, rectangular regions of a packed video frame are not allowed to overlap each other.

[0233] According to embodiments, the present disclosure can determine a value for packing_information() included in vps_packed_video_extension of v3c_parameter_set as follows.

[0234] pin_geometry_present_flag and pin_attribute_present_flag must be set to '1'. This means that the packed video consists of geometry and attributes. In addition, the value of pin_attribute_type_id is set to '0' to indicate that the attribute data consists of a texture.

[0235] That is, a V3C parameter set can contain vps_packed_video_extension(), which can be repeated for the value of vps_atlas_count_minus1, and can contain packing_information() when vps_packed_video_present_flag is true (i.e., there is a packed video in the corresponding atlas). For example, if vps_packed_video_present_flag is true for atlas ID j, then packing information (packing_information(j)) corresponding to atlas ID j is included.

[0236] FIG. 18a and FIG. 18b are diagrams showing an example of a syntax structure of packing information (packing_information()) according to embodiments.

[0237] The packing information in FIG. 18a and FIG. 18b is an example of packing information (packing_information(j)) corresponding to atlas ID j.

[0238] In FIG. 18a and FIG. 18b, pin_codec_id[ j ] represents a mapping index of a codec identifier of a video decoder used to decode a packed video sub-bitstream for atlas ID j.

[0239] If pin_occupancy_present_flag[ j ] is 0, it can indicate that the packed video frames of atlas ID j do not contain regions containing occupancy data, and if it is 1, it can indicate that they contain regions containing occupancy data.

[0240] If pin_geometry_present_flag[ j ] is 0, it can indicate that the packed video frames of atlas ID j do not contain regions containing geometry data, and if it is 1, it can indicate that they contain regions containing geometry data.

[0241] If pin_attribute_present_flag[ j ] is 0, it can indicate that the packed video frames of atlas ID j do not contain regions containing attribute data, and if it is 1, it can indicate that they contain regions containing attribute data.

[0242] In the present disclosure, in order to indicate that a packed video consists of geometry and attributes, pin_geometry_present_flag and pin_attribute_present_flag are both set to 1, as an example.

[0243] The value of pin_occupancy_2d_bit_depth_minus1[ j ] plus 1 represents the nominal 2D bit depth to which decoded regions containing occupancy data for the atlas with atlas ID j will be converted.

[0244] pin_occupancy_msb_align_flag[ j ] indicates how decoded regions containing occupancy samples associated with the atlas of atlas ID j are converted to the nominal occupancy bit depth.

[0245] pin_lossy_occupancy_compression_threshold[ j ] represents the threshold to be used to derive binary occupancy from decoded regions containing occupancy data for the atlas with atlas ID j.

[0246] The value of pin_geometry_2d_bit_depth_minus1[ j ] plus 1 represents the nominal 2D bit depth to which the decoded regions containing geometry data for the atlas with atlas ID j will be converted.

[0247] pin_geometry_msb_align_flag[ j ] indicates how decoded regions containing geometry samples associated with the atlas with atlas ID j are converted to nominal accommodative bit depth.

[0248] The value of pin_geometry_3d_coordinates_bit_depth_minus1[ j ] plus 1 represents the bit depth of the geometry coordinates of the reconstructed volume contents for the atlas with atlas ID j.

[0249] pin_attribute_count[ j ] represents the number of attributes with a unique attribute type that exist in the packed video frames for the atlas with atlas ID j.

[0250] pin_attribute_type_id[ j ][ i ] indicates the attribute type of the attribute having the i index in the atlas having the atlas ID j. For example, if the value of pin_attribute_type_id[ j ][ i ] is 0, it may indicate texture, if it is 1, it may indicate material ID, if it is 2, it may indicate transparency, if it is 3, it may indicate reflectance, and if it is 4, it may indicate normals. In the present disclosure, in order to indicate that the attribute data is composed of a texture, the value of pin_attribute_type_id is set to '0' as an embodiment.

[0251] The value of pin_attribute_2d_bit_depth_minus1[ j ][ i ] plus 1 represents the nominal 2D bit depth to which the decoded regions containing attribute data for the atlas with atlas ID j will be converted.

[0252] pin_attribute_msb_align_flag[ j ][ i ] indicates how decoded regions containing attribute samples associated with the atlas of atlas ID j are converted to nominal accommodative bit depth.

[0253] When pin_attribute_map_absolute_coding_persistence_flag[ j ][ i ] is 1, it indicates that attribute maps of attribute index i corresponding to the atlas of atlas ID j are coded without map prediction. When pin_attribute_map_absolute_coding_persistence_flag[ j ][ i ] is 0, it indicates that attribute maps of attribute index i corresponding to the atlas of atlas ID j are coded using the same map prediction method as that used for the geometry component corresponding to the atlas of atlas ID j.

[0254] The value of pin_regions_count_minus1[ j ] plus 1 represents the number of regions packed into one video frame for the atlas with atlas ID j. pin_regions_count_minus1 ranges from 0 to 255.

[0255] The following describes a method of signaling scalability information for packed video proposed in the present disclosure to packing information (packing_information()) with reference to FIGS. 19 to 21.

[0256] That is, the present disclosure proposes a signaling method for supporting the spatial resolution scalability function of the PVD data type. According to embodiments, when compressing packed video through a scalable video codec, information on a packed area for each layer may be signaled, or may be derived in the same manner by an encoder / decoder agreement. Information on the packed area may include the upper left x-coordinate of the packed area, the y-coordinate, and information on the width and height of the packed area.

[0257] The first embodiment is a method for signaling all information of a packing area corresponding to each layer in the packing information syntax, as shown in FIG. 19. When the packed video is encoded using a scalable codec, after signaling the number of layers of the scalable video codec, information of the packing area can be signaled in the number of layers. In the present disclosure, information of the packing area can be referred to as packing area information.

[0258] FIG. 19 is a diagram showing another example of the syntax structure of packing information (packing_information()) according to embodiments.

[0259] For convenience of explanation, syntax elements included in the packing information illustrated in FIGS. 18a and 18b are omitted in FIG. 19, and syntax elements included in the packing information of FIGS. 18a and 18b are also included in the packing information of FIG. 19, and detailed descriptions will refer to FIGS. 18a and 18b.

[0260] The packing information of Fig. 19 may include pin_max_layer_minus1[j] if the value of pin_codec_id[j] is 5 or 6.

[0261] In the present disclosure, pin_codec_id[ j ] represents a mapping index of a codec identifier of a video decoder used to decode a packed video sub-bitstream for atlas ID j. That is, the value of pin_codec_id[ j ] being 5 or 6 indicates that the codec of the video decoder is HEVC Scalable Main10 or VVC Multilayer Main10, as shown in FIG. 17. In other words, when pin_codec_id[j] == 5, the packed video is encoded with SHVC, and when pin_codec_id[j] == 6, the packed video is encoded with multilayer VVC.

[0262] pin_max_layer_minus1[j] represents the number of layers when packed video is encoded with a scalable video codec (SHVC, multilayer VVC).

[0263] According to embodiments, if packed video is encoded with a scalable video codec, the number of layers (num_layer) is set to pin_max_layer_minus1[j]+1, otherwise it is set to 1.

[0264] In addition, packing region information is included in the packing information as much as the value of pin_max_layer_minus1[j], and the packing region information is repeated as many times as pin_regions_count_minus1 and signaled to the packing information for each region. In the present disclosure, a layer may include one or more layers, or a layer may be associated with one or more regions.

[0265] The value of pin_regions_count_minus1[ j ] plus 1 represents the number of regions packed into one video frame for the atlas with ID j. pin_regions_count_minus1 ranges from 0 to 255.

[0266] The following is a description of the packing area information.

[0267] pin_region_tile_id[ j ][ l ][ i ] is the atlas tile ID associated with the ith region of the lth layer of the atlas with ID j, which is associated with the region corresponding to index i. That is, it represents the atlas tile ID of the ith region of the lth layer of the jth atlas.

[0268] The value of pin_region_type_id_minus2[ j ][ l ][ i ] plus 2 represents the ID of the ith region of the lth layer of the jth atlas. The value of pin_region_type_id_minus2[ j ][ l ][ i ] ranges from 0 to 2.

[0269] pin_region_top_left_x[ j ][ l ][ i ] specifies the horizontal position (i.e., x coordinate) of the top left sample of the ith region of the lth layer of the jth atlas in luma sample units in the packed video component frame.

[0270] pin_region_top_left_y[ j ][ l ][ i ] specifies the vertical position (i.e., y coordinate) of the top left sample of the ith region of the lth layer of the jth atlas in units of luma samples in the packed video component frame.

[0271] The value of pin_region_width_minus1[ j ][ l ][ i ] plus 1 specifies the horizontal width of the ith region of the lth layer of the jth atlas in luma samples.

[0272] The value of pin_region_height_minus1[ j ][ l ][ i ] plus 1 specifies the vertical height of the ith region of the lth layer of the jth atlas in luma samples.

[0273] pin_region_unpack_top_left_x[ j ][ l ][ i ] specifies the horizontal position of the top left sample of the ith region of the lth layer of the jth atlas in luma samples in the unpacked video component frame.

[0274] pin_region_unpack_top_left_y[ j ][ l ][ i ] specifies the vertical position of the top left sample of the ith region of the lth layer of the jth atlas in luma samples in the unpacked video component frame.

[0275] When pin_region_rotation_flag[ j ][ l ][ i ] is 0, it indicates that no rotation is performed on the ith region of the lth layer of the jth atlas. When pin_region_rotation_flag[ j ][ l ][ i ] is 1, it indicates that the ith region of the lth layer of the jth atlas is rotated 90 degrees.

[0276] pin_region_map_index[ j ][ l ][ i ] specifies the map index of the ith region of the lth layer of the jth atlas.

[0277] When pin_region_auxiliary_data_flag[ j ][ l ][ i ] is 1, it indicates that the ith region of the lth layer of the jth atlas contains only RAW and / or EOM coded points. When pin_region_auxiliary_data_flag is 0, it indicates that the ith region of the lth layer of the jth atlas can contain RAW and / or EOM coded points.

[0278] pin_region_attr_index[ j ][ l ][ i ] represents the attribute index of the ith region of the lth layer of the jth atlas. The value of pin_region_attr_index[ j ][ l ][ i ] ranges from 0 to pin_attribute_count[ j ] - 1.

[0279] pin_region_attr_partition_index[ j ][ l ][ i ] represents the attribute partition index of the ith region of the lth layer of the jth atlas. If not present, the value of pin_region_attr_partition_index[ j ][ l ][ i ] is inferred to be 0.

[0280] The second embodiment is a method for signaling a scale factor capable of deriving the width and height of the packing area of ​​an enhancement layer based on the base layer in the packing information syntax, as shown in FIG. 20. In the present disclosure, information on the packing area may be referred to as packing area information.

[0281] FIG. 20 is a diagram showing another example of the syntax structure of packing information (packing_information()) according to embodiments.

[0282] For convenience of explanation, syntax elements included in the packing information illustrated in FIGS. 18a and 18b are omitted in FIG. 20, and syntax elements included in the packing information of FIGS. 18a and 18b are also included in the packing information of FIG. 20, and detailed descriptions will be made with reference to FIGS. 18a and 18b.

[0283] In Fig. 20, the value of pin_regions_count_minus1[ j ] plus 1 represents the number of regions packed into one video frame for the atlas with ID j. pin_regions_count_minus1 ranges from 0 to 255.

[0284] The packing information of Fig. 20 may include pin_max_layer_minus1[j] if the value of pin_codec_id[j] is 5 or 6.

[0285] In the present disclosure, pin_codec_id[ j ] represents a mapping index of a codec identifier of a video decoder used to decode a packed video sub-bitstream for atlas ID j. That is, the value of pin_codec_id[ j ] being 5 or 6 indicates that the codec of the video decoder is HEVC Scalable Main10 or VVC Multilayer Main10, as shown in FIG. 17. In other words, when pin_codec_id[j] == 5, the packed video is encoded with SHVC, and when pin_codec_id[j] == 6, the packed video is encoded with multilayer VVC.

[0286] pin_max_layer_minus1[j] represents the number of layers when packed video is encoded with a scalable video codec (SHVC, multilayer VVC).

[0287] The present disclosure includes a loop that repeats as many times as the value of pin_max_layer_minus1[j], and the loop includes pin_region_scale_x[ j ][ l ][ i ] and pin_region_scale_y[ j ][ l ][ i ].

[0288] pin_region_scale_x represents the scale of the packed video width for each layer. Depending on the embodiment, it may be the scale of the enhancement layer relative to the base layer, or the scale of the enhancement layer relative to the base layer. pin_region_scale_x[ j ][ l ][ i ] represents the scale of the packed video width of the ith region of the lth layer of the jth atlas.

[0289] pin_region_scale_y represents the scale of the packed video height for each layer. Depending on the embodiment, it may be the scale of the enhancement layer relative to the base layer, or the scale of the enhancement layer relative to the base layer. pin_region_scale_y[ j ][ l ][ i ] represents the scale of the packed video height of the ith region of the lth layer of the jth atlas.

[0290] The following is a description of the packing area information.

[0291] pin_region_tile_id[ j ][ l ][ i ] is the atlas tile ID associated with the ith region of the atlas with ID j, which is associated with the region corresponding to index i. That is, it represents the atlas tile ID of the ith region of the jth atlas.

[0292] The value of pin_region_type_id_minus2[ j ][ i ] plus 2 represents the ID of the ith region of the jth atlas. The value of pin_region_type_id_minus2[ j ][ i ] ranges from 0 to 2.

[0293] pin_region_top_left_x[ j ]

[0000] [ i ] specifies the horizontal position (i.e., x coordinate) of the top left sample of the i-th region of the 0th layer (i.e., the base layer) of the j-th atlas in units of luma samples in the packed video component frame.

[0294] pin_region_top_left_y[ j ]

[0000] [ i ] specifies the vertical position (i.e., y coordinate) of the top left sample of the i-th region of the 0th layer (i.e., the base layer) of the j-th atlas in units of luma samples in the packed video component frame.

[0295] The value of pin_region_width_minus1[ j ]

[0000] [ i ] plus 1 specifies the horizontal width of the i-th region of the 0th layer (i.e., the base layer) of the j-th atlas in luma samples.

[0296] The value of pin_region_height_minus1[ j ]

[0000] [ i ] plus 1 specifies the vertical height of the i-th region of the 0th layer (i.e., the base layer) of the j-th atlas in luma samples.

[0297] The process of deriving packing region information using scale factors parsed from the decoder (e.g., pin_region_scale_x[ j ][ l ][ i ] and pin_region_scale_y[ j ][ l ][ i ]) will be explained later.

[0298] pin_region_unpack_top_left_x[ j ][ i ] specifies the horizontal position of the top left sample of the ith region of the jth atlas in luma samples in the unpacked video component frame.

[0299] pin_region_unpack_top_left_y[ j ][ i ] specifies the vertical position of the top left sample of the ith region of the jth atlas in luma samples in the unpacked video component frame.

[0300] When pin_region_rotation_flag[ j ][ i ] is 0, it indicates that no rotation is performed for the ith region of the jth atlas. When pin_region_rotation_flag[ j ][ i ] is 1, it indicates that the ith region of the jth atlas is rotated by 90 degrees.

[0301] pin_region_map_index[ j ][ i ] specifies the map index of the ith region of the jth atlas.

[0302] When pin_region_auxiliary_data_flag[ j ][ i ] is 1, it indicates that the ith region of the jth atlas contains only RAW and / or EOM coded points. When pin_region_auxiliary_data_flag is 0, it indicates that the ith region of the jth atlas may contain RAW and / or EOM coded points.

[0303] pin_region_attr_index[ j ][ i ] represents the attribute index of the ith region of the jth atlas. The value of pin_region_attr_index[ j ][ i ] ranges from 0 to pin_attribute_count[ j ] - 1.

[0304] pin_region_attr_partition_index[ j ][ i ] represents the attribute partition index of the ith region of the jth atlas. If not present, the value of pin_region_attr_partition_index[ j ][ i ] is inferred to be 0.

[0305] The following describes the process of deriving packing area information using the scale factor parsed from the decoder.

[0306] The pseudo code below is an example of deriving packing area information (upper left coordinate of packing area, width and height of packing area) of the enhancement layer based on the base layer.

[0307] for(l = 0; l < pin_max_layer_minus1[ j ]; l++) {

[0308] for( i = 0; i <= pin_regions_count_minus1[ j ]; i++ ) {

[0309] pin_region_width_minus1[ j ][ l + 1 ][ i ] =

[0310] pin_region_width_minus1[ j ][ l ][ i ] * pin_region_scale_x[ j ][ l ][ i ]

[0311] pin_region_height_minus1[ j ][ l + 1 ][ i ] =

[0312] pin_region_ height_minus1[ j ][ l ][ i ] * pin_region_scale_y[ j ][ l ][ i ]

[0313] pin_region_top_left_x[ j ][ l+1 ][ i ] = pin_region_top_left_x[ j ][ l ][ i ] * pin_region_scale_x[ j ][ l ][ i ]

[0314] pin_region_top_left_y[ j ][ l+1 ][ i ] = pin_region_top_left_y[ j ][ l ][ i ] * pin_region_scale_y[ j ][ l ][ i ]

[0315] }

[0316] }

[0317] For example, if packed video is encoded via a 2-layer scalable codec, pin_max_layer_minus1[ j ] can be 1 and pin_regions_count_minus1[ j ] can be 1.

[0318] pin_region_width_minus1[ j ][ l + 1 ][ i ] = pin_region_width_minus1[ j ][ l ][ i ] * pin_region_scale_x[ j ][ l ][ i ]

[0319] The above pseudo code is a process of calculating the width of the upper layer (pin_region_width_minus1[ j ][ l ][ i ]) by multiplying the width of the lower layer (pin_region_width_minus1[ j ][ l ][ i ]) by the scale factor (pin_region_scale_x[ j ][ l ][ i ]).

[0320] pin_region_height_minus1[ j ][ l + 1 ][ i ] = pin_region_ height_minus1[ j ][ l ][ i ] * pin_region_scale_y[ j ][ l ][ i ]

[0321] The above pseudo code is the process of calculating the height of the upper layer (pin_region_height_minus1[ j ][ l ][ i ]) by multiplying the height of the lower layer (pin_region_height_minus1[ j ][ l ][ i ]) by the scale factor (pin_region_scale_y[ j ][ l ][ i ]).

[0322] pin_region_top_left_x[ j ][ l+1 ][ i ] = pin_region_top_left_x[ j ][ l ][ i ] * pin_region_scale_x[ j ][ l ][ i ]

[0323] The above pseudo code is the process of multiplying the upper left x-coordinate of the lower layer (pin_region_top_left_x[ j ][ l ][ i ]) by the scale factor (pin_region_scale_x[ j ][ l ][ i ]) to obtain the upper left x-coordinate of the upper layer (pin_region_top_left_x[ j ][ l+1 ][ i ]).

[0324] pin_region_top_left_y[ j ][ l+1 ][ i ] = pin_region_top_left_y[ j ][ l ][ i ] * pin_region_scale_y[ j ][ l ][ i ]

[0325] The above pseudo code is the process of multiplying the upper left y-coordinate of the lower layer (pin_region_top_left_y[ j ][ l ][ i ]) by the scale factor (pin_region_scale_y[ j ][ l ][ i ]) to obtain the upper left y-coordinate of the upper layer (pin_region_top_left_y[ j ][ l+1 ][ i ]).

[0326] FIG. 21 is a diagram showing an example of packing information of a texture map and a displacement vector when there are two layers according to embodiments.

[0327] In Fig. 21, the upper left coordinates of the packed texture map and displacement vector of the base layer are defined as (x1, x1) and (x2, y2), respectively, and the upper left coordinates of the enhancement layer 1 are defined as (x3, y3) and (x4, y4), respectively. In addition, the areas of the base layer and the enhancement layer 1 in the packed area are defined as W1 and W2, respectively, and the heights of the base layer and the enhancement layer 1 in the packed area are defined as H1, H2, H3, and H4 for the texture map and the displacement vector, respectively. At this time, the values ​​of W1 and W2 may be the same, and the values ​​of H1, H2, H3, and H4 may be different from each other or some may be the same.

[0328] According to the embodiments, the scale factor of the width of the base layer and the enhancement layer of the texture map and the displacement vector can be obtained as W2 / W1.

[0329] According to embodiments, the scale factor of the height of the base layer and the enhancement layer of the texture map can be obtained as H3 / H1.

[0330] According to the embodiments, the scale factor of the height of the base layer and the enhancement layer of the displacement vector can be obtained as H4 / H2.

[0331] The third embodiment is a method for deriving the packing information of the enhancement layer from the decoder without signaling it. That is, the packing area information can be derived using the size information of the packed video output from the decoder. This method may require the constraint that the scales of the texture map area and the displacement vector area per layer must be identical.

[0332] The following is a description of the process of deriving packing area information according to the third embodiment.

[0333] Scale_width and Scale_height are the size information of the packed video.

[0334] Scale_width = Enhancement_layer_width / Base_layer_width

[0335] Scale_height = Enhancement_layer_height / Base_layer_height

[0336] for(l = 0; l < pin_max_layer_minus1[ j ]; l++) {

[0337] for( i = 0; i <= pin_regions_count_minus1[ j ]; i++ ) {

[0338] pin_region_width_minus1[ j ][ l + 1 ][ i ] = pin_region_width_minus1[ j ][ l ][ i ] * Scale_width

[0339] pin_region_height_minus1[ j ][ l + 1 ][ i ] = pin_region_ height_minus1[ j ][ l ][ i ] * scale_height

[0340] pin_region_top_left_x[ j ][ l+1 ][ i ] = pin_region_top_left_x[ j ][ l ][ i ] * Scale_width

[0341] pin_region_top_left_y[ j ][ l+1 ][ i ] = pin_region_top_left_y[ j ][ l ][ i ] * Scale_height

[0342] }

[0343] }

[0344] According to embodiments, the width of the upper layer (pin_region_width_minus1[ j ][ l ][ i ]) can be obtained by multiplying the width of the lower layer (pin_region_width_minus1[ j ][ l ][ i ]) by Scale_width.

[0345] According to embodiments, the height of the upper layer (pin_region_height_minus1[ j ][ l ][ i ]) can be obtained by multiplying the height of the lower layer (pin_region_height_minus1[ j ][ l ][ i ]) by Scale_height.

[0346] According to embodiments, the upper left x-coordinate of the upper layer (pin_region_top_left_x[ j ][ l ][ i ]) can be obtained by multiplying the upper left x-coordinate of the lower layer (pin_region_top_left_x[ j ][ l+1 ][ i ]) by Scale_width.

[0347] According to embodiments, the top left y-coordinate of the upper layer (pin_region_top_left_y[ j ][ l ][ i ]) can be obtained by multiplying the top left y-coordinate of the lower layer (pin_region_top_left_y[ j ][ l+1 ][ i ]) by Scale_height.

[0348] The following describes how to signal packing information for scalable coding-independent decoding.

[0349] Packing information signaling for scalable coding-independent decoding

[0350] According to embodiments, information for supporting the function of independently decoding PVD unit types can be signaled as an SEI message. That is, video codecs support the function of independently decoding, and HEVC can independently decode through the Motion-Constrained Tile Set (MCTS) function, and VVC can independently decode through the subpicture function. In the present disclosure, by signaling the tile id (in case of HEVC) and subpicture id (in case of VVC) of each packing area of ​​PVD data through an SEI message, the decoder can decode only the corresponding area through the tile id / subpicture id information.

[0351] According to embodiments, when scalable coding is performed, ID information of tiles or subpictures can be transmitted for each layer.

[0352] FIG. 22a and FIG. 22b are diagrams showing an example of a syntax structure of an SEI message for independent decoding of PVD data by region according to embodiments.

[0353] pir_num_packed_frames_minus1 plus 1 specifies the number of packed frames in which independently decodable region information is signaled.

[0354] pir_multilayer_enabled_flag is a flag indicating whether scalable codecs are available. If the value of pir_multilayer_enabled_flag is 1, scalable codecs are available, and if it is 0, they are not available.

[0355] pir_max_layer_minus1 refers to the number of layers when a scalable codec is available.

[0356] At this time, if a scalable codec is available, the number of layers (num_layer) is set to pir_max_layer_minus1+1, otherwise it is set to 1.

[0357] The SEI message of FIG. 22a and FIG. 22b is repeated as many times as the value of pir_num_packed_frames_minus1 and includes packing area information for each frame.

[0358] The following is a description of the packing area information.

[0359] pir_packed_frame_id[ j ] specifies the ID of the pack (i.e., packed frame) with index j. The value of pir_packed_frame_id[ j ] ranges from 0 to 15. And, below, k is set to pir_packed_frame_id[ j ].

[0360] pir_description_type_idc[ k ] indicates the type of video codec of the packed frame with index j. For example, a value of pir_description_type_idc[ k ] of 0 indicates HEVC, 1 indicates VVC, 2 indicates SHVC, and 3 indicates multilayer VVC.

[0361] More specifically, if pir_description_type_idc[ k ] == 2 , i.e. SHVC, the upper left index and lower right index information of the tile are signaled for each packing area, and if pir_description_type_idc[ k ] == 3 , i.e. multilayer VVC, the id information of the subpicture is signaled for each packing area.

[0362] The value of pir_num_regions_minus1[ k ] plus 1 is the number of rectangular regions of the kth packed frame, which specifies the number of rectangular regions of the packed frame with ID k for which independently decodable region information is signaled.

[0363] According to embodiments, the packing area information may include pir_top_left_tile_idx[ k ][ i ] and pir_bottom_right_tile_idx[ k ][ i ] if pir_description_type_idc[ k ] == 0.

[0364] pir_top_left_tile_idx[ k ][ i ] and pir_bottom_right_tile_idx[ k ][ i ] are the tile indices of the top left tile and the bottom right tile of the independently decodable region corresponding to the i-th region in the video sub-bitstream of the k-th packed frame, respectively.

[0365] According to embodiments, the packing area information may include pir_subpic_id[ k ][ i ] if pir_description_type_idc[ k ] == 1.

[0366] pir_subpic_id[ k ][ i ] represents the subpicture ID corresponding to the i-th rectangle in the video sub-bitstream of the k-th packed frame.

[0367] According to embodiments, the packing area information may include pir_top_left_tile_idx[ k ][ i ][ l ] and pir_bottom_right_tile_idx[ k ][ i ][ l ], repeated as many times as the numlayer value, if pir_description_type_idc[ k ] == 2.

[0368] pir_top_left_tile_idx represents the top left tile index of the packed area i. That is, pir_top_left_tile_idx[ k ][ i ][ l ] represents the top left tile index of the i-th area associated with the l-th layer in the video sub-bitstream of the k-th packed frame.

[0369] pir_bottom_right_tile_idx represents the tile index of the bottom right of the packed area i. That is, pir_bottom_right_tile_idx[ k ][ i ][ l ] represents the tile index of the bottom right of the i-th area associated with the l-th layer in the video sub-bitstream of the k-th packed frame.

[0370] According to embodiments, the packing area information may include pir_subpic_id[ k ][ i ][ l ] repeated as many times as the numlayer value if pir_description_type_idc[ k ] == 3.

[0371] pir_subpic_id represents the subpicture id of the packing region i. That is, pir_subpic_id[ k ][ i ][ l ] represents the subpicture id of the i-th region associated with the l-th layer in the video sub-bitstream of the k-th packed frame.

[0372] FIG. 23 is a diagram showing another example of a transmitter according to embodiments. The transmitter of FIG. 23 may correspond to the transmitter of FIG. 1, the transmitter of FIG. 6, the transmitter of FIG. 7, or the transmitter of FIG. 13. Therefore, for parts not described in FIG. 23, reference will be made to the description of the transmitter of FIG. 1, the transmitter of FIG. 6, the transmitter of FIG. 7, or the transmitter of FIG. 13. The elements of the transmitter illustrated in FIG. 23 may be implemented by hardware, software, a processor connected to a memory, and / or a combination thereof. That is, the elements of the transmitter of FIG. 23 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the elements of the transmitter of FIG. 23 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing operations and / or functions of elements of the transmitting device of FIG. 23.

[0373] Referring to FIG. 23, dynamic mesh video data acquired through the dynamic mesh video acquisition unit can be encoded in each encoder for each atlas component, base mesh component, displacement component, and attribute component.

[0374] For example, an atlas component is encoded into an atlas sub-bitstream in an atlas data encoder, and a basemesh component is encoded into a basemesh sub-bitstream in a basemesh encoder and then input to a multiplexer.

[0375] As another example, if packed video is not supported, the displacement component is encoded into a geometry video sub-bitstream in the displacement encoder, and the attribute component is encoded into an attribute video sub-bitstream in the attribute encoder. In this case, attribute video may be a term that includes texture video.

[0376] As another example, when supporting packed video, a packed video component (i.e., video data in which displacement data and texture map data are packed into one frame) is encoded into a packed video sub-bitstream by a packed encoder. According to embodiments, the packed encoder may perform video codec-based scalable encoding on packed video data in which displacement data and texture map data are packed into one frame. That is, data (texture map and displacement information) of a frame packed for each layer may be encoded using each independent video codec. For example, if the number of layers including the base layer is 2, data of a packed video frame of the corresponding layer may be encoded using two independent video codecs. Here, the packed encoder may be an SHVC encoder, a VVC encoder, or another encoder capable of scalable coding.

[0377] That is, if packed video is not supported, the geometry video sub-bitstream and the attribute video sub-bitstream are input to the multiplexer, and if packed video is supported, the packed video sub-bitstream is input to the multiplexer.

[0378] If the multiplexer does not support packed video, it multiplexes the atlas sub-bitstream, basemesh sub-bitstream, geometry video sub-bitstream, and attribute video sub-bitstream into a single bitstream of mesh data (e.g., a V3C bitstream).

[0379] If the multiplexer supports packed video, it multiplexes the atlas sub-bitstream, basemesh sub-bitstream, and packed video sub-bitstream into a single bitstream of mesh data (e.g., a V3C bitstream).

[0380] The bitstream of mesh data output from the multiplexer may be transmitted as is to the receiving device through the file / segment encapsulation and transmission unit, or may be encapsulated in a file format such as ISOBMFF through the file / segment encapsulation and transmission unit, or may be processed in the form of other DASH segments and then transmitted. The file / segment encapsulation and transmission unit may include mesh video-related metadata in the file format according to an embodiment. The mesh video-related metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. According to an embodiment, the file / segment encapsulation and transmission unit may encapsulate the mesh video-related metadata itself into a file. According to embodiments, the bitstream of mesh data multiplexed in the multiplexer may be transmitted through a network or may be stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.

[0381] In the present disclosure, mesh video related metadata may include packing information of FIGS. 18 to 22.

[0382] That is, in the dynamic mesh encoding process of the present disclosure, if Packed video is not supported depending on whether the packed video type is supported, displacement data and attribute frames can be encoded through a displacement encoder and an attribute encoder, respectively. The displacement encoder can perform transformation and quantization on the displacement data, and then perform encoding through video coding, arithmetic coding, etc. The attribute encoder can regenerate a texture map by performing a texture transfer process and a padding process on an input texture map, and can perform color space conversion, etc. on the regenerated texture map, and then encode it through video coding. That is, the texture transfer process is a process of regenerating a texture map of a current mesh based on an input mesh (i.e., a texture map (or attribute map) of an original mesh) and a restored mesh (i.e., a full-resolution geometry).

[0383] In contrast, when packed video is supported, a packed encoder can perform encoding by packing a displacement frame and an attribute frame into a single frame. In this case, the packed encoder can perform transformation and quantization on the displacement data, and then pack the quantized displacement vector transform coefficients into a frame. In addition, the packed encoder can perform video coding on the attribute frame, and then pack the displacement vector transform coefficients into a packed frame. At this time, the attribute frame may be a texture map generated through a texture transfer and padding process, a color space conversion, etc. According to embodiments, the displacement frame and the attribute frame may be packed into a single frame, or may be packed in the up, down, left, and right directions, and information on the corresponding packed area may be signaled through the packing information of FIGS. 18 to 22.

[0384] Additionally, as described above, packing information for multi-layers may be signaled at the VPS level as in FIGS. 19 to 21, or may be derived without signaling packing information for multi-layers.

[0385] When packing information for multilayers is signaled at the VPS level, i.e., when scalable coding of packed video is supported in the manner proposed in the present disclosure, packing information, which is information on the packed area for each component, can be signaled at the VPS level as shown in FIGS. 19 to 21. For details, refer to the descriptions of FIGS. 19 to 21.

[0386] Meanwhile, if the packing information for multilayers is derived without signaling, i.e., if scalable coding of packed video is supported in the manner proposed in the present disclosure, the frame size can be derived so that the packed area of ​​the displacement frame and the packed area of ​​the attribute frame can maintain the same ratio of resolution for each layer. Since a detailed description of this case has been provided above, it will be omitted here.

[0387] Fig. 24 is a drawing showing another example of a receiving device according to embodiments. Fig. 23 is a block diagram showing another example of a receiving device according to embodiments. The receiving device of Fig. 24 may be referred to as a dynamic mesh content receiving device. The receiving device of Fig. 24 may correspond to the receiving device of Fig. 1, the receiving device of Fig. 11, the receiving device of Fig. 12, or the receiving device of Fig. 14. Therefore, parts not described in Fig. 24 will refer to the description of the receiving device of Fig. 1, the receiving device of Fig. 11, the receiving device of Fig. 12, or the receiving device of Fig. 14. The elements of the receiving device illustrated in Fig. 24 may be implemented by hardware, software, a processor connected to a memory, and / or a combination thereof. That is, the elements of the receiving device of Fig. 24 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not illustrated in the drawing. One or more processors may perform at least one of the operations and / or functions of the elements of the receiving device of FIG. 24 described above. Furthermore, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the receiving device of FIG. 24.

[0388] In particular, Fig. 24 is an example of performing scalable decoding and restoring the texture map and displacement information by separating them (split) when the texture map and displacement information are packed into one frame as in Fig. 23 and transmitted by performing video codec-based encoding.

[0389] According to embodiments, a bitstream of mesh data encapsulated in a file at a transmitting device and delivered to a receiving device through a delivery module is decapsulated in a file decapsulation module. If the bitstream of mesh data is not encapsulated in a file format at the transmitting device, the decapsulation process at the receiving device is omitted. In the present disclosure, a base mesh sub-bitstream may be referred to as a base mesh bitstream, an atlas sub-bitstream may be referred to as an atlas bitstream, and a PVD video sub-bitstream may be referred to as a PVD video bitstream.

[0390] That is, the stored and / or received dynamic mesh content passes through the delivery module and the file decapsulation module and can be in a form similar to the bitstream structure of FIG. 15 (i.e., dynamic mesh bitstream).

[0391] And, in the mesh data decoding module, the base mesh sub-bitstream, the atlas sub-bitstream, and the PVD video sub-bitstream can be separated and then decoded through each decoder.

[0392] According to embodiments, a bitstream parser (not shown) of a mesh data decoding module may be responsible for parsing a dynamic mesh content bitstream. That is, a V3C unit header and a V3C unit payload constituting the bitstream may be parsed to obtain a base mesh bitstream, an atlas bitstream, and a PVD video bitstream, respectively, and each may be decoded by a respective decoder.

[0393] Additionally, the bitstream parser of the mesh data decoding module can obtain data corresponding to the unit type V3C_VPS defined in FIG. 16, i.e., VPS data, by parsing the V3C unit header and V3C unit payload.

[0394] The following is a detailed description of the process of performing scalable decoding based on signaling information for a PVD video bitstream in the mesh data decoding module of the receiving device of FIG. 24.

[0395] When scalable coding information is signaled as an extension to the V3C parameter set.

[0396] The mesh data decoding module of the receiving device can obtain codec information for signaled scalable coding by parsing the profile_tier_level() information contained in the VPS data acquired in the above process (see Fig. 17). Based on the acquired codec information for scalable coding, the mesh data decoding module can determine whether the receiver supports the corresponding codec, and predict and determine subsequent decoding operations.

[0397] Additionally, the mesh data decoding module can obtain extension information signaled in the VPS data (e.g., vps_packed_video_extension, packing_information) obtained in the above process, i.e., scalable coding information and packing area information as in Fig. 19 or Fig. 20. Alternatively, the mesh data decoding module can obtain scalable coding information and packing area information as in Fig. 22a and Fig. 22b from an SEI message.

[0398] As described above, when compressing packed video using a scalable video codec, information about a packed area for each layer may be signaled as in FIG. 19 or FIG. 20, or may be derived in the same manner by an encoder / decoder agreement. Information about a packed area may include the upper left x-coordinate of the packed area, the y-coordinate, and information about the width and height of the packed area.

[0399] Figure 19 illustrates a method for signaling all information about the packing area corresponding to each layer in the packing information syntax. When the packed video is encoded using a scalable codec, after signaling the number of layers of the scalable video codec, information about the packing area corresponding to the number of layers can be signaled.

[0400] Figure 20 is a method for signaling a scale factor that can derive the packing area width and height of an enhancement layer relative to a base layer in a packing information syntax.

[0401] The process of deriving packing area information using a scale factor parsed from packing information such as Fig. 20 in the mesh data decoding module is described in detail in Fig. 20, so a description thereof is omitted here.

[0402] In another embodiment, the present disclosure can derive the packing information of the enhancement layer from the mesh data decoding module of the receiving device without signaling it. That is, the packing area information can be derived using the size information of the packed video output from the mesh data decoding module. This method requires the constraint that the scales of the texture map area and the displacement vector area must be the same for each layer.

[0403] The present disclosure provides another embodiment in which information for supporting the function of independently decoding PVD unit types can be signaled as an SEI message, as in FIGS. 22a and 22b. By signaling the tile ID (in the case of HEVC) and subpicture ID (in the case of VVC) of each packing area of ​​PVD data as an SEI message, a mesh data decoding module can decode only the corresponding area through the tile / subpicture ID information. That is, when scalable coding is performed as in FIGS. 22a and 22b, ID information of tiles or subpictures can be received for each layer.

[0404] In this way, the receiving device can identify displacement vector and texture map data information encoded in the bitstream of the dynamic mesh content using scalable coding prior to the direct decoding process of the packed video data through the above process in the manner proposed in the present disclosure.

[0405] And, based on the information obtained above, scalable decoding can be performed on packed video.

[0406] Fig. 25 is a flowchart illustrating an example of a transmission method according to embodiments. The transmission method according to embodiments may include a step of encoding mesh data (S31011) and a step of transmitting a bitstream including the encoded mesh data (S31012). In one embodiment, the bitstream transmitted in step (S31012) includes an atlas bitstream, a base mesh bitstream, and a PVD video bitstream.

[0407] According to embodiments, the step of encoding mesh data (S31011) may include a process of encoding a base mesh, a process of packing and scalable encoding displacement vectors or displacement vector transformation coefficients and texture maps into one frame in units of layers, and a process of signaling related signaling information. Here, the layer units may be a base layer and one or more enhancement layers. In the step of encoding mesh data (S31011), the process of packing and scalable encoding displacement vectors or displacement vector transformation coefficients and texture maps into one frame in units of layers, and the process of signaling related signaling information will be described with reference to the descriptions of FIGS. 15 to 23, and are omitted here to avoid redundant description.

[0408] In the step (S31012) of transmitting a bitstream including the above mesh data, the atlas bitstream, base mesh bitstream, and PVD video bitstream generated as described above in the step (S31011) of encoding the above mesh data are generated as one bitstream as shown in FIG. 15, and transmitted to a receiving device through a transmitting unit.

[0409] Fig. 26 is a flowchart showing an example of a receiving method according to embodiments. The receiving method according to embodiments may include a step of receiving a bitstream including mesh data (S32011) and a step of decoding the mesh data included in the bitstream (S32012). The step of receiving a bitstream including mesh data (S32011) or the step of decoding the mesh data (S32012) separates an atlas bitstream, a basemesh bitstream, and a PVD video bitstream from the received bitstream according to type information of a V3C unit header. In the step of decoding the mesh data (S32012), scalable decoding is performed on the PVD video bitstream based on signaling information as in Fig. 18 or Fig. 22, and a process of dividing a texture map and displacement information from a target frame and reconstructing dynamic mesh data is performed. In the step of decoding mesh data (S32012), the process of decoding packed video data and reconstructing dynamic mesh data is described in, for example, FIGS. 18 to 22 and 24, and is omitted here to avoid redundant description.

[0410] As explained so far, the method proposed in this disclosure defines the syntax required when applying scalable coding to data constituting dynamic mesh content, such as PVD type displacement vectors and texture maps, by packing them into one frame and according to predefined criteria and / or newly defined LoD values.

[0411] According to embodiments, the displacement vector and texture map packed data of dynamic mesh content can be encoded as layer-based scalable coded PVD data in a spatial or temporal manner according to the scalable coding syntax.

[0412] According to embodiments, scalable coding-related information such as codec information, layer information, LoD, etc. related to scalable coded displacement vectors and texture maps can be signaled at the parameter set level of the bitstream constituting the dynamic mesh content.

[0413] Accordingly, the dynamic mesh content receiver can efficiently access the bitstreams that make up the dynamic mesh content, since it can determine whether there is data with scalable coding applied within the content before actually decoding the geometry data and texture data encoded with the video codec.

[0414] Additionally, by packing geometry data and texture data into one frame at a time, synchronization issues between data can be minimized, enabling efficient scalable coding services to be provided.

[0415] In addition, the dynamic mesh content receiver can effectively decode and render scalable coded data in whole or in a selective manner, depending on the hardware constraints of the receiver, such as its resources, display, etc., and / or the intent and profile definitions of the content creator, user, or receiver itself.

[0416] Each of the parts, modules, or units described above may be software, processors, or hardware parts that execute sequential execution processes stored in memory (or storage units). Each of the steps described in the embodiments described above may be performed by processors, software, or hardware parts. Each of the modules / blocks / units described in the embodiments described above may operate as a processor, software, or hardware. In addition, the methods presented in the embodiments may be implemented as code. This code may be written on a processor-readable storage medium and thus may be read by a processor provided by an apparatus.

[0417] Furthermore, throughout the specification, when a part is said to "include" a component, this does not exclude other components, unless otherwise specifically stated, but rather implies the inclusion of other components. Furthermore, terms such as "part" described in the specification mean a unit that processes at least one function or operation, which may be implemented using hardware, software, or a combination of hardware and software.

[0418] For convenience of explanation, this specification has been described separately in each drawing. However, it is also possible to design new embodiments by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by those skilled in the art, is also within the scope of the embodiments.

[0419] The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.

[0420] Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which the present disclosure pertains without departing from the spirit or scope of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.

[0421] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. The components according to the embodiments may be implemented by separate chips. At least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.

[0422] In this document, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" means "at least one of A, B, and / or C". Additionally, "or" in this document is interpreted as "and / or". For example, "A or B" can mean 1) "A" only, 2) "B" only, or 3) "A and B". In other words, "or" in this document can mean "additionally or alternatively".

[0423] Various elements of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be implemented on a single chip, such as a hardware circuit. In some embodiments, the embodiments may optionally be implemented on separate chips. In some embodiments, at least one of the elements of the embodiments may be implemented within one or more processors that include instructions for performing operations according to the embodiments.

[0424] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including one or more memories and / or one or more processors according to the embodiments. One or more memories may store programs for processing / controlling the operations according to the embodiments, and one or more processors may control various operations described in this document. One or more processors may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in a processor or a memory.

[0425] Terms such as "first" and "second" may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted in a limited manner by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not necessarily mean the same user input signals unless the context clearly indicates otherwise.

[0426] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of the terms. The expression “comprises” or “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.

[0427] As described above, the relevant contents have been described in the best form for carrying out the embodiments.

[0428] As described above, the embodiments may be applied, in whole or in part, to mesh data transmission and reception devices and systems. Those skilled in the art will appreciate that various modifications and variations may be made to the embodiments within the scope of the embodiments. The embodiments may include modifications and variations, and such modifications and variations do not depart from the scope of the claims and their equivalents.

Claims

1. A step of receiving a bitstream containing mesh data; and A step of decoding the above mesh data; comprising: How to decode.

2. In the first paragraph, the step of decoding the mesh data A base mesh processing step for restoring a base mesh from a base mesh bitstream included in the above bitstream; A packed video data processing step for separating and restoring displacement information and attribute information from the packed video data bitstream included in the bitstream based on signaling information; and A decoding method comprising a restoration step of restoring a mesh based on the base mesh and the displacement information.

3. In the second paragraph, the packed video data processing step A step of scalable decoding packed video data from the packed video data bitstream based on layers including a base layer and one or more enhancement layers; and A decoding method comprising a step of separating and restoring displacement information and attribute information from packed video data of the decoded target packed video frame based on the signaling information.

4. In the third paragraph, the signaling information is Contains packing area information for the packed area for each layer, A decoding method in which the above packing area information includes information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

5. In the third paragraph, the signaling information is Contains packing area information for the packed area for the above base layer, A decoding method in which the above packing area information includes information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

6. In paragraph 5, The packing area information for the above enhancement layer is derived from the packing area information of the base layer based on scale factor information, A decoding method wherein the above scale factor information is included in the signaling information or is derived directly from the decoding device.

7. Memory; and comprising at least one processor connected to said memory, At least one processor of the above: Receive a bitstream containing mesh data; and Decode the above mesh data; configured to do so, Decoding device.

8. In the 7th paragraph, the at least one processor A base mesh processing unit that restores a base mesh from a base mesh bitstream included in the above bitstream; A packed video data processing unit that separates and restores displacement information and attribute information from the packed video data bitstream included in the bitstream based on signaling information; and A decoding device including a mesh restoration unit that restores a mesh based on the base mesh and the displacement information.

9. In the 8th paragraph, the packed video data processing unit A decoding device that scalably decodes packed video data from the packed video data bitstream based on layers including a base layer and one or more enhancement layers, and separates and restores displacement information and attribute information from the packed video data of the decoded target packed video frame based on the signaling information.

10. In paragraph 8, the signaling information is Contains packing area information for the packed area for each layer, A decoding device wherein the above packing area information includes information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

11. In paragraph 8, the signaling information is Contains packing area information for the packed area for the above base layer, A decoding device wherein the above packing area information includes information on the upper left position of the packed area, information on the width of the packed area, and information on the height of the packed area.

12. In paragraph 11, The packing area information for the above enhancement layer is derived from the packing area information of the base layer based on scale factor information, A decoding device wherein the above scale factor information is included in the signaling information or is derived directly from the decoding device.

13. Step of encoding mesh data; and A step of transmitting a bitstream including the encoded mesh data; comprising: Encoding method.

14. A computer-readable storage medium storing a bitstream generated by the method according to Article 13.

15. Step of obtaining a bitstream for video information; wherein the bitstream is generated based on a step of encoding mesh data and a step of transmitting a bitstream including the encoded mesh data; and A method comprising the step of transmitting data including the bitstream.

Citation Information

Patent Citations

  • Composite active material for negative electrode, method for manufacturing the same, negative electrode and secondary battery comprising the same

    KR1020230060965A

  • Base Mesh Data and Motion Information Sub-Stream Format for Video-Based Dynamic Mesh Compression

    US20240022765A1

  • A method, an apparatus and a computer program product for video encoding and video decoding

    WO2024012765A1

  • 3D data transmission device, 3D data transmission method, 3D data reception device, and 3D data reception method

    WO2024063544A1