Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method
The method efficiently encodes and decodes mesh data by restoring base meshes and displacement information, optionally omitting texture map coding, addressing latency and complexity issues in dynamic mesh data transmission for VR, AR, and autonomous driving.
Patent Information
- Application Number
- PCT/KR2024/019872
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2024-12-05
- Publication Date
- 2025-07-03
AI Technical Summary
The challenge lies in efficiently transmitting and receiving large amounts of mesh data, such as point cloud or mesh data, which is complex due to the high latency and encoding/decoding requirements, particularly in applications like Virtual Reality (VR), Augmented Reality (AR), and autonomous driving, where dynamic mesh data changes over time.
A method and device for encoding and decoding mesh data by restoring a base mesh, displacement information, and texture maps, with the option to omit texture map coding when similarity is high, using SEI messages for signaling texture map restoration information, and employing V3C unit types for efficient compression and transmission.
This approach reduces the bit rate and encoding/decoding complexity by omitting texture map coding, achieving significant bit savings and improved transmission speed while maintaining quality in dynamic mesh data applications.
Smart Images

Figure KR2024019872_03072025_PF_FP_ABST
Abstract
Description
Mesh data transmission device, mesh data transmission method, mesh data reception device, and mesh data reception method
[0001] The embodiments provide a method for providing 3D content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services.
[0002] Among 3D content, point cloud data and mesh data are collections of points in 3D space. However, the sheer number of points in 3D space makes it difficult to generate point cloud or mesh data.
[0003] That is, there is a problem that a lot of processing is required to transmit and receive 3D data with a large amount of points, such as point cloud data or mesh data.
[0004] The technical problem according to the embodiments is to provide a device and method for efficiently transmitting and receiving mesh data in order to solve the problems described above.
[0005] The technical problem according to the embodiments is to provide a device and method for resolving latency and encoding / decoding complexity of mesh data.
[0006] A technical problem according to embodiments is to provide a device and method for efficiently performing encoding and decoding of texture map data constituting mesh data.
[0007] However, the scope of the embodiments is not limited to the aforementioned technical tasks, and the scope of the embodiments may be expanded to other technical tasks that can be inferred by a person skilled in the art based on the entire contents of this document.
[0008] To achieve the above-described purpose and other advantages, a decoding method according to embodiments may include a step of receiving a bitstream including mesh data, and a step of decoding the mesh data.
[0009] According to embodiments, the decoding step may include a base mesh processing step of restoring a base mesh from a base mesh bitstream, a displacement information processing step of restoring displacement information from a displacement vector bitstream, a texture map processing step of restoring texture maps from a texture map bitstream, a step of determining whether at least one texture map is omitted from the texture map bitstream, and a step of restoring the at least one omitted texture map based on texture map restoration-related information and at least one reference frame if it is determined that there is the at least one omitted texture map.
[0010] According to embodiments, the texture map restoration related information may include information for identifying whether at least one texture map is omitted and reference texture map information related to the at least one reference frame.
[0011] According to embodiments, the texture map restoration related information may be carried through a SEI (Supplemental enhancement information) message.
[0012] According to embodiments, the texture map bitstream is composed of one or more data units, each data unit is composed of a data unit header including type information and a data unit payload carrying texture map data, and the texture map restoration related information can be carried through the data unit header.
[0013] According to embodiments, a decoding device may include a memory and at least one processor connected to the memory, wherein the at least one processor may be configured to receive a bitstream including mesh data and decode the mesh data.
[0014] According to embodiments, the at least one processor may include a base mesh processing unit for restoring a base mesh from a base mesh bitstream, a displacement information processing unit for restoring displacement information from a displacement vector bitstream, a texture map processing unit for restoring texture maps from a texture map bitstream, and an omitted texture map restoration unit for determining whether at least one texture map is omitted from the texture map bitstream and, if it is determined that there is at least one omitted texture map, restoring the at least one omitted texture map based on texture map restoration-related information and at least one reference frame.
[0015] According to embodiments, the texture map restoration related information may include information for identifying whether at least one texture map is omitted and reference texture map information related to the at least one reference frame.
[0016] According to embodiments, the texture map restoration related information may be carried through an SEI message.
[0017] According to embodiments, the texture map bitstream is composed of one or more data units, each data unit is composed of a data unit header including type information and a data unit payload carrying texture map data, and the texture map restoration related information can be carried through the data unit header.
[0018] According to embodiments, the encoding method may include a step of encoding mesh data and a step of transmitting a bitstream including the encoded mesh data.
[0019] According to embodiments, the encoding step may include a step of encoding a base mesh generated by simplifying an original mesh to generate a base mesh bitstream, a step of encoding displacement information generated based on the base mesh to generate a displacement vector bitstream, a step of restoring a mesh based on the encoded base mesh and the encoded displacement information, a step of generating texture maps based on the original mesh and the restored mesh, and determining whether to omit at least one texture map among the generated texture maps, a step of encoding remaining texture maps excluding at least one texture map for which omission has been determined to be determined to be omitted to generate a texture map bitstream, and a step of generating texture map restoration-related information for restoration of the at least one texture map for which omission has been determined.
[0020] According to embodiments, the texture map restoration related information may include information for identifying whether at least one texture map is omitted and reference texture map information related to the at least one reference frame.
[0021] According to embodiments, the texture map restoration related information may be carried through an SEI message.
[0022] According to embodiments, the texture map bitstream is composed of one or more data units, each data unit is composed of a data unit header including type information and a data unit payload carrying texture map data, and the texture map restoration related information can be carried through the data unit header.
[0023] According to embodiments, a computer program stored on a computer-readable recording medium can be combined with a computer, which is hardware, to perform the above method.
[0024] According to embodiments, a transmission method may include a step of obtaining a bitstream for image information, wherein the bitstream is generated based on a step of encoding mesh data and a step of transmitting a bitstream including the encoded mesh data, and a step of transmitting data including the bitstream.
[0025] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide a quality 3D service.
[0026] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can achieve various video codec methods.
[0027] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can provide general-purpose 3D content such as autonomous driving services.
[0028] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can obtain the effects of reducing bits of the transmitted texture map, improving transmission speed, and reducing the complexity of the encoder / decoder by omitting coding and transmission of the current texture map when the similarity between the reference texture map of the dynamic mesh and the current texture map is high.
[0029] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can obtain a significant bit saving effect by omitting coding and transmission of texture maps, which occupy the largest data proportion among mesh data components, on a frame-by-frame basis without a significant difference in quality.
[0030] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments can perform efficient compression of a texture map by signaling texture map restoration-related information necessary for restoration of a texture map whose coding has been omitted on the transmission side through NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI of V3C_AD, thereby obtaining a great effect in terms of bit savings and transmission speed.
[0031] The mesh data transmission method, mesh data transmission device, mesh data reception method, and mesh data reception device according to the embodiments signal texture map restoration related information using the existing V3C unit type V3C_AVD or by creating a new V3C unit type called V3C_TVD and signaling texture map restoration related information through it, thereby omitting texture maps on a frame-by-frame basis to achieve a significant bit-saving effect, and also restoring better quality data by restoring the omitted texture maps based on a weight basis. In addition, by proposing a method of transmitting texture map restoration related information (i.e., syntax) that is suitable for the V3C bitstream structure, a great effect can be obtained in terms of efficient compression and transmission speed.
[0032] The drawings are included to further understand the embodiments, and the drawings illustrate the embodiments together with the description related to the embodiments. For a better understanding of the various embodiments described below, reference should be made to the following description of the embodiments in conjunction with the following drawings, in which like reference numerals correspond to corresponding parts throughout the drawings.
[0033] FIG. 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0034] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0035] Figure 3 illustrates pre-processing of V-MESH compression according to embodiments.
[0036] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0037] Figure 5 illustrates a displacement generation process according to embodiments.
[0038] Figure 6 illustrates an intra-frame encoding process of V-MESH data according to embodiments.
[0039] Figure 7 illustrates an inter-frame encoding process of V-MESH data according to embodiments.
[0040] Figure 8 illustrates a lifting conversion process for displacement according to embodiments.
[0041] Figure 9 illustrates a process of packing transformation coefficients into a 2D image according to embodiments.
[0042] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.
[0043] Figure 11 illustrates an intra-frame decoding process of V-MESH data according to embodiments.
[0044] Figure 12 shows an inter-frame decoding processor of V-MESH data.
[0045] Fig. 13 is a drawing showing an example of a transmitting device according to embodiments.
[0046] Fig. 14 is a drawing showing an example of a receiving device according to embodiments.
[0047] FIG. 15 is a drawing showing another example of a transmitting device according to embodiments.
[0048] FIG. 16 is a diagram showing an example of an encoding process when texture map coding is omitted according to embodiments.
[0049] Fig. 17 is a drawing showing another example of a receiving device according to embodiments.
[0050] Fig. 18 is a drawing showing another example of a receiving device according to embodiments.
[0051] FIG. 19 is a diagram showing an example of a texture map decoding process in the case of omission of texture map coding according to embodiments.
[0052] FIG. 20(a) and FIG. 20(b) are diagrams showing examples of a reconstruction process and syntax mapping according to the texture map coding omission process of the present disclosure.
[0053] FIG. 21 is a drawing showing an example of a V3C bitstream structure according to embodiments, and is an example of a V3C bitstream in a V3C unit stream format.
[0054] Figure 22 is a table defining NAL types and conformance types for each SEI message purpose according to embodiments.
[0055] FIG. 23 is a diagram showing an example of NAL unit semantics according to embodiments.
[0056] FIG. 24 is a diagram showing an example of a syntax structure of sei_rbsp() including an SEI message according to embodiments.
[0057] FIG. 25 is a diagram showing an example of the syntax structure of sei_message() according to embodiments.
[0058] FIG. 26 is a diagram showing an example of the syntax structure of an SEI message payload (sei_payload(payloadType, payloadSize)) according to embodiments.
[0059] FIG. 27 is a diagram showing an example of the syntax structure of skipped_frame_indication(payloadSize) according to embodiments.
[0060] FIG. 28 is a table showing examples of types of V3C units assigned to the vuh_unit_type field in the V3C unit header according to embodiments.
[0061] Figure 29 shows an example of the syntax structure of a V3C unit header according to embodiments.
[0062] Figure 30 shows another example of the syntax structure of a V3C unit header according to embodiments.
[0063] Figure 31 is a flowchart showing an example of a transmission method according to embodiments.
[0064] Figure 32 is a flowchart showing an example of a receiving method according to embodiments.
[0065] Preferred embodiments of the embodiments are described in detail, examples of which are illustrated in the accompanying drawings. The following detailed description, with reference to the accompanying drawings, is intended to illustrate preferred embodiments of the embodiments, rather than merely show embodiments that can be implemented according to the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without these details.
[0066] While most of the terms used in the examples are commonly used in the field, some terms were arbitrarily selected by the applicant, and their meanings are described in detail in the following descriptions as needed. Therefore, the examples should be understood based on the intended meaning of the terms, not simply their names or meanings.
[0067] Recently, with the development of 3D data modeling and rendering technology, research on creating and processing 3D data is being conducted in various fields such as Virtual Reality (VR), Augmented Reality (AR), autonomous driving, Computer-Aided Design (CAD) / Computer-Aided Manufacturing (CAM), and Geographic Information System (GIS). 3D data can be represented as a point cloud, mesh, etc., depending on the representation format. Among these, a mesh is composed of geometric information expressing the coordinate values of each vertex (or point), connection information indicating the connection relationship between vertices, a texture map expressing the color information of the mesh surface as 2D image data, and texture coordinates indicating mapping information between the surface of the mesh and the texture map. In the present disclosure, a mesh is defined as a dynamic mesh if one or more of the elements constituting the mesh change over time, and a static mesh if they do not change.
[0068] Because dynamic mesh data has a large amount of data for elements that constitute the mesh compared to two-dimensional image data, technologies have been developed to efficiently compress this large amount of mesh data to store and transmit it.
[0069] FIG. 1 illustrates a system for providing dynamic mesh content according to embodiments.
[0070] The system of FIG. 1 includes a transmitting device (100) and a receiving device (110) according to embodiments. The transmitting device (100) may include a mesh video acquisition unit (101), a mesh video encoder (102), a file / segment encapsulator (103), and a transmitter (104). The receiving device (110) may include a receiving unit (111), a file / segment decapsulator (112), a mesh video decoder (113), and a renderer (114). Each component of FIG. 1 may correspond to hardware, software, a processor, and / or a combination thereof. Hereinafter, the mesh data transmitting device according to embodiments may be interpreted as a term referring to a 3D data transmitting device or transmitting device (100), or a mesh video encoder (hereinafter, referred to as an encoder) (102). The mesh data receiving device according to the embodiments may be interpreted as a term referring to a 3D data receiving device or receiving device (110), or a mesh video decoder (hereinafter, decoder) (113).
[0071] The system of FIG. 1 can perform video-based dynamic mesh compression and decompression.
[0072] Advances in 3D capture, modeling, and rendering have enabled users to consume diverse forms of 3D content, such as AR, XR, metaverse, and holograms, across multiple platforms and devices. 3D content increasingly represents objects with greater precision and realism, enabling users to enjoy immersive experiences. To achieve this, the creation and use of 3D models requires a significant amount of data. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. Embodiments include a series of processing steps in a system that utilizes such mesh content.
[0073] First, the method of compressing dynamic mesh data starts from the V-PCC (Video-based point cloud compression) standard technology for point cloud data. Point cloud data is data that has color information at the coordinates (X, Y, Z) of a vertex (or point). In the present disclosure, the coordinates (i.e., position information) of a vertex are referred to as geometry information, the color information of a vertex is referred to as attribute information, and the geometry information and attribute information are referred to as vertex information or point cloud data. The vertex information to which connectivity information between vertices is added is referred to as mesh data. When creating content, it can be created in the form of mesh data from the beginning. Alternatively, it can be used by converting it into mesh data by adding connectivity information to point cloud data.
[0074] Currently, the MPEG standards body defines two types of dynamic mesh data: Category 1: Mesh data with texture maps as color information. Category 2: Mesh data with vertex colors as color information.
[0075] Mesh coding standards for Category 1 data are currently under development, and work on Category 2 data standards is also planned for the future. The overall process for providing mesh content services may include acquisition, encoding, transmission, decoding, rendering, and / or feedback, as shown in Figure 1.
[0076] To provide mesh content services, 3D data acquired through multiple cameras or specialized cameras can be processed into mesh data types through a series of processes and then converted into video. The generated mesh video is then transmitted through a series of processes, and the receiving end can then reprocess the received data into mesh video and render it. This allows mesh video to be presented to users, who can then interact with the mesh content according to their intended intent.
[0077] A mesh compression system may include a transmitting device (100) and a receiving device (110) as shown in FIG. 1. The transmitting device (100) may encode mesh video to output a bitstream, and transmit the bitstream to the receiving device (110) in the form of a file or streaming (streaming segment) via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0078] In the above transmitting device (100), the encoder may be called a mesh video / video / picture / frame encoding device, and in the receiving device (110), the decoder may be called a mesh video / video / picture / frame decoding device. The transmitter may be included in a mesh video encoder. The receiver may be included in a mesh video decoder. The renderer (114) may include a display unit, and the renderer and / or the display unit may be configured as separate devices or external components. The transmitting device (100) and the receiving device (110) may further include separate internal or external modules / units / components for a feedback process.
[0079] Mesh data represents the surface of an object as a number of polygons. Each polygon is defined by vertices in 3D space and connection information that describes how the vertices are connected. It can also contain vertex attributes such as vertex color and normal. Mapping information that allows the surface of the mesh to be mapped to a 2D planar area can also be included in the attributes of the mesh. The mapping can be described as a set of parameter coordinates, generally called UV coordinates or texture coordinates, associated with the mesh vertices. Meshes contain 2D attribute maps, which can be used to store high-resolution attribute information such as textures, normals, and displacement. Here, displacement can be used interchangeably with displacement, displacement information, or displacement vectors (i.e., displacement vectors).
[0080] The mesh video acquisition unit (101) may include processing 3D object data acquired through a camera, etc. into a mesh data type having the attributes described above through a series of processes and generating a video composed of such mesh data. The mesh video may have attributes of the mesh, such as vertices, polygons, connection information between vertices, colors, normals, etc., that may change over time. A mesh video having attributes and connection information that change over time in this way may be expressed as a dynamic mesh video.
[0081] A mesh video encoder (102) can encode an input mesh video into one or more video streams. One video can include multiple frames, and one frame can correspond to a still image / picture. In this document, a mesh video can include a mesh image / frame / picture, and a mesh video can be used interchangeably with a mesh image / frame / picture. The mesh video encoder (102) can perform a Video-based Dynamic Mesh (V-Mesh) Compression procedure. The mesh video encoder (102) can perform a series of procedures such as prediction, transformation, quantization, and entropy coding for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0082] The file / segment encapsulator (103) can encapsulate encoded mesh video data and / or mesh video-related metadata in the form of a file, etc. Here, the mesh video-related metadata may be received from a metadata processing unit, etc. The metadata processing unit may be included in the mesh video encoder (102) or may be configured as a separate component / module. The file / segment encapsulator (103) can encapsulate the corresponding data in a file format such as ISOBMFF, or process it in the form of other DASH segments, etc. The file / segment encapsulator (103) may include mesh video-related metadata in the file format according to an embodiment. The mesh video metadata may be included in boxes at various levels in the ISOBMFF file format, for example, or may be included as data in a separate track within the file. Depending on the embodiment, the file / segment encapsulator (103) may encapsulate the mesh video related metadata itself into a file.
[0083] The transmission processing unit can process encapsulated mesh video data for transmission according to the file format. The transmission processing unit can be included in the transmission unit (104) or can be configured as a separate component / module. The transmission processing unit can process mesh video data according to any transmission protocol. The processing for transmission can include processing for transmission through a broadcast network or processing for transmission through broadband. According to an embodiment, the transmission processing unit can receive not only mesh video data but also mesh video-related metadata from the metadata processing unit and process it for transmission.
[0084] The transmission unit (104) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (111) of the reception device (110) via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (104) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (111) can extract the bitstream and transmit it to a decoding device.
[0085] The receiving unit (111) can receive mesh video data transmitted by a mesh data transmission device. Depending on the channel through which it is transmitted, the receiving unit (111) can receive mesh video data through a broadcast network, through a broadband, or through a digital storage medium.
[0086] The receiving processing unit can perform processing according to the transmission protocol on the received mesh video data. The receiving processing unit can be included in the receiving unit (111) or can be configured as a separate component / module. In order to correspond to the processing performed for transmission on the transmitting side, the receiving processing unit can perform the reverse process of the aforementioned transmission processing unit. The receiving processing unit can transfer the acquired mesh video data to the file / segment decapsulator (112) and transfer the acquired mesh video-related metadata to the metadata parser. The mesh video-related metadata acquired by the receiving processing unit can be in the form of a signaling table.
[0087] The file / segment decapsulator (112) can decapsulate mesh video data in the form of a file received from a receiving processing unit. The file / segment decapsulator (112) can decapsulate files according to ISOBMFF, etc., to obtain a mesh video bitstream or mesh video-related metadata (metadata bitstream). The obtained mesh video bitstream can be transmitted to the mesh video decoder (113), and the obtained mesh video-related metadata (metadata bitstream) can be transmitted to the metadata processing unit. The mesh video bitstream may include metadata (metadata bitstream). The metadata processing unit may be included in the mesh video decoder (113) or may be configured as a separate component / module. The mesh video-related metadata obtained by the file / segment decapsulator (112) may be in the form of a box or track within a file format. The file / segment decapsulator (112) may receive metadata required for decapsulation from the metadata processing unit, if necessary. The mesh video related metadata may be passed to the mesh video decoder (113) and used in the mesh video decoding procedure, or may be passed to the renderer (114) and used in the mesh video rendering procedure.
[0088] The mesh video decoder (113) can receive a bitstream and perform a reverse operation corresponding to the operation of the mesh video encoder (102) to decode the video / image. The decoded mesh video / image can be displayed through the display unit of the renderer (114). The user can view all or part of the rendered result through a VR / AR display or a general display.
[0089] The feedback process may include a process of transmitting various feedback information that may be acquired during the rendering / display process to the transmitter or to the decoder on the receiver. Interactivity may be provided in mesh video consumption through the feedback process. Depending on the embodiment, head orientation information, viewport information indicating the area that the user is currently viewing, etc. may be transmitted during the feedback process. Depending on the embodiment, the user may interact with things implemented in the VR / AR / MR / autonomous driving environment, in which case information related to the interaction may be transmitted to the transmitter or the service provider during the feedback process. Depending on the embodiment, the feedback process may not be performed.
[0090] Head orientation information can refer to information about the user's head position, angle, and movement. Based on this information, information about the area the user is currently viewing within the mesh video, i.e. viewport information, can be calculated.
[0091] Viewport information can be information about the area the user is currently viewing in the mesh video. This can be used to perform gaze analysis to determine how the user consumes the mesh video, which area of the mesh video they are gazing at, and for how long. Gaze analysis can be performed on the receiving side and transmitted to the transmitting side through a feedback channel. Devices such as VR / AR / MR displays can extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0092] Depending on the embodiment, the aforementioned feedback information may not only be transmitted to the transmitter but may also be consumed by the receiver. That is, the aforementioned feedback information may be utilized to perform decoding, rendering, and other processes on the receiver. For example, head orientation information and / or viewport information may be utilized to preferentially decode and render only the mesh video for the area currently being viewed by the user.
[0093] This document relates to embodiments of dynamic mesh video compression as described above. The method / embodiment disclosed in this document can be applied to the Video-based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or the next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connection information and attributes that change over time, and it can perform lossy and lossless compression for various applications such as real-time communication, storage, free-viewpoint video, and AR / VR.
[0094] The dynamic mesh video compression method described below is based on MPEG's V-Mesh method.
[0095] In this document, picture / frame can generally mean a unit representing one video of a specific time period.
[0096] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, only the pixel / pixel value of the chroma component, or only the pixel / pixel value of the depth component.
[0097] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0098] As described above, the encoding process of Fig. 1 is as follows.
[0099] That is, the video-based dynamic mesh compression (V-Mesh) compression method can provide a method of compressing dynamic mesh video data based on 2D video codecs such as HEVC (High Efficiency Video Coding) and VVC (Versatile Video Coding). The V-Mesh compression process receives the following data as input and performs compression.
[0100] Input mesh: Contains the 3D coordinates of the vertices that make up the mesh, normal information for each vertex, mapping information that maps the mesh surface to a 2D plane, and connection information between the vertices that make up the surface. The mesh surface can be expressed as triangles or more polygons, and connection information between the vertices that make up each surface is stored according to a set shape. The input mesh can be saved in the OBJ file format.
[0101] Attribute map: (Hereinafter, texture map is also used in the same meaning): Contains information about the attributes of the mesh (color, normal, displacement, etc.), and stores data in the form of a mapping of the surface of the mesh onto a 2D image. Mapping which part (surface or vertex) of the mesh each data of this attribute map corresponds to is based on the mapping information contained in the input mesh. Since the attribute map has data for each frame of the mesh video, it can also be expressed as an attribute map video. The attribute map in the V-Mesh compression method mainly contains the color information of the mesh, and is saved in an image file format (PNG, BMP, etc.).
[0102] Material Library File: Contains information about the material attributes used in a mesh, and in particular, information that links the input mesh to its corresponding attribute map. It is saved in the Wavefront Material Template Library (MTL) file format.
[0103] In the V-Mesh compression method, the following data and information can be generated through the compression process.
[0104] Base mesh: The input mesh is simplified (decimated) through a pre-processing process, thereby expressing the objects of the input mesh using the minimum number of vertices determined by the user's standards.
[0105] Displacement: Displacement information used to express the input mesh as similarly as possible to the base mesh, and is expressed in the form of 3D coordinates.
[0106] Atlas information: This is the metadata required to reconstruct a mesh using base mesh, displacement, and attribute map information. It can be created and utilized as sub-mesh units (such as patches) that make up the mesh.
[0107] Referring to FIGS. 2 to 7, a method for encoding mesh position information (or vertex position information) is described, and referring to FIGS. 6 to 10, etc., a method for encoding attribute information (attribute map) by restoring mesh position information is described.
[0108] Figure 2 illustrates a V-MESH compression method according to embodiments.
[0109] Fig. 2 illustrates the encoding process of Fig. 1, and the encoding process may include a pre-processing process and an encoding process. The mesh video encoder (102) of Fig. 1 may include a pre-processor (200) and an encoder (201) as in Fig. 2. In addition, the transmitting device of Fig. 1 may be broadly referred to as an encoder, and the mesh video encoder (102) of Fig. 1 may be referred to as an encoder. The V-Mesh compression method may include a pre-processing process (Pre-processing, 200) and an encoding process (Encoding, 201) as in Fig. 2. The pre-processor (200) of Fig. 2 may be located in front of the encoder (201) of Fig. 2. The pre-processor (200) and the encoder (201) of Fig. 2 may be referred to as a single encoder.
[0110] The pre-processor (200) can receive a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). The pre-processor (200) can generate a base mesh (m(i)) and / or a displacement (or displacement) (d(i)) through pre-processing. The pre-processor (200) can receive feedback information from the encoder (201) and generate the base mesh and / or the displacement based on the feedback information.
[0111] The encoder (201) can receive a base mesh (m(i)), a displacement (d(i)), a static of a dynamic mesh (M(i)), and / or an attribute map (A(i)). In the present disclosure, at least one of the base mesh (m(i)), the displacement (d(i)), the static of a dynamic mesh (M(i)), and / or the attribute map (A(i)) can be referred to as mesh-related data. The encoder (201) can encode the mesh-related data to generate a compressed bitstream.
[0112] Figure 3 illustrates a pre-processing process of V-MESH compression according to embodiments.
[0113] Fig. 3 illustrates the configuration and operation of the preprocessor of Fig. 2. In Fig. 3, the input mesh may include a static of a dynamic mesh (M(i)) and / or an attribute map (A(i)). In addition, the input mesh may include three-dimensional coordinates of vertices constituting the mesh, normal information of each vertex, mapping information for mapping the mesh surface to a 2D plane, connection information between vertices constituting the surface, etc.
[0114] Fig. 3 shows a process of performing pre-processing on an input mesh. The pre-processing process (200) may largely include four steps: 1) GoF (Group of Frame) generation, 2) Mesh Decimation, 3) UV parameterization, and 4) Fitting subdivision surface (300). According to embodiments, GoF generation may be referred to as a GoF generation process or a GoF generation unit, mesh simplification may be referred to as a mesh simplification process or a mesh simplification unit, UV parameterization may be referred to as a UV parameterization process or a UV parameterization unit, and the fitting subdivision surface may be referred to as a fitting subdivision surface process or a fitting subdivision surface unit. The pre-processor (200) can generate displacement and / or base meshes from the received input mesh and transmit them to the encoder (201). The pre-processor (200) can transmit GoF information associated with GoF generation to the encoder (201).
[0115] Below, each step of Fig. 3 is described.
[0116] GoF Generation: This is the process of generating a reference structure for mesh data. If the number of vertices, the number of texture coordinates, the vertex connection information, and the texture coordinate connection information of the mesh of the previous frame and the current mesh are all the same, the previous frame can be set as the reference frame. That is, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, the encoder (201) can perform inter frame encoding. Otherwise, intra frame encoding is performed for the corresponding frame.
[0117] Mesh Decimation: This process simplifies the input mesh to create a simplified mesh, or base mesh. Vertices to be removed from the original mesh are selected based on user-defined criteria, and the selected vertices and the triangles connected to them can be removed.
[0118] In the process of performing mesh simplification (Mesh decimation), the input mesh (voxelized), target triangle ratio (TTR), and minimum triangle component (CCCount) information are passed as input, and the simplified mesh (decimated mesh) can be obtained as output. In this process, connected triangle components smaller than the set minimum triangle component (CCCount) can be removed.
[0119] UV parameterization: This is the process of mapping a 3D surface of a decimated mesh into a texture domain. Parameterization can be performed using the UVAtlas tool. This process generates mapping information, which indicates where each vertex of the decimated mesh can be mapped to on a 2D image. This mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is created.
[0120] Fitting subdivision surface (300): This is a process of performing subdivision on a decimated mesh (i.e., a simplified mesh having texture coordinates). The displacement and base mesh generated through this process are output to the encoder (201). A user-defined method, such as a mid-edge method, may be applied as the subdivision method. A fitting process is performed so that the input mesh and the mesh on which the subdivision has been performed become similar to each other. In the present disclosure, the mesh on which the fitting process has been performed is referred to as a fitted subdivision mesh (or fitted subdivision mesh).
[0121] Figure 4 illustrates a mid-edge subdivision method according to embodiments.
[0122] Figure 4 illustrates the mid-edge method of the fitting subdivision surface described in Figure 3. Referring to Figure 4, an original mesh containing four vertices is subdivided to generate a sub-mesh. A sub-mesh can be generated by creating a new vertex in the middle of the edge between the vertices. Then, a fitting process is performed so that the input mesh and the sub-mesh become similar to each other, thereby generating a fitted sub-division mesh.
[0123] When a fitted subdivided mesh (hereinafter referred to as a fitted subdivided mesh) is generated, displacement is calculated using this result and a pre-compressed and decoded base mesh (hereinafter referred to as a reconstructed base mesh). That is, the reconstructed base mesh is subdivided in the same way as the fitting subdivision surface. The difference in position of each vertex between this result and the fitted subdivided mesh is the displacement for each vertex. Since displacement represents the position difference in three-dimensional space, it is also expressed as a value in the (x, y, z) space of a Cartesian coordinate system. Depending on the user input parameters, the (x, y, z) coordinate values can be converted to (normal, tangential, bi-tangential) coordinate values of the local coordinate system.
[0124] Fig. 5 illustrates a displacement generation process according to embodiments. The displacement generation process of Fig. 5 may be performed in a pre-processor (200) or in an encoder (201).
[0125] FIG. 5 illustrates in detail the displacement calculation method of the fitting subdivision surface (300) as described in FIG. 4.
[0126] An encoder and / or pre-processor according to embodiments may include 1) a subdivision unit, 2) a local coordinate system calculation unit, and 3) a displacement calculation unit. The subdivision unit may perform subdivision on a reconstructed base mesh to generate a subdivided reconstructed base mesh. Here, the restoration of the base mesh may be performed in the pre-processor (200) or in the encoder (201). The local coordinate system calculation unit may receive a fitted subdivision mesh and a subdivided reconstructed base mesh, and may convert a coordinate system of the mesh into a local coordinate system based on the fitted subdivision mesh and the subdivided reconstructed base mesh. The local coordinate system calculation operation may be optional. The displacement calculation unit may calculate a positional difference between the fitted subdivision mesh and the subdivided reconstructed base mesh. For example, a positional difference value between vertices of two input meshes may be generated. The vertex positional difference value becomes a displacement.
[0127] The mesh data transmission method and device according to the embodiments can encode mesh data as follows. Mesh data is a term including point cloud data. Point cloud data (which may be referred to as point cloud for short) according to the embodiments can refer to data including vertex coordinates (or geometry information) and color information (or attribute information). In addition, geometry images, attribute images, occupancy maps, and additional information (or patch information) generated through patch generation and packing based on vertex coordinates and color information are also referred to as point cloud data. Therefore, point cloud data including connection information can be referred to as mesh data. In this document, point cloud and mesh data can be used interchangeably.
[0128] The V-Mesh compression (reconstruction) method according to the embodiments may include intra frame encoding (Fig. 6) and inter frame encoding (Fig. 7).
[0129] Based on the results of the GoF generation described above, intra-frame encoding or inter-frame encoding is performed. In the case of intra-encoding, the data to be compressed may be a base mesh, displacement, attribute map, etc. In the case of inter-encoding, the data to be compressed may be a displacement, attribute map, and a motion field between a reference base mesh and the current base mesh.
[0130] Fig. 6 illustrates an intra-frame encoding process of a V-MESH compression method according to embodiments. Each component for the intra-frame encoding process of Fig. 6 corresponds to hardware, software, a processor, and / or a combination thereof.
[0131] The encoding process of FIG. 6 details the encoding of the mesh video encoder (102) of FIG. 1. That is, it shows the configuration of the mesh video encoder (102) when the encoding of FIG. 1 is an intra-frame method. The encoder of FIG. 6 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of FIG. 6 may correspond to the pre-processor (200) and encoder (201) of FIG. 3.
[0132] The preprocessor (200) can receive an input mesh and perform the preprocessing described above. The preprocessing can generate a base mesh and / or a fitted subdivision mesh.
[0133] The quantizer (411) of the encoder (201) can quantize the base mesh and / or the fitted subdivided mesh. The static mesh encoder (412) can encode the static mesh (i.e., the quantized base mesh) and generate a bitstream (i.e., a compressed base mesh bitstream) including the encoded base mesh. The static mesh decoder (413) can decode the encoded static mesh (i.e., the encoded base mesh). The inverse quantizer (414) can inversely quantize the quantized static mesh (i.e., the base mesh) to output a reconstructed (or restored) base mesh. The displacement calculation unit (415) can generate displacements (or displacements) based on the reconstructed static mesh (i.e., the base mesh) and the fitted subdivided mesh. According to embodiments, the displacement calculation unit (415) calculates displacement, which is the position difference between each vertex of the subdivided base mesh and the fitted subdivided mesh after subdividing (or refining) the restored base mesh. In other words, the displacement is a displacement vector, which is the position difference between the vertices of the two meshes so that the fitted subdivided (or refining) mesh becomes similar to the original mesh. The forward linear lifting unit (416) can perform lifting transformation on the input displacement to generate lifting coefficients (or transform coefficients). The quantizer (417) can quantize the lifting coefficients. The image packing unit (418) can pack an image based on the quantized lifting coefficients. The video encoder (419) can encode the packed image. That is, the quantized lifting coefficients are packed into one frame as a 2D image by the image packing unit (418), compressed through the video encoder (419), and output as a displacement bitstream (i.e., compressed displacement bitstream).
[0134] A video decoder (420) decodes a compressed displacement bitstream. An image unpacking unit (421) can perform unpacking on the decoded displacement frame to output quantized lifting coefficients. A dequantizer (422) can dequantize the quantized lifting coefficients. An inverse linear lifting unit (423) applies inverse lifting to the inverse quantized lifting coefficients to generate restored displacement. A mesh restoration unit (424) reconstructs and deforms a mesh using the restored displacement output from the inverse linear lifting unit (423) and the restored base mesh (or subdivided restored base mesh) output from the inverse quantization unit (414). The present disclosure refers to the reconstructed and deformed mesh as a restored deformed mesh.
[0135] The attribute transfer (425) receives an input mesh and / or an input attribute map, and regenerates an attribute map based on the restored deformed mesh. The attribute map refers to a texture map corresponding to attribute information among mesh data components, and in the present disclosure, the attribute map and the texture map may be used interchangeably. The push-pull padding (426) may pad data in the attribute map based on the push-pull method. The color space conversion unit (427) may convert the space of the color component of the attribute map. For example, the attribute map may be converted from an RGB color space to a YUV color space. The video encoder (428) may encode the attribute map and output it as a compressed attribute bitstream.
[0136] A multiplexer (430) can generate a compressed bitstream by multiplexing a compressed base mesh bitstream, a compressed displacement bitstream, and a compressed attribute bitstream.
[0137] In Fig. 6, the displacement calculation unit (415) may be included in the pre-processor (200). In addition, at least one of the quantizer (411), the static mesh encoder (412), the static mesh decoder (413), and the inverse quantizer (414) may be included in the pre-processor (200).
[0138] As described in FIG. 6, the intra-frame encoding method includes base mesh encoding (also called static mesh encoding). That is, when performing intra-frame encoding on the current input mesh frame, the base mesh generated in the pre-processing process of the pre-processor (200) can be encoded using a static mesh compression technology in a static mesh encoder (412) after undergoing a quantization process in a quantizer (411). In the V-Mesh compression method, for example, Draco technology is applied to base mesh encoding, and vertex position information, mapping information (texture coordinates), vertex connection information, etc. of the base mesh become compression targets.
[0139] The encoder of Fig. 6 generates a bitstream by compressing the base mesh, displacement, and attributes within the frame, and the encoder of Fig. 7 generates a bitstream by compressing the motion, displacement, and attributes between the current frame and the reference frame.
[0140] Fig. 7 illustrates an inter-frame encoding process of a V-MESH compression method according to embodiments. Each component for the inter-frame encoding process of Fig. 7 corresponds to hardware, software, a processor, and / or a combination thereof.
[0141] The encoding process of Fig. 7 details the encoding of Fig. 1. That is, it shows the configuration of an encoder when the encoding of Fig. 1 is an inter-frame method. The encoder of Fig. 7 may include a pre-processor (200) and / or an encoder (201). The pre-processor (200) and encoder (201) of Fig. 7 may correspond to the pre-processor (200) and encoder (201) of Fig. 3.
[0142] For a description of the components corresponding to the encoding operation of FIG. 6 among the encoding operations of FIG. 7, refer to the description of FIG. 6. That is, the operation of the quantizer (511), displacement calculation unit (515), wavelet transformer (516), quantizer (517), image packing unit (518), video encoder (519), video decoder (520), image unpacking unit (521), inverse quantizer (522), inverse wavelet transformer (523), mesh restoration unit (524), attribute transfer (525), push-pull padding (526), color space conversion unit (527), video encoder (528), and multiplexer (530) of FIG. 7 is similar to that of the quantizer (411), static mesh encoder (412), static mesh decoder (413), inverse quantizer (414), displacement calculation unit (415), forward linear lifting unit (416), quantizer (417), image Since the operations described in the packing unit (418), video encoder (419), video decoder (420), image unpacking unit (421), inverse quantizer (422), inverse linear lifting unit (423), mesh restoration unit (424), attribute transfer (425), push-pull padding (426), color space conversion unit (427), video encoder (428), and multiplexer (430) are the same or similar, a detailed description thereof is omitted in FIG. 7 to avoid redundant description.
[0143] In Fig. 7, for inter-frame based encoding, the motion encoder (512) can obtain a motion vector between the two base meshes based on the restored quantized reference base mesh and the quantized current base mesh, and then encode the motion vector to output a compressed motion bitstream. The motion encoder (512) can be referred to as a motion vector encoder. The base mesh restoration unit (513) can restore the base mesh based on the restored quantized reference base mesh and the encoded motion vector. The restored base mesh is dequantized in the dequantizer (514) and then output to the displacement calculation unit (515).
[0144] In Fig. 7, the displacement calculation unit (515) may be included in the pre-processor (200). In addition, at least one of the quantizer (511), the motion encoder (512), the base mesh restoration unit (513), and the inverse quantizer (514) may be included in the pre-processor (200).
[0145] As described in Fig. 7, the inter-frame encoding method may include motion field encoding (also called motion vector encoding). Inter-frame encoding may be performed when a one-to-one correspondence of vertices is established between a reference mesh and a current input mesh, and only the position information of the vertices is different. When performing inter-frame encoding, instead of compressing the base mesh, the difference between the vertices of the reference base mesh and the current base mesh, i.e., the motion field (also called motion vector), may be calculated and encoded to encode this information. The reference base mesh is the result of quantizing the already decoded base mesh data and is determined according to the reference frame index determined in the GoF generation. The motion field may also be encoded as a value. Alternatively, the predicted motion field can be calculated by averaging the motion fields of the restored vertices among the vertices connected to the current vertex, and the residual motion field, which is the difference between the predicted motion field value and the motion field value of the current vertex, can be encoded. This residual motion field value can be encoded using entropy coding.The process of encoding displacement and attribute maps, excluding the motion field encoding process of inter frame encoding, is the same as the structure of the intra frame encoding method except for the base mesh encoding.
[0146] Figure 8 illustrates a lifting conversion process for displacement according to embodiments.
[0147] Figure 9 illustrates a process of packing transformation coefficients (or lifting coefficients) according to embodiments into a 2D image.
[0148] Figures 8 and 9 illustrate the process of transforming displacement and packing transform coefficients of the encoding process of Figures 6 and 7, respectively.
[0149] The encoding method according to the embodiments includes displacement encoding.
[0150] After base mesh encoding and / or motion field encoding, a reconstructed base mesh is generated through restoration and dequantization, and the displacement between the result of performing subdivision on the reconstructed base mesh and the fitted subdivided mesh generated through the fitting subdivision surface can be calculated (see 415 in FIG. 6 or 515 in FIG. 7). For effective encoding, a data transform process such as wavelet transform can be applied to the displacement information (see 416 in FIG. 6 or 516 in FIG. 7).
[0151] FIG. 8 shows a process of transforming displacement information using a lifting transform in the forward linear lifting unit (416) of FIG. 6 or the wavelet transformer (516) of FIG. 7. For example, a linear wavelet-based lifting transform may be performed. The transform coefficients generated through the transform process are quantized in a quantizer (417 or 517) and then packed into a 2D image through an image packing unit (418 or 518) as in FIG. 9. The transform coefficients are configured as one block for every 256 (= 16×16) units, and each block can be packed in a z-scan order. The horizontal number of blocks is fixed to 16, but the vertical number of blocks can be determined according to the number of vertices of the subdivided base mesh. Transform coefficients can be packed by aligning them with Morton codes within a single block. The packed images generate displacement videos for each GoF unit, and these displacement videos can be encoded using a conventional video compression codec in a video encoder (419 or 519).
[0152] Referring to FIG. 8, the base mesh (original) may include vertices and edges for LoD0. A first subdivision mesh generated by dividing (or subdividing) the base mesh includes vertices generated by further dividing (or subdividing) edges of the base mesh. The first subdivision mesh includes vertices for LoD0 and vertices for LoD1. LoD1 includes the subdivided vertices and the vertices of the base mesh (LoD0). The first subdivision mesh may be further divided (or subdivided) to generate a second subdivision mesh. The second subdivision mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 including vertices further divided (or subdivided) from LoD0, and vertices further divided (or subdivided) from LoD1. LoD is a level of detail that indicates the degree of detail of mesh data content. As the level index increases, the distance between vertices becomes closer and the level of detail increases. In other words, the smaller the LoD value, the lower the detail of the mesh data content, and the larger the LoD value, the higher the detail of the mesh data content. LoD N contains the vertices included in the previous LoDN-1 as is. When a mesh (or vertex) is further divided through subdivision, the mesh can be encoded based on a prediction and / or update method by considering the previous vertices v1, v2 and the subdivided vertex v. Instead of directly encoding information about the current LoD N, a residual value between the previous LoD N-1 can be generated and the mesh can be encoded using the residual value to reduce the size of the bitstream. The prediction process means the operation of predicting the current vertex v using the previous vertices v1 and v2. Since adjacent subdivision meshes have similar data, this property can be utilized for efficient encoding.Current vertex position information is predicted as a residual of previous vertex position information, and the previous vertex position information is updated through the residual. In the present disclosure, vertex, apex, and point may be used with the same meaning. In addition, LoDs may be defined during the subdivision process of the base mesh. According to embodiments, the subdivision process of the base mesh may be performed in the pre-processor (200) or in a separate component / module.
[0153] Referring to FIG. 9, a vertex has a transform coefficient (also called a lifting coefficient) generated through a lifting transformation. The transform coefficient of a vertex related to a lifting transformation can be packed into an image by an image packing unit (418 or 518) and then encoded by a video encoder (419 or 519).
[0154] Figure 10 illustrates an attribute transfer process of a V-MESH compression method according to embodiments.
[0155] According to the embodiments, FIG. 10 shows the detailed operation of the attribute transfer (425 or 525) of the encoding of FIG. 6, FIG. 7, etc.
[0156] Encoding according to embodiments includes attribute map encoding. According to embodiments, attribute map encoding may be performed in the video encoder (428) of FIG. 6 or the video encoder (528) of FIG. 7.
[0157] According to embodiments, in the present disclosure, the encoder compresses information about the input mesh through base mesh encoding (i.e., intra encoding), motion field encoding (i.e., inter encoding), and displacement encoding. In the encoding process, the compressed input mesh is restored through base mesh decoding (intra frame), motion field decoding (inter frame), and displacement video decoding processes, and the restored result, the reconstructed deformed mesh (hereinafter referred to as Recon. deformed mesh), is used to compress the input attribute map as shown in FIGS. 6 and 7. The reconstructed deformed mesh (Recon. deformed mesh) has position information of vertices, texture coordinates, and corresponding connection information, but does not have color information corresponding to the texture coordinates. Therefore, as shown in Fig. 10, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of attribute transfer (425 or 525).
[0158] According to embodiments, attribute transfer (425 or 525) first checks whether each point P(u, v) of a 2D texture domain belongs to a texture triangle of a reconstructed deformed mesh, and if it exists in a texture triangle T, the barycentric coordinate of P(u, v) according to the triangle T ( , , ) is calculated. And the 3D vertex positions of triangle T and ( , , ) is used to compute the 3D coordinates M(x, y, z) of P(u, v). Find the vertex coordinates M'(x', y', z') and the triangle T' containing this vertex that corresponds to the most similar position to the calculated M(x, y, z) in the input mesh domain. Then, the center of mass coordinates of M'(x', y', z') in this triangle T' ( ', ', ') is calculated. The texture coordinates corresponding to the three vertices of Triangle T' and ( ', ', ') is used to calculate the texture coordinates (u', v'), and the color information corresponding to these coordinates is found in the input attribute map. The color information found in this way is immediately assigned to the pixel location (u, v) of the new attribute map. If P(u, v) does not belong to any triangle, the pixel at that location in the new attribute map can be filled with a color value using a padding algorithm, such as the push-pull algorithm of push-pull padding (426 or 526).
[0159] The new attribute map generated through attribute transfer (425 or 525) is grouped into GoF units to form an attribute map video, which is compressed using the video codec of the video encoder (428 or 528).
[0160] Referring to Figure 10, the reference relationship between the input mesh, the input attribute map, the reconstructed deformed mesh, and the regenerated attribute map can be seen.
[0161] The decoding process of Fig. 1 can perform the reverse process of the corresponding process of the encoding process of Fig. 1. The specific decoding process is as follows.
[0162] FIG. 11 illustrates an intra-frame decoding (or intra-decoding) process of V-Mesh technology according to embodiments.
[0163] Fig. 11 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1. In addition, Fig. 11 can restore mesh data by performing the reverse process of the intra-frame encoding process of Fig. 6. Each component for the intra-frame decoding process of Fig. 11 corresponds to hardware, software, and / or a combination thereof.
[0164] First, the bitstream (i.e., compressed bitstream) received and input to the demultiplexer (611) of the intra frame decoding unit (610) can be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information of the mesh, such as V-PCC / V3C. The term V-PCC (Video-based Point Cloud Compression) used in this document can be used with the same meaning as V3C (Visual Volumetric Video-based Coding), and the two terms can be used interchangeably. Therefore, the term V-PCC in this document can be interpreted as the term V3C.
[0165] According to embodiments, the mesh sub-stream may be input to a static mesh decoder (612) and decoded, the displacement sub-stream may be input to a video decoder (613) and decoded, and the attribute map sub-stream may be input to a video decoder (617) and decoded.
[0166] According to embodiments, the mesh sub-stream is decoded through a decoder (612) of a static mesh codec used in encoding, such as Google Draco, and as a result, a reconstructed quantized base mesh, for example, connection information, vertex geometry information, vertex texture coordinates, etc. of the base mesh can be reconstructed.
[0167] According to embodiments, the displacement sub-stream is decoded into displacement video through a decoder (613) of a video compression codec used in encoding, and is restored as displacement information for each vertex (i.e., Recon. displacements) through an image unpacking process of an image unpacking unit (614), an inverse quantization process of an inverse quantizer (615), and an inverse transform process of an inverse linear lifting unit (616).
[0168] According to embodiments, the base mesh restored by the static mesh decoder (612) is inverse quantized by the inverse quantizer (620) and then output to the mesh restoration unit (630). The mesh restoration unit (630) reconstructs and restores the deformed mesh (i.e., decoded mesh) through the restored displacement output from the inverse linear lifting unit (616) and the restored base mesh output from the inverse quantizer (620). That is, the inverse quantized restored base mesh is combined with the restored displacement information to generate the final decoded mesh. In the present disclosure, the final decoded mesh is referred to as a reconstructed deformed mesh.
[0169] According to embodiments, an attribute map sub-stream is decoded through a decoder (617) corresponding to a video compression codec used in encoding, and then restored to a final attribute map (i.e., decoded attribute map) through a color conversion unit (640) through processes such as color format conversion and color space conversion.
[0170] According to embodiments, the restored decoded mesh and decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0171] Referring to FIG. 11, the received compressed bitstream includes patch information, a mesh substream, a displacement substream, and an attribute map substream. A substream is interpreted as a term referring to a part of a bitstream included in a bitstream. The bitstream includes patch information (data), mesh information (data), displacement information (data), and attribute map information (data).
[0172] As described above, the decoder of FIG. 11 performs the following intra-frame decoding operations. The static mesh decoder (612) decodes the mesh sub-stream to generate a reconstructed quantized base mesh, and the inverse quantizer (620) applies the quantization parameters of the quantizer inversely to generate the reconstructed base mesh. The video decoder (613) decodes the displacement sub-stream, the image unpacking unit (614) unpacks the images of the decoded displacement video, and the inverse quantizer (615) inversely quantizes the quantized images. The inverse linear lifting unit (616) applies a lifting transform in the reverse process of the encoder to generate the reconstructed displacement. The mesh restoration unit (630) generates a reconstructed deformed mesh based on the reconstructed base mesh and the reconstructed displacement. The video decoder (617) decodes the attribute map sub-stream, and the color conversion unit (640) converts the color format and / or space of the decoded attribute map to generate a decoded attribute map.
[0173] Figure 12 illustrates the inter-frame decoding (or inter-decoding) process of V-Mesh technology.
[0174] Fig. 12 illustrates the configuration and operation of the mesh video decoder (113) of the receiving device of Fig. 1. In addition, Fig. 12 can restore mesh data by performing the reverse process of the inter-frame encoding process of Fig. 7. Each component for the inter-frame decoding process of Fig. 12 corresponds to hardware, software, and / or a combination thereof.
[0175] First, the bitstream received and input to the demultiplexer (711) of the intra frame decoding unit (710) can be separated into a motion sub-stream (also called a motion sub-stream or motion vector sub-stream), a displacement sub-stream, an attribute map sub-stream, and a sub-stream including patch information of a mesh such as V3C / V-PCC.
[0176] According to embodiments, a motion sub-stream may be input to a motion decoder (712) and decoded, a displacement sub-stream may be input to a video decoder (713) and decoded, and an attribute map sub-stream may be input to a video decoder (717) and decoded.
[0177] According to embodiments, a motion sub-stream is decoded through entropy decoding and inverse prediction processes in a motion decoder (712) and restored into motion information (or motion vector information). A base mesh restoration unit (718) combines the restored motion information with a reference base mesh that has already been restored and stored to generate a reconstructed quantized base mesh for the current frame. An inverse quantizer (720) applies inverse quantization to the restored quantized base mesh to generate a reconstructed base mesh. A video decoder (713) decodes a displacement sub-stream, an image unpacking unit (714) unpacks an image of the decoded displacement video, and an inverse quantizer (715) inversely quantizes a quantized image. The reverse linear lifting unit (716) applies a lifting transformation in the reverse process of the encoder to generate a restored displacement. The mesh restoration unit (730) generates a reconstructed deformed mesh, i.e., a final decoded mesh, based on the restored base mesh and the restored displacement.
[0178] According to embodiments, the video decoder (717) decodes the attribute map sub-stream in the same manner as intra decoding, and the color conversion unit (740) converts the color format and / or space of the decoded attribute map to generate a decoded attribute map. The decoded mesh and the decoded attribute map can be utilized by the receiver as final mesh data that can be utilized by the user.
[0179] Referring to Fig. 12, the bitstream includes motion information (also called motion vectors), displacement, and an attribute map. Since Fig. 12 performs inter-frame decoding, it further includes a process of decoding inter-frame motion information. The motion information is decoded, and a restored quantized base mesh for the motion information is generated based on the reference base mesh, thereby generating a restored base mesh. For a description of the operation of Fig. 12, which is identical to that of Fig. 11, refer to the description of Fig. 11.
[0180] Fig. 13 illustrates a mesh data transmission device according to embodiments.
[0181] FIG. 13 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 1, the encoder (preprocessor and encoder) of FIG. 2, FIG. 6, or FIG. 7, and / or a transmitting encoding device corresponding thereto. Each component of FIG. 13 corresponds to hardware, software, a processor, and / or a combination thereof.
[0182] The operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in Fig. 13. The transmitter of Fig. 13 may perform an intra-frame encoding (or intra-encoding or intra-screen encoding) process and / or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
[0183] The pre-processor (811) receives the original mesh as input and generates a simplified mesh (decimated mesh) (or base mesh) and a fitted decimated mesh (or subdivision). Simplification can be performed based on the target number of vertices or target number of polygons that constitute the mesh. Parameterization, which generates texture coordinates and texture connection information per vertex, can be performed on the simplified mesh. For example, parameterization is a process of mapping a 3D surface to a texture domain for the decimated mesh. If parameterization is performed using the UVAtlas tool, mapping information is generated that can identify where each vertex of the decimated mesh can be mapped on a 2D image. The mapping information is expressed and stored as texture coordinates, and the final base mesh is generated through this process. In addition, the work of quantizing the mesh information in floating-point form into fixed-point form can be performed. This result can be output as a base mesh to a motion vector encoder (813) or a static mesh encoder (814) through a switching unit (812). The pre-processor (811) can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. The pre-processor (811) can generate a fitted subdivided mesh by adjusting the vertex positions so that the subdivided mesh becomes similar to the original mesh.
[0184] According to embodiments, the base mesh is output to a motion vector encoder (813) via a switching unit (812) when performing inter-encoding for the corresponding mesh frame, and is output to a static mesh encoder (814) via a switching unit (812) when performing intra-encoding for the corresponding mesh frame. The motion vector encoder (813) may be referred to as a motion encoder.
[0185] For example, when performing intra-encoding (or intra-frame encoding) on the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (814). In this case, encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. The base mesh bitstream generated through encoding is transmitted to a multiplexer (823).
[0186] As another example, when performing inter-encoding (or inter-frame encoding) for the corresponding mesh frame, the motion vector encoder (813) can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as input, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (813) can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through encoding is transmitted to the multiplexer (823).
[0187] The base mesh restoration unit (815) can receive the base mesh encoded by the static mesh encoder (814) or the motion vector encoded by the motion vector encoder (813) and generate a reconstructed base mesh. For example, the base mesh restoration unit (815) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (814) to restore the base mesh. At this time, quantization can be applied before the static mesh decoding, and inverse quantization can be applied after the static mesh decoding. As another example, the base mesh restoration unit (815) can restore the base mesh based on the reconstructed quantized reference base mesh and the motion vector encoded by the motion vector encoder (813). The reconstructed base mesh is output to the displacement calculation unit (816) and the mesh restoration unit (820).
[0188] The displacement calculation unit (816) can perform mesh refinement on the restored base mesh. The displacement calculation unit (816) can calculate a displacement vector, which is a difference value between the vertex positions of the restored base mesh and the fitted subdivision (or refined) mesh generated by the pre-processor (811). At this time, the displacement vector can be calculated as many times as the number of vertices of the refined mesh. The displacement calculation unit (816) can convert the displacement vector calculated in the 3D Cartesian coordinate system into a local coordinate system based on the normal vector of each vertex.
[0189] The displacement vector video generation unit (817) may include a linear lifting unit, a quantizer, and an image packing unit. That is, in the displacement vector video generation unit (817), the linear lifting unit may transform the displacement vector for effective encoding. The transformation may be performed by a lifting transformation, a wavelet transformation, etc., according to embodiments. In addition, quantization may be performed in a quantizer on the transformed displacement vector value, i.e., the transform coefficient. At this time, a different quantization parameter may be applied to each axis of the transform coefficient, and the quantization parameter may be derived according to an encoder / decoder agreement. The transformed and quantized displacement vector information may be packed into a 2D image in the image packing unit. The displacement vector video generation unit (817) may generate a displacement vector video by bundling packed 2D images for each frame, and the displacement vector video may be generated for each GoF (Group of Frame) unit of the input mesh.
[0190] The displacement vector video encoder (818) can encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is transmitted to a multiplexer (823).
[0191] The displacement vector restoration unit (819) may include a video decoder, an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. That is, the displacement vector restoration unit (819) performs decoding on an encoded displacement vector in the video decoder, performs image unpacking in the image unpacking unit, performs inverse quantization in the inverse quantizer, and then performs inverse transformation in the inverse linear lifting unit to restore the displacement vector. The restored displacement vector is output to the mesh restoration unit (820). The mesh restoration unit (820) restores a deformed mesh based on the base mesh restored by the base mesh restoration unit (815) and the displacement vector restored by the displacement vector restoration unit (819). The restored mesh (or referred to as a restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates.
[0192] The texture map video generation unit (821) can regenerate a texture map based on the texture map (or attribute map) of the original mesh and the restored deformed mesh output from the mesh restoration unit (820). According to embodiments, the texture map video generation unit (821) can assign color information per vertex of the texture map of the original mesh to the texture coordinates of the restored deformed mesh. According to embodiments, the texture map video generation unit (821) can generate a texture map video by grouping the regenerated texture maps by GoF unit for each frame.
[0193] The generated texture map video can be encoded using a video compression codec of a texture map video encoder (822). The texture map video bitstream generated through encoding is transmitted to a multiplexer (823).
[0194] A multiplexer (823) multiplexes a motion vector bitstream (e.g., in case of inter encoding), a base mesh bitstream (e.g., in case of intra encoding), a displacement vector bitstream, and a texture map bitstream into a single bitstream. The single bitstream can be transmitted to a receiver via a transmitter (824). Alternatively, the motion vector bitstream, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream can be generated as a file with one or more track data or encapsulated into segments and transmitted to a receiver via the transmitter (824).
[0195] Referring to FIG. 13, a transmitting device (encoder) can encode a mesh in an intra-frame or inter-frame manner. A transmitting device according to intra-encoding can generate a base mesh, a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map). A transmitting device according to inter-encoding can generate a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), and a texture map (or referred to as attribute map). The texture map obtained from the data input unit is generated and encoded based on the restored mesh. Displacement is generated and encoded through the difference in vertex positions between the base mesh and the divided (or subdivided or subdivided) mesh. More specifically, the displacement is the difference in position between the fitted sub-divided mesh and the sub-divided restored base mesh, i.e., the difference in vertex positions between the two meshes. In addition, the base mesh is generated by simplifying and encoding the original mesh through pre-processing. Motion is generated as motion vectors for the mesh of the current frame based on the reference base mesh of the previous frame.
[0196] Fig. 14 illustrates a mesh data receiving device according to embodiments.
[0197] Fig. 14 corresponds to the receiving device (110) or mesh video decoder (113) of Fig. 1, the decoder of Fig. 11 or Fig. 12, and / or the receiving decoding device corresponding thereto. Each component of Fig. 14 corresponds to hardware, software, a processor, and / or a combination thereof. The receiving (decoding) operation of Fig. 14 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of Fig. 13.
[0198] The bitstream of the mesh data received by the receiver (910) is demultiplexed into a compressed motion vector bitstream (e.g., inter decoding) or a base mesh bitstream (e.g., intra decoding), a displacement vector bitstream, and a texture map bitstream after file / segment decapsulation in the demultiplexer (911). For example, if the current mesh has inter-screen encoding (i.e., inter encoding) applied, the motion vector bitstream is received, demultiplexed, and then output to the motion vector decoder (913) via the switching unit (912). As another example, if the current mesh has intra-screen encoding (i.e., intra encoding) applied, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (914) via the switching unit (912). Here, the motion vector decoder (913) may be referred to as a motion decoder.
[0199] According to embodiments, if the current mesh has inter-screen encoding applied according to frame header information, the motion vector decoder (913) can perform decoding on the motion vector bitstream. According to embodiments, the motion vector decoder (913) can reconstruct the final motion vector by adding the previously decoded motion vector as a predictor to the residual motion vector decoded from the bitstream.
[0200] According to embodiments, if the current mesh has been subjected to in-screen encoding according to frame header information, the static mesh decoder (914) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates, normal information, etc. of the base mesh.
[0201] According to embodiments, the base mesh restoration unit (915) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by adding the decoded motion vector to the reference base mesh and then performing inverse quantization. As another example, if the current mesh has intra-screen encoding applied, the base mesh restoration unit (915) can generate a restored base mesh by performing inverse quantization on the base mesh decoded through the static mesh decoder (914).
[0202] According to embodiments, the displacement vector video decoder (917) can decode the displacement vector bitstream as a video bitstream using a video codec.
[0203] According to embodiments, the displacement vector restoration unit (918) extracts displacement vector transform coefficients from the decoded displacement vector video, and restores the displacement vector by applying inverse quantization and inverse transformation processes to the extracted displacement vector transform coefficients. To this end, the displacement vector restoration unit (918) may include an image unpacking unit, an inverse quantizer, and an inverse linear lifting unit. If the restored displacement vector is a value in a local coordinate system, a process of inversely transforming it into a Cartesian coordinate system may be performed.
[0204] The mesh restoration unit (916) can generate additional vertices by performing subdivision on the restored base mesh. Through subdivision, vertex connection information including the added vertices, texture coordinates, and texture coordinate connection information can be generated. At this time, the mesh restoration unit (916) can generate a final restored mesh (or a restored deformed mesh) by combining the subdivided restored base mesh with the restored displacement vector.
[0205] According to embodiments, the texture map video decoder (919) can decode the texture map bitstream as a video bitstream using a video codec to restore the texture map. The restored texture map has color information for each vertex contained in the restored mesh, and the color value of each vertex can be obtained from the texture map using the texture coordinates of each vertex.
[0206] According to embodiments, the mesh restored by the mesh restoration unit (916) and the texture map restored by the texture map video decoder (919) are shown to the user through a rendering process in the mesh data renderer (920).
[0207] Referring to FIG. 14, a receiving device (decoder) can decode a mesh in an intra-frame or inter-frame manner. A receiving device according to intra-decoding can receive a base mesh, a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map. A receiving device according to inter-decoding can receive a motion vector (or referred to as motion), a displacement vector (or referred to as displacement), a texture map (or referred to as attribute map), and render mesh data based on the restored mesh and the restored texture map.
[0208] A mesh data transmission device and method according to embodiments may pre-process mesh data, encode the pre-processed mesh data, and transmit a bitstream including the encoded mesh data. A point mesh data reception device and method according to embodiments may receive a bitstream including mesh data and decode the mesh data. The mesh data transmission and reception method / device according to embodiments may be abbreviated as method / device according to embodiments. The mesh data transmission and reception method / device according to embodiments may also be referred to as 3D data transmission and reception method / device or point cloud data transmission and reception method / device.
[0209] As mentioned above, the current V-DMC mesh data is largely composed of a base mesh, displacements, and texture map data (or information). At this time, the base mesh is coded using a static mesh codec, the displacements are coded using a video codec or arithmetic coding method, and the texture map (or attribute map) is coded using a video codec method. In other words, each component (base mesh, displacement, texture map) is compressed using its own efficient coding method, and in this process, some frames or pictures may be skipped.
[0210] However, there is no technology that omits texture map coding among the technologies applied to V-DMC to date.
[0211] In the present disclosure, the current texture map coding can be omitted when the similarity between the reference texture map and the current texture map is high. Furthermore, when texture map coding omission is performed only on a frame-by-frame basis, there may be cases where texture map coding omission is not performed due to differences in some areas while most of the texture maps are similar. Even in such cases, to enable efficient compression through texture map coding omission, the present disclosure can also omit texture maps on a submesh, patch, tile, or other basis.
[0212] In this way, the present disclosure can skip the current texture map coding when the similarity between the reference texture map of the dynamic mesh and the current texture map is high. At this time, the unit for skipping the texture map coding may be a frame, a submesh, a patch, a tile, etc. That is, the present disclosure can obtain the effect of reducing the bits of the transmitted texture map and the encoding / decoding complexity by omitting the texture map that occupies a considerably large proportion of the V-DMC bitstream data. In other words, the present disclosure can increase the compression efficiency of the V-DMC codec by increasing the compression efficiency of the texture map that occupies the largest bit proportion among the base mesh, displacement vector, and texture map, which are the main components of the V-DMC data.
[0213] In one embodiment, the present disclosure proposes a method for applying a texture map coding skip technique in a V-DMC codec and a weight-based processing method based on a reference distance. In addition, even when the texture map skip technique is applied, an efficient signaling method is required to convey necessary information. To this end, a signaling method, syntax, and semantics through a SEI (Supplementary enhancement information) message are proposed so that the texture map skip technique can be applied in the V-DMC codec. That is, the present disclosure proposes an encoding and decoding process in a V-DMC codec for the operation of the V-DMC-based frame-level texture map coding skip technique, a method for restoring a texture map skipped at the frame level and a weight-based restoration method based on a reference distance, and a method for signaling information on whether texture map coding is skipped per frame (e.g., a flag) through an SEI message and reference information for deriving an omitted texture map.
[0214] In another embodiment, the present disclosure proposes a method for restoring an omitted texture map based on a multi-reference structure when a frame with an omitted texture map is transmitted by applying a texture map coding omission technique in a V-DMC codec. In addition, a method for signaling a related syntax for restoring an omitted texture map through a V3C (visual volumetric video-based coding) unit header is proposed. The present disclosure proposes a method for signaling through V3C_AVD (Attribute Video Data) of an existing V3C unit header and a method for generating a new V3C unit type, V3C_TVD (Texture map Video Data), and signaling through it. In the present disclosure, V3C can be used in the same sense as V-PCC (video-based Point Cloud Coding). That is, the present disclosure proposes a method for restoring omitted texture map frames based on a multi-reference structure, a method for signaling information (e.g., a flag) on whether texture map coding is omitted through a V3C unit header and reference direction information syntaxes for deriving omitted frames, and a V3C_TVD structure, which is a new V3C unit type for texture maps other than the V3C_AVD structure of the existing V3C unit header.
[0215] To this end, the transmitting device simplifies the original mesh and generates a base mesh through a parameterization process. The generated base mesh is quantized and passes through a static mesh encoder to generate a base mesh sub-bitstream. In addition, the mesh data after the process of subdividing and fitting the simplified mesh from the original mesh and the mesh data restored from the previously encoded mesh are compared to calculate a displacement vector, which is the difference between each vertex. In order to efficiently encode the calculated displacement vector, the displacement vector coordinate system is transformed into a local coordinate system, and the displacement vector transformed into the local coordinate system is converted into displacement vector coefficients, quantized, and then encoded to generate a displacement vector sub-bitstream. Then, the encoded base mesh and displacement vector are restored again to restore the mesh, and a texture map of the restored mesh is generated through the texture coordinates and connection information of the restored mesh and the relationship between the original mesh and the texture map of the original mesh, and encoded to generate a texture map sub-bitstream. The three sub-bitstreams generated in the above process are then multiplexed to create a single V-DMC bitstream and transmitted. Figure 15 is a drawing to explain this in more detail.
[0216] Fig. 15 illustrates a transmitting device according to embodiments. The transmitting device of Fig. 15 may be referred to as a mesh data transmitting device or an encoder or an encoder of a transmitting device or a V-Mesh encoder or a dynamic mesh encoder.
[0217] FIG. 15 corresponds to the transmitting device (100) or mesh video encoder (102) of FIG. 1, the encoder (preprocessor and encoder) of FIG. 2, FIG. 6, or FIG. 7, the transmitting device of FIG. 13, and / or the transmitting encoding device corresponding thereto. The components of FIG. 15 may be implemented by hardware, software, a processor, and / or a combination thereof. That is, the components of the transmitting device of FIG. 15 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not shown in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the components of the transmitting device of FIG. 15 described above. In addition, one or more processors may operate or execute a set of software programs and / or instructions for performing the operations and / or functions of the elements of the transmitting device of FIG. 15. In Fig. 15, the execution order of each block may be changed, some blocks may be omitted, and some blocks may be newly added.
[0218] In the present disclosure, the operation process of a transmitter for compressing and transmitting dynamic mesh data using V-Mesh compression technology may be as shown in FIG. 15. The transmitter of FIG. 15 may support both an intra-frame encoding (or intra-encoding or intra-screen encoding) process and / or an inter-frame encoding (or inter-encoding or inter-screen encoding) process.
[0219] In Fig. 15, the mesh simplification unit (11011) simplifies the input original mesh through a mesh simplification algorithm to generate a base mesh (or simplified base mesh or simplified mesh). At this time, mesh simplification can be performed based on the number of target vertices or target polygons constituting the mesh. For example, a method such as decimation can be used as a mesh simplification algorithm that simplifies the original mesh. That is, the decimation method can be a process of selecting vertices to be removed from the original mesh using a certain reference point, and then removing the selected vertices and the triangles connected to the selected vertices.
[0220] That is, the mesh simplification unit (11011) can simplify the input mesh by the target number of vertices or the target number of faces. At this time, the simplification process can be performed through various methods such as triangle collapse and edge collapse.
[0221] According to embodiments, the base mesh simplified in the mesh simplification unit (11011) is provided to the mesh parameterization unit (11012) and the mesh refinement unit (11018).
[0222] The mesh parameterization unit (11012) performs a process of mapping a 3D surface to a texture domain for a simplified mesh (decimated mesh). That is, the mesh parameterization unit (11012) generates texture coordinates and texture connection information of the input mesh. In one embodiment, the mesh parameterization unit (11012) may perform parameterization using a UV Atlas tool. Through this process, mapping information is generated regarding which location on a 2D image each vertex of the simplified mesh (decimated mesh) can be mapped to. The mapping information is expressed and stored as texture coordinates, and through this process, the final base mesh is generated. That is, the mesh parameterization unit (11012) performs parameterization to generate texture coordinates (UV coordinates) and texture connection information per vertex of the input mesh (i.e., simplified mesh or simplified base mesh).
[0223] The final base mesh (or base mesh with texture map) generated in the above parameterization unit (11012) is input to the mesh quantization unit (11013) and quantized.
[0224] According to embodiments, the mesh quantization unit (11013) may perform a task of quantizing floating-point type mesh information (e.g., geometry information (x, y, z) or / and texture coordinates (u, v), normal information (nx, ny, nz), etc.) into fixed-point type. That is, the mesh quantization unit (11013) may quantize vertex coordinates and texture coordinates of the base mesh. According to embodiments, quantization for specific components may be omitted.
[0225] The above mesh subdivision unit (11018) subdivides the base mesh simplified by the mesh simplification unit (11011). That is, the mesh subdivision unit (11018) can perform mesh subdivision on the base mesh to generate additional vertices. Depending on the subdivision method, vertex connection information, texture coordinates, and texture coordinate connection information including the added vertices can be generated. At this time, depending on the subdivision method, geometry information connection information, texture coordinate connection information, and texture coordinates can be implicitly derived and generated. According to embodiments, the mesh subdivision unit (11018) can perform subdivision through a method such as mid-edge, Loop, or Catmul&Clark.
[0226] In addition, in the mesh subdivision unit (11018), mesh subdivision can be performed n times by user parameters or an agreement between the encoder / decoder, and when the vertex of the base mesh is defined as R9, the vertex newly created by performing subdivision 1 time is defined as R1, and the vertex created by performing subdivision n times is defined as Rn, LoD n can be defined as follows:
[0227] LOD n = R9 R9 , ..., R n
[0228] According to embodiments, the mesh fitting unit (11019) may perform fitting by adjusting vertex positions so that the mesh subdivided by the mesh subdivision unit (11018) becomes similar to the original mesh, thereby generating a fitted subdivided mesh. That is, the mesh fitting unit (11019) performs vertex position adjustment so that the subdivided mesh becomes similar to the original mesh. According to embodiments, the mesh simplification unit (11011), the mesh parameterization unit (11012), the mesh subdivision unit (11018), and the mesh fitting unit (11019) may be omitted, and when the corresponding processes are omitted, the original mesh may be applied as an input to the mesh quantization unit (11013).
[0229] At this time, coordinate information of the original mesh can be applied as input to the displacement vector calculation unit (11020), and according to embodiments, the displacement vector encoding process (i.e., displacement vector calculation unit (11020), displacement vector coordinate system conversion unit (11021), displacement vector encoder (11022)) can be omitted.
[0230] The present disclosure may be referred to as a pre-processor, including a mesh simplification unit (11011), a mesh parameterization unit (11012), a mesh refinement unit (11018), and a mesh fitting unit (11019). According to embodiments, the pre-processor may further include a displacement vector calculation unit (11020).
[0231] According to embodiments, the quantized base mesh in the mesh quantization unit (11013) can be output to a motion vector encoder (11015) or a static mesh encoder (11016) through a switching unit (11014).
[0232] According to embodiments, the base mesh is output to a motion vector encoder (11015) through a switching unit (11014) when performing inter-encoding for the corresponding mesh frame, and is output to a static mesh encoder (11016) through a switching unit (11014) when performing intra-encoding for the corresponding mesh frame. The motion vector encoder (11015) may be referred to as a motion encoder.
[0233] For example, when performing intra encoding or intra frame encoding for the corresponding mesh frame, the base mesh can be compressed through a static mesh encoder (11016). In this case, encoding can be performed on connection information, vertex geometry information, vertex texture information, normal information, etc. of the base mesh. That is, vertex coordinates, vertex connection information, texture coordinates, texture connection information, etc. of the mesh can be encoded in the static mesh encoder (11016). The base mesh bitstream generated through encoding is transmitted to a multiplexer (not shown).
[0234] As another example, when performing inter-encoding (or inter-frame encoding) on the corresponding mesh frame, the motion vector encoder (11015) can receive a base mesh and a reference reconstructed base mesh (or a reconstructed quantized reference base mesh) as input, calculate a motion vector between the two meshes, and encode the value. In addition, the motion vector encoder (11015) can perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and entropy-encode a differential motion vector (or residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated through the encoding is transmitted to a multiplexer (not shown) as a base mesh bitstream. That is, in the case of intra-frame encoding, the static mesh bitstream is input to the multiplexer as a base mesh bitstream, and in the case of inter-frame encoding, the motion vector bitstream is input to the multiplexer as a base mesh bitstream.
[0235] In FIG. 15, the base mesh decoder (11017, or base mesh restoration unit) can receive a base mesh encoded by a static mesh encoder (11016) or a motion vector encoded by a motion vector encoder (11015) and generate a reconstructed base mesh. The base mesh decoder (11017) performs restoration of the base mesh according to the encoding type (inter-screen encoding or intra-screen encoding) of the current mesh. For example, the base mesh decoder (11017) can perform static mesh decoding on the base mesh encoded by the static mesh encoder (11016) to restore the base mesh. At this time, quantization can be applied before static mesh decoding, and inverse quantization can be applied after static mesh decoding. That is, when intra-screen encoding is performed, the current base mesh can be restored by performing dequantization on the quantized base mesh through the mesh quantization unit (11013). As another example, the base mesh decoder (11017) can restore the base mesh based on the restored quantized reference base mesh and the motion vector encoded by the motion vector encoder (11015). That is, when inter-screen encoding is performed, the current base mesh can be generated by decoding the motion vector using the motion vector decoding method and then applying (i.e., adding) the decoded motion vector to the reference restored base mesh. At this time, when the motion vector is not quantized, the motion vector restoration process is omitted and the current base mesh can be restored using the motion vector calculated by the motion vector encoder (11015). The restored base mesh is output to the displacement vector calculation unit (11020) and the mesh dequantization unit (11024).
[0236] According to embodiments, the displacement vector calculation unit (11020) can perform mesh refinement on the restored base mesh. In addition, the displacement vector calculation unit (11020) can calculate a displacement vector, which is a difference value of vertex positions between the restored base mesh that has been refined and the fitted subdivision (or refined) mesh generated by the mesh fitting unit (11019). At this time, the displacement vector can be calculated as many times as the number of vertices of the refined mesh. That is, the displacement vector of the number of vertices of the refined mesh can be calculated through the displacement vector calculation unit (11020).
[0237] According to embodiments, the displacement vector coordinate system conversion unit (11021) may output the vertex displacement vector calculated in the 3D Cartesian coordinate system (i.e., (x, y, z) space) (or canonical coordinate system) as it is, or may convert it into a local coordinate system (i.e., normal, tangential, bi-tangential coordinate system) based on the normal vector of each vertex. At this time, the normal vector may be calculated for each subdivided vertex based on the geometry information and connection information of the surrounding vertices. According to embodiments, among the normal, tangential, and bi-tangential components of the local coordinate system, only the normal component may be encoded. When the coordinate system conversion is applied according to the agreement between the encoder / decoder, encoding of only the normal component is always performed, or the encoder may decide to signal a 1-bit flag. At this time, the normal vector can be calculated for each subdivided vertex based on the geometric information and / or connection information of the surrounding vertices.
[0238] And, in the displacement vector coordinate system conversion unit (11021), whether or not to perform displacement vector coordinate system conversion is determined by an agreement between the encoder (i.e., transmitting device) / decoder (i.e., receiving device), or a coordinate system conversion status flag (e.g., (asps_vmc_ext_displacement_coordinate_system)) may be transmitted in units such as sequence, GOF (Group of frame), frame, and sub-mesh to determine whether or not to perform coordinate system conversion. As an example, the displacement vector coordinate system conversion status flag (asps_vmc_ext_displacement_coordinate_system), which is information that can identify whether or not to perform displacement vector coordinate system conversion, may be signaled in signaling information (e.g., atlas sequence parameter set, ASPS) and transmitted to the receiving device. For example, if the value of the displacement vector coordinate system transformation flag (asps_vmc_ext_displacement_coordinate_system) syntax (or field) is 0, the canonical coordinate system is used as is, and if it is 1, it indicates that the transformation is performed to the local coordinate system.
[0239] According to embodiments, the displacement vector of the cartesian coordinate system or the displacement vector converted to the local coordinate system in the displacement vector coordinate system conversion unit (11021) is encoded into a displacement vector bitstream (or displacement vector video bitstream) in the displacement vector encoder (11022). In one embodiment, the displacement vector encoder (11022) can encode the displacement vector of the cartesian coordinate system or the displacement vector converted to the local coordinate system into the displacement vector bitstream using a 2D video encoder such as H.264, HEVC, or VVC.
[0240] That is, the displacement vector encoder (11022) can perform encoding on a displacement vector or a displacement vector coefficient (or displacement vector transform coefficient). In the present disclosure, the displacement vector encoder (11022) can perform encoding through a video codec-based encoder, a zero run-length encoder, an arithmetic encoder, etc. For example, when the encoding method is video codec-based encoding, the displacement vector encoder (11022) can encode the displacement vector or the displacement vector coefficient by packing it into a frame. That is, in the displacement vector encoder (11022), the displacement vector coefficients can be packed into a 2D image and then encoded using a 2D video codec (i.e., a video compression codec), or zero run-length encoded, or arithmetic encoded to generate a displacement vector video bitstream.
[0241] According to embodiments, a displacement vector video bitstream encoded and generated by a displacement vector encoder (11022) is transmitted to a multiplexer (not shown). According to embodiments, a method for selecting encoding of the displacement vector encoder (11022) may use a displacement vector encoder promised in an encoder (i.e., a transmitting side) / decoder (i.e., a receiving side), or may analyze the characteristics of a displacement vector in an encoder on the transmitting side and transmit the type of a selected displacement vector encoder to a decoder on the receiving side.
[0242] According to embodiments, the displacement vector restoration unit (11023) can restore the displacement vector by performing the reverse process of displacement vector encoding on the displacement vector or displacement vector coefficient encoded by the displacement vector encoder (11022). That is, the displacement vector restoration unit (11023) can perform displacement vector depacking (or unpacking) depending on the method of encoding the displacement vector, for example, when encoded based on a video codec. More specifically, the displacement vector restoration unit (11023) decodes a bitstream encoded through a 2D video encoder by being packed into a 2D image / video using a 2D video decoder, and performs depacking. Then, dequantization is performed on the quantized transform coefficients on which depacking was performed, and inverse transformation is performed to calculate the restored displacement vector. That is, the displacement vector restoration unit (11023) can additionally perform inverse quantization, inverse transformation, etc. processes depending on whether quantization and transformation processes are performed during the displacement vector encoding process.
[0243] According to embodiments, the mesh dequantization unit (11024) can dequantize vertex coordinates or texture coordinates of the restored base mesh using the reverse process of quantization. If the quantization process is omitted in the mesh quantization unit (11013), the dequantization process is also omitted in the mesh dequantization unit (11024).
[0244] According to embodiments, the mesh restoration unit (11025) can restore a mesh based on a restored displacement vector output from the displacement vector restoration unit (11023) and a restored base mesh (or a dequantized restored base mesh) output from the mesh dequantization unit (11024). More specifically, the mesh restoration unit (11025) can perform subdivision on the restored base mesh output from the mesh dequantization unit (11024) and add the restored displacement vector from the displacement vector restoration unit (11023) to generate a reconstructed deformed mesh. According to embodiments, the mesh restored by the mesh restoration unit (11025) (or referred to as a restored mesh or a restored deformed mesh) has restored vertices, connection information between vertices, texture coordinates, and connection information between texture coordinates. The restored mesh (or restored mesh or restored deformed mesh) generated in the above mesh restoration unit (11025) is provided to the texture map generation unit (11026).
[0245] According to embodiments, the texture map generation unit (11026) can regenerate the texture map of the current mesh based on the texture map (or attribute map) of the original mesh and the mesh restored by the mesh restoration unit (11025). That is, the texture map generation unit (11026) can generate the texture map of the restored mesh through the texture map of the original mesh and the relationship between the original mesh and the restored mesh. In other words, the texture map generation unit (11026) generates the texture map of the restored mesh through the texture coordinates and connection information of the restored mesh and the relationship between the original mesh and the texture map of the original mesh.
[0246] According to embodiments, the texture map generation unit (11026) can assign color information per vertex of the texture map of the original mesh to the texture coordinates of the restored base mesh (or the restored deformed mesh). According to embodiments, the texture map generation unit (11026) can generate a texture map (or texture map video) by grouping the regenerated texture maps by GoF unit for each frame.
[0247] According to embodiments, after the texture map generation unit (11026) generates a texture map, the texture map omission decision unit (11027) determines whether to omit encoding of the current texture map. In one embodiment of the present disclosure, if the similarity between the texture map of the reference frame and the texture map of the current frame is high, encoding of the current texture map is omitted. Here, omitting encoding of the current texture map (i.e., the texture map of the current frame) includes not only not encoding the current texture map but also not transmitting the current texture map to the receiving device. The present disclosure refers to this case as a texture map coding omission mode. In the present disclosure, the texture map omission decision unit (11027) may be included in the texture map encoder (11028) or may be configured as a separate block or module as shown in FIG. 15.
[0248] In the present disclosure, the similarity between the two frames can be determined by comparing and measuring the mesh data of the original mesh and the current frame, and the mesh data of the original mesh and the reference frame, using a point cloud-based metric, etc., and comparing the difference value with a threshold value.
[0249] Fig. 16 is a diagram showing an example of an encoding process when texture map coding is omitted according to embodiments. As shown in Fig. 16, texture map data of frames other than frames for which texture map coding omission has been determined are stored in a texture map buffer and can be encoded into a texture map sub-bitstream (or texture map bitstream) through a video encoder.
[0250] For example, assuming that there are 8 frames in a GOP (Group Of Picture) and that texture map coding of 3 of these frames is omitted, only the texture maps of 5 frames are stored in the texture map buffer, encoded through the video encoder (i.e., V-DMC texture map encoder), and then transmitted.
[0251] At this time, the number of output texture maps may differ from the input due to omission of texture map coding during the texture map encoding process. Fig. 16 shows an example where eight frames of texture maps are input, but the number of texture map frames output after encoding is five due to omission of texture map coding for three frames.
[0252] According to embodiments, when the encoder (e.g., the texture map skip decision unit (11027)) determines to skip texture map coding of the current frame, it may signal texture map coding skip information (texture_skip_flag), texture map reference direction information (refDirection_idx), and / or reference texture frame index information (ref_texture_idx).
[0253] And, as information about a texture map reference target, texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx) can be signaled at the same time, and depending on the implementation, only one of the two pieces of information may be signaled. That is, information about a texture map reference target can include texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx).
[0254] At this time, if it is determined that the coding of the current texture map is skipped, texture_skip_flag=1 can be signaled, and if the coding is not skipped, texture_skip_flag=0 can be signaled. That is, the texture map coding skip information (texture_skip_flag) indicates whether the coding of the texture map is skipped. For example, if the value of the texture map coding skip information (texture_skip_flag) is 1, it can indicate that the coding of the texture map of that frame is skipped, and if it is 0, it can indicate that it is not skipped.
[0255] And, if the coding of the current texture map is decided to be omitted, the texture map in the direction with a smaller difference is referenced based on the calculated PSNR (Peak-Signal-to-Noise Ratio), and the texture map reference direction information (refDirection_idx) according to it can be signaled together. For example, if the texture map reference direction is a frame before (past) the current frame, refDirection_idx=0 can be signaled. If the texture map reference direction is a frame after (future) the current frame, refDirection_idx=1 can be signaled. If the texture map reference direction is a frame in both directions, refDirection_idx=2 can be signaled. And, if the texture map coding of the current frame is not omitted, refDirection_idx=3 can be signaled or not signaled. That is, the texture map reference direction information (refDirection_idx) indicates reference texture map direction information. For example, if the value of the texture map reference direction information (refDirection_idx) is 00, it can indicate that there is no reference direction because there is no texture map coding omission. In addition, if the value of the texture map reference direction information (refDirection_idx) is 01 to 10, it indicates that there is texture map coding omission, and the referenced frame changes depending on the value. For example, if the value of the texture map reference direction information (refDirection_idx) is 01, it can indicate that a past frame is referenced based on the current frame, if it is 10, it can indicate that a future frame is referenced based on the current frame, and if it is 11, it can indicate that a bidirectional frame is referenced based on the current frame.
[0256] According to embodiments, the present disclosure can signal texture map coding skip information (texture_skip_flag) and texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx) through an SEI message (sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI) of a NAL sample stream in a V3C_AD (Atlas Data) atlas sub-bitstream of a V3C sample stream.
[0257] Meanwhile, a texture map that is determined not to be omitted in the texture map omission decision unit (11027) may be encoded in a texture map encoder (11028). For example, the texture map encoder (11028) may encode the texture map using a 2D video codec-based encoder, a zero run length encoder, an entropy coding-based arithmetic encoder, etc. In addition, the texture map encoder (11028) may further perform color space conversion of the texture map. Then, a texture map substream (or texture map video bitstream) generated through texture map encoding is transmitted to a multiplexer (not shown).
[0258] According to embodiments, the type of texture map encoder (11028) may include a video encoder (e.g., VVC, HEVC, etc.), an entropy coding-based encoder, etc. In addition, a method for selecting a texture map encoder (11028) may use a texture map encoder promised in an encoder (i.e., a transmitting side) / decoder (i.e., a receiving side), or may transmit the type of texture map encoder selected by the encoder on the transmitting side to the decoder on the receiving side.
[0259] As described above, the texture map generation unit (11026) performs a process of generating a new texture map having color information corresponding to the texture coordinates of the restored mesh. In addition, the texture map omission determination unit (11027) can determine whether to omit texture map coding of the current frame by calculating the difference in texture PSNR between the mesh of the current frame and the original mesh and the texture PSNR between the mesh of the reference frame and the original mesh using a method such as a point-based metric. When it is determined that the texture map coding of the current frame is omitted, information on whether to omit texture map coding (texture_skip_flag) can be signaled and transmitted to the receiving device via an SEI message. In the present disclosure, the information on whether to omit texture map coding is also referred to as a texture map coding omission flag. In addition, direction information (refDirection_idx) of the reference texture map having a smaller PSNR difference calculated in advance and / or index information (ref_texture_idx) of the reference texture map can be signaled and transmitted to the receiving device. At this time, the direction information (refDirection_idx) of the reference texture map and / or the index information (ref_texture_idx) of the reference texture map can also be signaled and transmitted to the receiving device via an SEI message. In addition, the original sequence (GoF) size information (num_original_gof_size) before performing texture map coding skip and information on whether to continuously apply the SEI message during the current frame or sequence (e.g., persistence_flag) can be signaled and transmitted to the receiving device via an SEI message.Information signaled through the above SEI message can be transmitted to the receiving device through the SEI message (sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI) of the NAL sample stream in the V3C_AD (Atlas Data) atlas sub bitstream (atlas_sub_bitstream) of the V3C sample stream. Then, the texture map encoder (11028) performs encoding using a 2D video encoder such as H.264, HEVC, or VVC for the remaining texture maps except for the texture maps for which texture map omission has been determined in the texture map omission determination unit (11027) to generate a texture map bitstream (or sub-bitstream).
[0260] According to embodiments, a multiplexer (not shown) may multiplex the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream generated through the entire process into a single bitstream and then transmit the multiplexer to a receiving device. That is, a sub-bitstream may be generated for each component and transmitted as a single bitstream from the transmitting end through the multiplexer. Alternatively, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream may be encapsulated into a file / segment and transmitted to the receiving device. Here, the bitstream may be referred to as a sub-bitstream.
[0261] According to embodiments, the bitstream multiplexed in the multiplexer may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD.
[0262] Fig. 17 illustrates another example of a receiving device according to embodiments. In the present disclosure, the receiving device of Fig. 17 may be referred to as a mesh data receiving device or decoder or a decoder of a receiving device or a V-Mesh decoder or a dynamic mesh decoder.
[0263] FIG. 17 corresponds to the receiving device (110) or mesh video decoder (113) of FIG. 1, the decoder of FIG. 11 or FIG. 12, the receiving device of FIG. 14, and / or the receiving decoding device corresponding thereto. The receiving (decoding) operation of FIG. 17 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of FIG. 15. The components of FIG. 17 may be implemented by hardware, software, a processor, and / or a combination thereof. That is, the components of the receiving device of FIG. 17 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not shown in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the components of the receiving device of FIG. 17 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing operations and / or functions of elements of the receiving device of FIG. 17. The execution order of each block in FIG. 17 may be changed, some blocks may be omitted, and some new blocks may be added.
[0264] FIG. 17 may largely include a base mesh decoding unit, a displacement information decoding unit, and a texture map decoding unit. According to embodiments, the base mesh decoding unit may include a switching unit (15011), a motion vector decoder (15012), a static mesh decoder (15013), a base mesh restoration unit (15014), a mesh refinement unit (15015), and a mesh restoration unit (15016). According to embodiments, the displacement information decoding unit may include a displacement vector decoder (15017), a displacement vector inverse quantization unit (15018), a displacement vector inverse transformation unit (15019), and a displacement vector coordinate system inverse transformation unit (15020).
[0265] According to embodiments, a bitstream of mesh data received by a receiver (not shown) may be demultiplexed into a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream in a demultiplexer (not shown) after file / segment decapsulation. If the current mesh has inter-screen encoding (i.e., inter-encoding) applied, the base mesh bitstream may be a motion vector bitstream. Here, the bitstream may be referred to as a sub-bitstream.
[0266] According to embodiments, the base mesh bitstream is provided to a motion vector decoder (15012) via a switching unit (15011) or to a static mesh decoder (15013).
[0267] For example, if the current mesh has inter-screen encoding (i.e., inter encoding) applied, the base mesh bitstream, i.e., the motion vector bitstream, is received, demultiplexed, and then output to the motion vector decoder (15012) through the switching unit (15011). As another example, if the current mesh has intra-screen encoding (i.e., intra encoding) applied, the base mesh bitstream is received, demultiplexed, and then output to the static mesh decoder (15013) through the switching unit (15011). Here, the motion vector decoder (15012) may be referred to as a motion decoder.
[0268] According to embodiments, the motion vector decoder (15012) can perform decoding on a motion vector bitstream on a vertex-by-vertex basis or a subgroup basis.
[0269] According to embodiments, the motion vector decoder (15012) can reconstruct a final motion vector by adding a differential motion vector (i.e., a residual motion vector) decoded from a bitstream using a previously decoded motion vector as a predictor. That is, the motion vector decoder (15012) can decode a differential motion vector (or a residual motion vector) in units of vertices or subgroups (or subblocks) through a motion vector bitstream, and perform prediction based on connection information using a previously decoded motion vector as a predictor to decode the motion vector by adding it to the residual motion vector.
[0270] According to embodiments, the static mesh decoder (15013) can decode the base mesh bitstream to restore connection information, vertex geometry information, texture coordinates (i.e., attribute geometry information), normal information, etc. of the base mesh.
[0271] According to embodiments, the base mesh restoration unit (15014) can restore the current base mesh based on the decoded motion vector or the decoded base mesh. For example, if the current mesh has inter-screen encoding applied, the base mesh restoration unit (15014) can add the decoded (or restored) motion vector to the reference base mesh and then perform inverse quantization to generate a restored base mesh (i.e., the current base mesh). As another example, if the current mesh has intra-screen encoding applied, the base mesh restoration unit (15014) can perform inverse quantization on the base mesh decoded (or restored) through the static mesh decoder (15012) to generate a restored base mesh (i.e., the current base mesh).
[0272] According to embodiments, the mesh subdivision unit (15015) can generate additional vertices by performing subdivision on the base mesh. The present disclosure can implicitly derive and generate geometry information connection information, texture coordinate connection information, and texture coordinates according to the subdivision method.
[0273] According to embodiments, the mesh refinement unit (15015) can perform refinement through methods such as mid-edge, Loop, and Catmul&Clark.
[0274] According to embodiments, mesh subdivision in the mesh subdivision unit (15015) may be performed n times by user parameters or encoder / decoder promises. According to embodiments, the vertices of the base mesh are R9, the newly generated vertices are R1 by performing subdivision 1, … The vertices generated by performing subdivision n times are R n When defined as LoD n can be defined as follows:
[0275] LOD n = R9 R9 , ..., R n
[0276] According to embodiments, the displacement vector decoder (15017) may perform video codec-based decoding on the demultiplexed displacement vector bitstream as a video bitstream, or perform zero run-length decoding, or perform arithmetic decoding. In the present disclosure, the displacement vector decoder may be used interchangeably with the displacement vector transform decoder.
[0277] According to embodiments, the displacement vector decoder (15017) can restore the displacement vector by decoding the displacement vector in a reverse process of the displacement vector encoding method of the transmitting side.
[0278] According to embodiments, the displacement vector dequantization unit (15018) can dequantize displacement vector coefficients restored by the displacement vector decoder (15017). According to embodiments, the displacement vector transform coefficients can be quantized through different quantization parameters for each axis, and in this case, the displacement vector dequantization unit (15018) can determine the quantization rate for each LoD level by deriving a quantization parameter or a scaling parameter by an encoder / decoder agreement.
[0279] According to embodiments, the displacement vector inverse transform unit (15019) performs an inverse transform of the transform performed in the encoder of the transmitting device on the inverse quantized displacement vector coefficients to output displacement vectors. According to embodiments, a lifting inverse transform, a wavelet inverse transform, etc. may be performed.
[0280] If the lifting inverse transformation is performed, the vertex R of the kth subdivision level k R as a predictor when performing prediction t (t <k 또는 t<=k)의 세분화 정점 변위 벡터를 통해 k번째 세분화 레벨의 변위벡터 예측을 수행할 수 있다.
[0281] According to embodiments, when performing prediction of a displacement vector, an average or distance-based weighted average prediction can be performed on n points near the current vertex based on connection information among vertices with a lower level of detail than the current vertex.
[0282] According to embodiments, prediction can be performed based on the displacement vectors of n vertices used to generate the current vertex in the mesh refinement step.
[0283] Additionally, when the lifting inverse transformation is performed, a process of updating the displacement vector of the vertex used for prediction in the encoder can be performed through the parsed residual signal.
[0284] According to embodiments, the displacement vector coordinate system inverse transformation unit (15020) parses the coordinate system transformation flag included in the signaling information in units of sequence or GOF (Group of Frames) or frame or submesh, and if the value is 1, the inverse quantized (or inversely transformed) restored displacement vector can be inversely transformed from the local coordinate system (n, t, b) to the canonical coordinate system (x, y, z) to restore the final displacement vector.
[0285] According to embodiments, coordinate system inversion can always be performed without flag transmission.
[0286] The output of the displacement vector coordinate system inverse transformation unit (15020) is provided to the mesh restoration unit (15016).
[0287] That is, in the encoder of the transmitting device, the vertex displacement vector calculated in the (x,y,z) space can be converted to a (normal, tangential, bi-tangential) coordinate system (or local coordinate system) based on the normal vector of each vertex. At this time, the normal vector can be calculated for each subdivided vertex based on the geometric information and connection information of the surrounding vertices.
[0288] According to embodiments, the mesh restoration unit (15016) restores the mesh based on the mesh refined by the mesh refinement unit (15015) and the final restored displacement vector output from the displacement vector coordinate system inversion unit (15020). That is, the process of re-refining the restored base mesh is performed, and the final restored displacement vector is added to restore the final mesh geometry information.
[0289] According to embodiments, the received and demultiplexed texture map bitstream is input to a texture map decoder (15021). According to embodiments, the texture map decoder (15021) can decode the texture map bitstream to restore the texture map. For example, the texture map decoder (15021) can decode the texture map bitstream to restore the texture map using a 2D video codec-based decoder, a zero run length decoder, an entropy coding-based arithmetic decoder, or the like.
[0290] Then, the final restored mesh is restored based on the texture map restored by the texture map decoder (15021) and the geometry information restored by the mesh restoration unit (15016).
[0291] At this time, if texture map coding is omitted on the transmitting side, a process of restoring the omitted texture map using reference information for the omitted texture map (e.g., reference texture map information) is required. For example, the omitted texture map restoration unit (15022) can generate the omitted texture map on the transmitting side based on the reference texture map information received and included in the signaling information. The omitted texture map restoration unit (15022) in the present disclosure may be included in the texture map decoder (15021) or may be configured as a separate block or module as shown in FIG. 17. The process of restoring a texture map whose coding is omitted on the transmitting side by the omitted texture map restoration unit (15022) will be described in detail below.
[0292] Meanwhile, the present disclosure enables the encoding and decoding of media representing dynamic meshes using V3C technology. This can be achieved by converting an input dynamic mesh representation into the following multiple V3C components: a base mesh, a set of displacements, a 2D representation of attributes, and an atlas. Here, attributes can be used in the same or similar sense as texture maps. Furthermore, displacements (or displacement vectors) can be used in the same or similar sense as geometry.
[0293] At this time, the base mesh can be encoded as described above through a static mesh encoder in the case of intra coding, or through a motion vector encoder in the case of inter coding, to generate a base mesh sub-bitstream. In addition, the displacement (or displacement vector) can be encoded through an encoder specified by a profile or SEI message (e.g., a video codec-based encoder, a zero run length encoder, an arithmetic encoder, etc.) to generate a displacement vector sub-bitstream. In addition, the attribute (or texture map) can be encoded through a 2D video codec-based encoder, a zero run length encoder, an arithmetic encoder, etc. to generate an attribute sub-bitstream. At this time, the attribute to be encoded is an attribute for which coding omission is determined not to be performed, that is, a texture map. A detailed description of the coding omission of the attribute will be referred to the descriptions of FIGS. 15 and 16, and will be omitted here to avoid redundant description.
[0294] An atlas is a collection of 2D bounding boxes and their associated information, arranged on a rectangular frame, representing the space in which volumetric data is rendered in 3D space. Furthermore, an atlas is a list of metadata corresponding to a portion of a mesh surface in 3D space.
[0295] In particular, the atlas component provides information to a V3C decoding and / or rendering system on how to perform inverse reconstruction. For example, this information includes how to subdivide the base mesh, how to apply displacement vectors to the subdivided mesh vertices, and how to apply attributes to the reconstructed mesh.
[0296] The sub-bitstreams generated in the above process (e.g., atlas sub-bitstream, base mesh sub-bitstream, geometry video (or displacement) sub-bitstream, attribute video (or texture map) sub-bitstream, packed video sub-bitstream) are finally generated as a single V3C bitstream (or V-DMC bitstream) through a multiplexer and transmitted to a receiving device.
[0297] The receiving device of Fig. 18 receives one V3C bitstream (or V-DMC bitstream), divides it into each sub-bitstream, and then performs decoding and reconstruction on each.
[0298] That is, Fig. 18 illustrates another example of a receiving device according to embodiments. In the present disclosure, the receiving device of Fig. 18 may be referred to as a mesh data receiving device or decoder or a decoder of a receiving device or a V-Mesh decoder or a dynamic mesh decoder.
[0299] FIG. 18 corresponds to the receiving device (110) or mesh video decoder (113) of FIG. 1, the decoder of FIG. 11 or FIG. 12, the receiving device of FIG. 14, and / or the receiving decoding device corresponding thereto. The receiving (decoding) operation of FIG. 18 may follow the reverse process of the corresponding process of the transmitting (encoding) operation of FIG. 15. The components of FIG. 18 may be implemented by hardware, software, a processor, and / or a combination thereof. That is, the components of the receiving device of FIG. 18 may be implemented by hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits configured to communicate with one or more memories, although not shown in the drawing. One or more processors may perform at least one or more of the operations and / or functions of the components of the receiving device of FIG. 18 described above. Additionally, one or more processors may operate or execute a set of software programs and / or instructions for performing operations and / or functions of elements of the receiving device of FIG. 18. The execution order of each block in FIG. 18 may be changed, some blocks may be omitted, and some new blocks may be added.
[0300] That is, in FIG. 18, the demultiplexer (16100) extracts an atlas sub-bitstream, a base mesh sub-bitstream, a geometry video (or displacement) sub-bitstream, an attribute video (or texture map) sub-bitstream, and a packed video sub-bitstream from the received V3C bitstream (or V-DMC bitstream). At this time, the atlas component may be divided into tiles and encapsulated into a NAL (Network Abstraction Layer) unit. In addition, the base mesh component may also be divided into sub-meshes and encapsulated into NAL units. Each sub-bitstream extracted from the demultiplexer (16100) is provided to a corresponding decoder of the decoding unit (16200). The operation of each decoder will be omitted here, with reference to the decoding description of FIG. 1, FIG. 11, FIG. 12, and / or FIG. 17.
[0301] Additional processing may be performed on the base mesh, geometry, and attributes decoded by the decoding unit (16200) in the processor (16300). For example, normal format conversion, such as normal coordinate system, bit depth, and resolution conversion, may be performed on at least one of the decoded base mesh, geometry, and attribute streams. In addition, components may be temporally aligned (atlas composition alignment), and conversion to a nominal video format or a nominal base mesh format may be performed.
[0302] To this end, the processor (16300) is provided with decoded atlas information, decoded base mesh information, decoded displacement information (if available), a set of decoded video sub-bitstreams corresponding to attributes and displacements (if available), and information from a VPS (V3C parameter set) (if available). Additional processes performed on the attribute data in the processor (16300) (i.e., map extraction, bit depth conversion, resolution conversion, map reconstruction, atlas composition alignment, attribute dimension packing, chroma-up sampling, etc.) will be described later.
[0303] In the above processor (16300), a reconstruction-related process (e.g., nominal format conversion, pre-reconstruction, reconstruction, post-reconstruction, adaptation) may be performed on the atlas, base mesh, geometry, attributes, etc. for which additional processing has been performed or not performed in the above processor (16300). That is, the mesh content is obtained through the nominal format conversion, pre-reconstruction, reconstruction, post-reconstruction, and adaptation steps using V3C components in the nominal format. For example, in the reconstruction step, the base mesh component in the nominal format, GeoFramesNF, which is a video component in the nominal format, and AttrFramesNF, if applicable, and the decoded atlas data are processed to reconstruct the mesh content and related information as outputs.
[0304] At this time, if texture map coding is omitted on the transmitting side, the map reconstruction step restores the omitted texture map based on reference information (e.g., reference texture map information) for the texture map whose coding is omitted on the transmitting side. For example, the map reconstruction step may generate the omitted texture map on the transmitting side based on the reference texture map information received and included in the signaling information. The map reconstruction step of FIG. 18 corresponds to the omitted texture map restoration unit (15022) of FIG. 17, and the present disclosure describes the process of restoring the omitted texture map in the map reconstruction step of FIG. 18 as an embodiment. Therefore, the process of restoring the omitted texture map in the omitted texture map restoration unit (15022) of FIG. 17 will refer to the process of restoring the omitted texture map in the map reconstruction step of FIG. 18.
[0305] The content related to the texture map (or attribute) in Fig. 18 is described in more detail as follows. That is, a texture map sub-bitstream (or attribute sub-bitstream) separated from a V-DMC bitstream (or V3C bitstream) is decoded through a texture map decoder of a decoding unit (16200), and attribute data can be generated as a result of the decoding.
[0306] The map extraction step extracts texture maps from the decoded attribute data. Normally, one texture map is transmitted and restored per frame, but if MultipleMapStreamsPresentFlag is non-zero, multiple texture maps may be restored per frame.
[0307] In the bit depth conversion step, a process of adjusting the bit depth of the texture map data is performed. For example, the bit depth may be adjusted to the number of digits specified according to the bit depth parameter, or the conversion may be performed to the smallest or largest bit depth.
[0308] In the resolution conversion step, the resolution of the texture map data is adjusted. The texture map resolution can be converted to a size equal to the specified width and height.
[0309] The Map Reconstruction step restores texture maps that were not transmitted due to coding omissions on the transmitting side. The process of restoring texture maps with coding omitted in the Map Reconstruction step is described in detail below.
[0310] In the atlas composition alignment step, buffering can be performed for a certain period of time based on the reference structure for each mesh element, and synchronization of mesh elements can be performed so that they can be output according to the POC.
[0311] In the attribute dimension packing step, packing can be performed for each dimension depending on which dimension the attribute data is included in to restore it to a 3D mesh.
[0312] In the Chroma Up-sampling step, the chroma up-sampling process can be performed if the chroma format of the decoded attribute data is a format other than 4:4:4 and has three components.
[0313] The following is a detailed explanation of how to restore a texture map where coding was omitted during the map reconstruction step.
[0314] For example, if the texture maps of some frames are omitted and only the remaining texture maps are transmitted, the texture maps decoded by the video decoder may be stored in the texture map buffer in a form in which some texture maps are omitted, as in FIG. 19.
[0315] That is, when the coding of the texture map of some frames is omitted as in FIG. 16, the receiving device decodes the texture map sub-bitstream through a video decoder as in FIG. 19 and then stores the decoded texture map in the texture map buffer. FIG. 19 is a diagram showing an example of a texture map decoding process in the case of omission of texture map coding according to embodiments.
[0316] For example, assuming that there are 8 frames in a GOP as in FIG. 16 and that texture map coding of 3 of these frames is skipped, only the texture maps of 5 frames are received by the receiving device, and only the texture maps of 5 frames are decoded by the video decoder and stored in a texture map buffer (e.g., DPB). Then, the texture maps of the decoded frames stored in the texture map buffer are provided to a skipped texture map derivation unit.
[0317] The texture map derivation unit parses the information on whether to skip texture map coding (texture_skip_flag) signaled and received in the SEI message and the information on the texture map reference target, and for frames in which texture map coding is skipped (e.g., three frames), the texture map information of the corresponding frame is copied as is from the reference frame according to the reference direction and stored in the texture map buffer. In the present disclosure, the information on the texture map reference target may include texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx). In the present disclosure, the reference texture frame index information (ref_texture_idx) may be referred to as reference texture map index information or index information of the reference texture map.
[0318] According to embodiments, in the present disclosure, whether texture maps of some frames in the current sequence are omitted can be determined by parsing sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI of the V3C_AD stream.
[0319] That is, when Sei_rbsp() is parsed, sei_message() is performed, and at this time, if the payloadType of sei_payload is 69, skipped_frame_indication(payloadSize) is performed, which indicates that some texture maps in the current sequence are omitted, and the related syntaxes for restoring the omitted texture maps can be parsed.
[0320] And, in order to restore the omitted texture map, the texture map derivation unit can parse the information on whether to skip texture map coding (Texture_skip_flag) and the texture map reference direction information (refDirection_idx) and / or the index information of the reference texture map (ref_texture_idx) in the transmitted SEI message. In addition, the original GoF size (num_original_gof_size) before the texture map was omitted and information on whether to continue applying the SEI message (e.g., persistence_flag) can be parsed. And, using the parsed information, the texture map information of the corresponding frame can be copied from the reference frame according to the reference direction or the reference texture map index information, thereby restoring the omitted texture map.
[0321] At this time, if the texture map reference direction information (RefDirection_idx) is bidirectional, the frame can be restored using the average value of the reference texture maps on both sides. Alternatively, the frame can be restored using the weighted average value based on the distance of the reference frames (the difference in POC (Picture Order Count) or index values). For example, a reference frame with a closer distance among bidirectional reference frames can be given more weight.
[0322] FIG. 20(a) and FIG. 20(b) are diagrams showing examples of a reconstruction process and syntax mapping according to a texture map coding omission process of the present disclosure. That is, FIG. 20(a) and FIG. 20(b) are embodiments showing texture map omission-related indexes and syntax values according to the overall encoding / decoding and map reconstruction process of a texture map coding omission method with a GOP of 8. In particular, FIG. 20(a) shows an example of a process for determining whether to omit a texture map and reference information in a texture map encoder on the transmitting side, and a video encoding process of a texture map in which coding is not omitted. Taking the transmitting device of FIG. 15 as an example, the process for determining whether to omit a texture map and reference information in FIG. 20(a) may be performed in a texture map omission decision unit (11027) or may be performed in a texture map encoder (11028).
[0323] To explain more specifically, as shown in Fig. 20(a), the texture map encoder (or texture map skip decision unit) can calculate the PSNR between the original mesh and the mesh data of each frame using a point-based metric for each frame to determine whether to skip the current texture map coding (e.g., texture_skip_flag) and the reference direction (e.g., ref_direction_idx). As shown in Fig. 20(a), when coding skipping of frames 1, 3, and 6 is determined, the values of the texture map coding skip decision information (texture_skip_flag) corresponding to frames 1, 3, and 6 are set to 1, and the values of the texture map coding skip decision information (texture_skip_flag) corresponding to the remaining frames 0, 2, 4, 5, and 7 are set to 0. Then, when the reference direction is determined, the reference direction can be signaled in the ref_direction_idx of the corresponding frame. In Figure 20(a), it can be seen that the reference direction of frame 1 is past, the reference direction of frame 3 is future, and the reference direction of frame 6 is bidirectional.
[0324] When frames for which coding is omitted (e.g., frames 1, 3, and 6) and their reference directions are determined, as in Fig. 20(a), video encoding is performed on frames for which coding is not omitted (e.g., frames 0, 2, 4, 5, and 7), as in Fig. 20(b). That is, in Fig. 20(b), encoding is performed into a texture map bitstream through a video encoder for the texture maps of the remaining frames, excluding the frames for which texture map coding omission is determined.
[0325] Meanwhile, the texture map decoder of the receiving device performs decoding on the texture map bitstream transmitted from the transmitting side. If the texture map coding of some frames is omitted and transmitted as in Fig. 20(a), the texture map decoder restores the texture maps of the remaining frames except for the texture maps of the frames for which the coding is omitted, as in Fig. 20(b). For example, if the GOP is 8 and the texture map coding of 3 frames is omitted, the texture map decoder restores the texture maps of 5 frames.
[0326] Therefore, the original GOP size may be different and the frame index may be different from the POC until the texture map reconstruction process is performed.
[0327] In order to restore this in POC order, a process of performing restoration is required along with the syntax transmitted through the SEI message (e.g. texture_skip_flag, refDirection_idx, ref_texture_idx, num_original_gof_size, persistence_flag, etc.).
[0328] In the map reconstruction step (or omitted texture map restoration part) of the present disclosure, as shown in the right figure of FIG. 20(b), restoration can be performed on texture maps whose coding has been omitted on the transmitting side from decoded video data, as shown in the left figure of FIG. 20(b), through the example method of implementing the map reconstruction algorithm below based on syntax information related to omission of texture maps transmitted through SEI messages. In Fig. 20(b) and the map reconstruction algorithm below, f is the current output frame index, skipcount is the number of frames whose coding has been skipped accumulated so far, f-skipcount is the index of the input frame excluding the frame whose coding has been skipped, inputordidx is the original order index of the input frame decoded at the receiver because the coding has not been skipped at the transmitter, refidx[f][0] is the index of the input frame to be referenced in the past direction from the current output frame, refidx[f][1] is the index of the input frame to be referenced in the future direction from the current output frame, and outordidx is the index according to the actual output order in the output frame array, i.e., the actual output order of the restored frame.
[0329] The following shows an implementation example of a map reconstruction algorithm according to the present disclosure.
[0330] Step 1 derives reference index information based on the mapping information between the input frame (i.e., decoded frame) index and the actual output frame index. That is, Step 1 calculates which input frame each output frame should refer to based on the relationship between the decoded input frame index and the actual output frame index.
[0331] In the code of Step 1, numOutFrames is the total number of output frames, skipCount is the number of frames whose coding was skipped, and texture_skip_flag[f] indicates whether the texture map coding of the corresponding frame is skipped. For example, if the value of texture_skip_flag[f] is 1, it means that the coding of the corresponding frame is skipped. In addition, ref_direction_idx[f] indicates the reference direction of the corresponding frame. For example, if the value of ref_direction_idx[f] is 0, it indicates a past direction reference, 1 indicates a future direction reference, and 2 indicates a bidirectional reference. In the code below, if the value of texture_skip_flag[f] is 1 (i.e., true), the skipCount value is increased by 1 and the frame index to be referenced (refIdx[f]) is set according to the value of ref_direction_idx[f]. In addition, if the value of texture_skip_flag[f] is 0 (i.e., false), the output frame index is set to be the same as the input frame index. In the code below, when the value of texture_skip_flag[f] is 1, outOrdIdx[f] = outOrdIdx[f - 1] + 1 specifies the order of output frames whose coding is skipped by adding +1 to the index of the previous output frame. Then, when the value of texture_skip_flag[f] is 1, outOrdIdx[f] = inputOrdIdx[f - skipCount] + skipCount calculates the order of output frames by adding the number of frames whose coding is skipped (skipCount) to the order of the input frames.
[0332] Step 1
[0333] skipCount = 0
[0334] for( f = 0; f < numOutFrames; f++ ) {
[0335] if (texture_skip_flag[f]) {
[0336] skipCount++
[0337] outOrdIdx[f] = outOrdIdx[f - 1] + 1
[0338] if (ref_direction_idx[f] == 0) {
[0339] refIdx[f][0] = inputOrdIdx[f - skipCount]
[0340] }
[0341] else if (ref_direction_idx[f] == 1) {
[0342] refIdx[f][1] = inputOrdIdx[f - skipCount] + 1
[0343] }
[0344] else { / ref_direction_idx[f] == 2
[0345] refIdx[f][0] = inputOrdIdx[f - skipCount]
[0346] refIdx[f][1] = inputOrdIdx[f - skipCount] + 1
[0347] }
[0348] }
[0349] else {
[0350] outOrdIdx[f] = inputOrdIdx[f - skipCount] + skipCount
[0351] refIdx[f][0] = inputOrdIdx[f - skipCount]
[0352] }
[0353] }
[0354] Step 2-1 derives an output frame based on the reference index (in the case of bidirectional reference - average value restoration). That is, in step 2-1, in the case of bidirectional reference, the frame whose coding is omitted on the transmitting side is restored using the pixel values of the adjacent frame based on the reference index information (refIdx[f]). For example, the restoration is performed by calculating the average of the pixel values of the previous (refIdx[f][0]) and subsequent (refIdx[f][1]) frames. That is, if the texture map reference direction information (RefDirection_idx) is bidirectional, the frame can be restored using the average value of the reference texture maps corresponding to both sides. In addition, the frame whose coding is not omitted can use the pixel values of the reference frame as is.
[0355] Step 2-1
[0356] for( f = 0; f < numOutFrames; f++ ) {
[0357] ordIdx = outOrdIdx[ f ]
[0358] for( c=0; c < iNumComp; c++ ) {
[0359] for( y=0; y < iHeight[c]; y++ ) {
[0360] for( x=0; x < iWidth[c]; x++ ) {
[0361] if(texture_skip_flag[f])
[0362] if(ref_direction_idx == 2) {
[0363] outputFrames[ordIdx][c][y][x]
[0364] = (inputFrames[ refIdx[f][0] ][ c ][ y ][ x ]
[0365] + inputFrames[ refIdx[f][1] ][ c ][ y ][ x ]) / 2
[0366] }
[0367] }
[0368] else {
[0369] outputFrames[ ordIdx ][ c ][ y ][ x ]
[0370] = inputFrames[ refIdx[f][0] ][ c ][ y ][ x ]
[0371] }
[0372] }
[0373] }
[0374] }
[0375] Step 2-2 derives the output frame based on the reference index (in the case of bidirectional reference - weighted average restoration). Step 2-2 is an example of restoring the texture map of the frame whose coding was omitted on the transmitting side based on the weighted average in the case of bidirectional reference. That is, in bidirectional reference, the texture map of the frame whose coding was omitted on the transmitting side is restored by considering the distance between the two frames. In other words, for each pixel, the relative distance (distanceWeight0, distanceWeight1) between the previous and next frames is calculated, and the average weighted by applying a weight to the pixel value of each reference frame in proportion to the calculated distance is calculated. In the code below, texture_skip_flag[prev] and texture_skip_flag[next] are checked, and the distance (distanceWeight0, distanceWeight1) to the previous frame (prev) and the next frame (next) can be calculated. In this way, in step 2-2, the frame can be reconstructed using a weighted average value based on the distance (POC (Picture Order Count) or index value difference) of the reference frame. For example, a reference frame with a closer distance among the bidirectional reference frames can be given more weight.
[0376] Step 2-2
[0377] for( f = 0; f < numOutFrames; f++ ) {
[0378] ordIdx = outOrdIdx[ f ]
[0379] for( c=0; c < iNumComp; c++ ) {
[0380] for( y=0; y < iHeight[c]; y++ ) {
[0381] for( x=0; x < iWidth[c]; x++ ) {
[0382] if(texture_skip_flag[f])
[0383] if(ref_direction_idx == 2) {
[0384] prev=f
[0385] next=f
[0386] distanceWeight0=0
[0387] distanceWeight1=0
[0388] while(texture_skip_flag[prev]) {
[0389] distanceWeight0++
[0390] prev--
[0391] }
[0392] while(texture_skip_flag[next]) {
[0393] distanceWeight1++
[0394] next++
[0395] }
[0396] outputFrames[ ordIdx ][ c ][ y ][ x ]
[0397] = (distanceWeight0
[0398] * inputFrames [ refIdx[f][1] ][ c ][ y ][ x ]
[0399] + distanceWeight1
[0400] * inputFrames [ refIdx[f][0] ][ c ][ y ][ x ] )
[0401] / (distanceWeight0 + distanceWeight1)
[0402] }
[0403] }
[0404] else {
[0405] outputFrames[ordIdx][c][y][x]
[0406] = inputFrames[ refIdx[f][0] ][ c ][ y ][ x ]
[0407] }
[0408] }
[0409] }
[0410] }
[0411] In this way, in step 1, reference index information (refIdx[f]) is generated for each output frame, in step 2-1, the texture map of the frame with omitted coding is restored through the average value of the bidirectional reference frames, and in step 2-2, the texture map of the frame with omitted coding is restored through the weighted average value of the bidirectional reference frames.
[0412] Meanwhile, in the present disclosure, the V3C bitstream can be transmitted / received in either the V3C unit stream format or the V3C sample stream format. The V-DMC bitstream of the present disclosure can follow the V3C bitstream structure defined in the V3C codec specification (ISO / IEC 23090-5) to be described later. In this case, the existing V3C bitstream structure may be followed, but some V3C units may not be used, and some structures for V-DMC only, such as V-DMC extensions, may be followed.
[0413] FIG. 21 is a drawing showing an example of a V3C bitstream structure according to embodiments, and is an example of a V3C bitstream in a V3C unit stream format.
[0414] In Fig. 21, the V3C bitstream can be composed of V3C sample stream precision information and multiple sample stream V3C units. Each sample stream V3C unit is composed of V3C sample stream size information and a V3C unit. The V3C unit is again composed of a V3C unit header (V3C_unit_header) and a V3C unit payload (V3C_unit_payload). The V3C sample stream precision information can represent the precision of the V3C sample stream size information in all sample stream V3C units in bytes.
[0415] The above V3C sample stream size information specifies the size of the subsequent V3C unit in bytes.
[0416] The above V3C unit header includes type information (vuh_unit_type) indicating the type of data carried by the corresponding V3C unit payload. The V3C unit payload carries one of V3C parameter set data, atlas data, base mesh data, displacement data, occupancy video data, geometry video data, attribute video data, common atlas data, or packed video data according to the type information (vuh_unit_type). For example, if the type information (vuh_unit_type) of the V3C unit header indicates a V3C parameter set (V3C_VPS), the V3C unit payload includes a V3C parameter set (v3c_parameter_set()) including overall encoding information of a bitstream, and if it indicates atlas data (V3C_AD), the V3C unit payload includes an atlas sub-bitstream (atlas_sub_bitstream()) carrying the atlas data. And, in one embodiment, if the type information (vuh_unit_type) of the V3C unit header indicates accumuncity video data (V3C_OVD), the V3C unit payload includes an accumuncity video sub-bitstream (video_sub_bitstream()) carrying the accumuncity video data, if it indicates geometry video data (V3C_GVD), it includes a geometry video sub-bitstream (video_sub_bitstream()) carrying the geometry video data, and if it indicates attribute video data (V3C_AVD), it includes an attribute video sub-bitstream (video_sub_bitstream()) carrying the attribute video data.In addition, if the type information (vuh_unit_type) of the V3C unit header indicates base mesh data (V3C_BMD), the V3C unit payload may include a base mesh sub-bitstream carrying base mesh data, and if it indicates displacement data (V3C_DD), the V3C unit payload may include a displacement sub-bitstream carrying displacement data (Displacements Data).
[0417] Here, the V3C parameter set (V3C_VPS) may include parameter set information such as decoder configuration information related to mesh encoding / decoding and a sequence header. The atlas data (V3C_AD) may include additional information such as 2D mapping or texture mapping for a 3D object. In addition, the geometry video data (V3C_GVD) is displacement information compressed by a video codec. In the present disclosure, the geometry video data may be referred to as displacement video data. The attribute video data (V3C_AVD) is attribute or texture information compressed by a video codec. In addition, the base mesh data (V3C_BMD) is compressed base mesh information for mesh encoding / decoding. The displacement data (V3C_DD) is displacement information compressed by arithmetic coding. However, in the case of lossless encoding, displacement data (or displacement information) may not be included.
[0418] As described above, if the type information (vuh_unit_type) of the V3C unit header indicates atlas data (V3C_AD), the corresponding V3C unit payload includes an atlas sub bitstream (atlas_sub_bitstream()) carrying the atlas data.
[0419] At this time, the atlas sub bitstream (atlas_sub_bitstream()) carries the atlas data in NAL sample stream format.
[0420] In the present disclosure, the NAL sample stream carrying atlas data is composed of NAL unit size precision information and multiple sample stream NAL units, similar to the V3C sample stream. Each sample stream NAL unit is composed of NAL unit size information and a NAL unit. The NAL unit is further composed of a NAL unit header (NAL_unit_header) and a NAL unit payload (NAL_unit_payload).
[0421] The above NAL unit size precision information can indicate the precision of NAL unit size information in all sample stream NAL units in bytes.
[0422] The above NAL unit size information specifies the size of the subsequent NAL unit in bytes.
[0423] The NAL unit header includes type information indicating the type of data carried by the corresponding NAL unit payload. The NAL unit payload carries one of an atlas sequence parameter set (ASPS), an atlas adaptation parameter set (AAPS), an atlas frame parameter set (AFPS), a prefix essential SEI message, atlas tile layer information, and a suffix essential SEI message according to the type information of the NAL unit header. In the present disclosure, the prefix essential SEI message and the suffix essential SEI message are transmitted via sel_rbsp(). Here, rbsp (Raw Byte Sequence Payload) is specified as a sequential sequence of bytes.
[0424] In the present disclosure, if the type information of the NAL unit header indicates NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI, the corresponding NAL unit payload carries texture map coding skip information (texture_skip_flag) and texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx) through sel_rbsp().
[0425] Fig. 22 is a table defining NAL types and conformance types for each SEI message purpose according to embodiments. That is, in order to signal information for restoring a texture map for which coding has been omitted using an SEI message, an SEI message, a NAL type, and a conformance type can be defined as shown in Fig. 22.
[0426] In this disclosure, SEI messages provide additional information for performing processes related to decoding, reconstruction, display, etc., and are divided into essential SEI and non-essential SEI types.
[0427] Non-essential SEI messages are types of information that are not required in the decoding process, and essential SEI messages are types of information that are an essential part of the V3C bitstream and must not be removed.
[0428] And, Type A is an ESEI message of the type required for conformance point A of Fig. 19, and Type B is an ESEI message of the type required for conformance point B of Fig. 19.
[0429] Whether to skip texture map coding of the present disclosure and reference information for deriving skipped frames (e.g., (texture_skip_flag, refDirection_idx, ref_texture_idx)) can be transmitted via an SEI message. According to embodiments, whether to skip texture map coding and reference information for deriving skipped frames (e.g., (texture_skip_flag, refDirection_idx, ref_texture_idx)) can be transmitted in the form of a skipped frame indication SEI message as in FIG. 22. The NAL type of the skipped frame indication SEI message may be NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI. In addition, the Conformance Type of the skipped frame indication SEI message may be Type-A or Type-B.
[0430] In this way, the present disclosure can signal texture map coding skip information (texture_skip_flag) and texture map reference direction information (refDirection_idx) and / or reference texture frame index information (ref_texture_idx) through the SEI message (sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI) of the NAL sample stream in the V3C_AD (Atlas Data) atlas sub-bitstream of the V3C sample stream.
[0431] Fig. 23 is a diagram showing an example of NAL unit semantics according to embodiments. In Fig. 23, nal_unit_type indicates type information included in the NAL unit header. The nal_unit_type field indicates the type of the RBSP data structure included in the NAL unit.
[0432] In Fig. 23, if the value of nal_unit_type is 43 (i.e., NAL_PREFIX_NSEI) or 44 (i.e., NAL_SUFFIX_NSEI), a non-essential SEI is included in the NAL unit, the RBSP syntax structure of the NAL unit is sei_rbsp(), and the type class of the NAL unit is non-ACL (Atlas Coding Layer).
[0433] And, if the value of nal_unit_type is 45 (i.e., NAL_PREFIX_ESEI) or 46 (i.e., NAL_SUFFIX_ESEI), essential SEI is included in the NAL unit, the RBSP syntax structure of the NAL unit is sei_rbsp(), and the type class of the NAL unit is non-ACL.
[0434] Fig. 24 is a diagram showing an example of a syntax structure of sei_rbsp() including an SEI message according to embodiments. In Fig. 24, sei_message() carries an SEI message including reference information (e.g., (texture_skip_flag, refDirection_idx, ref_texture_idx)) for indicating whether to skip texture map coding and for deriving skipped frames.
[0435] FIG. 25 is a diagram showing an example of the syntax structure of sei_message() according to embodiments.
[0436] In Fig. 25, after determining the type (payloadType) and size (payloadSize) of the payload, sei_payload(payloadType, payloadSize) is called to process the payload data. Here, the type of the payload (payloadType) indicates the type of the SEI message, and the size of the payload (payloadSize) indicates the payload size of the SEI message. The called sei_payload(payloadType, payloadSize) carries the actual SEI message based on the determined payloadType and payloadSize.
[0437] FIG. 26 is a diagram showing an example of the syntax structure of an SEI message payload (sei_payload(payloadType, payloadSize)) according to embodiments.
[0438] In one embodiment, the SEI message payload may include sei(payloadSize) according to PayloadType if the NAL unit type (nal_unit_type) is NAL_PREFIX_NSEI or NAL_PREFIX_ESEI.
[0439] In another embodiment, the SEI message payload may include sei(payloadSize) according to PayloadType if the NAL unit type (nal_unit_type) is NAL_SUFFIX_NSEI or NAL_SUFFIX_ESEI.
[0440] For example, if the NAL unit type (nal_unit_type) is NAL_PREFIX_NSEI or NAL_PREFIX_ESEI and the payloadType is 69, the SEI message payload may contain a skipped_frame_indication(payloadSize).
[0441] Fig. 27 is a diagram showing an example of the syntax structure of skipped_frame_indication(payloadSize) according to embodiments. That is, Fig. 27 shows an example of the syntax and semantics of the Skipped_frame_indication SEI message.
[0442] In Fig. 27, persistence_flag is a flag indicating that the SEI message persists for the current texture map. For example, if the value of persistence_flag is 0, it means that the SEI message is applied only for the currently decoded texture map frame, and if the value of persistence_flag is 1, it means that the SEI message persists for the current texture map in output order until one of various conditions is satisfied.
[0443] The following are examples of the various conditions mentioned above: when a new CAS (Coded Atlas Sequence) begins, and when the bitstream ends.
[0444] num_original_gof_size represents the size of the original GoF (Group of Frames) (before the texture map was omitted).
[0445] texture_skip_flag[j] is a flag indicating whether the texture map of the j-th frame is skipped. If the value of texture_skip_flag[j] is 1, it means that the coding and transmission of the j-th texture map are skipped, and if the value of texture_skip_flag[j] is 0, it means that the coding and transmission of the ith texture map are not skipped.
[0446] ref_direction_idx[j] represents the reference direction information of the texture map of the j-th frame. If the value of ref_direction_idx[j] is 0, it means that the skipped frame is derived by referencing the past frame, and if the value of ref_direction_idx[j] is 1, it means that the skipped frame is derived by referencing the future frame. In addition, if the value of ref_direction_idx[j] is 2, it means that the skipped frame is derived using a bidirectional texture map. If the value of ref_direction_idx[j] is 3, it means that the j-th texture map is not omitted.
[0447] The following is a more specific description of the operation of the receiving device of FIG. 18 when the skipped_frame_indication(payloadSize) is transmitted from the transmitting device to the receiving device via the SEI message payload as described above.
[0448] That is, the receiving device of FIG. 18 receives the V-DMC (or V3C) bitstream transmitted from the transmitting device, separates the base mesh sub-bitstream (or bitstream), the displacement vector sub-bitstream (or bitstream), and the texture map sub-bitstream (or bitstream) through a demultiplexer, and then performs a process of decoding the separated base mesh bitstream, displacement vector bitstream, and texture map bitstream, respectively.
[0449] First, the base mesh sub-bitstream is decoded into a motion vector, which is the difference between base meshes, through a motion vector decoder in the case of an inter-frame, and then decoded into a base mesh by adding the decoded motion vector value to the reference base mesh. In the case of an intra-frame, it is decoded into a base mesh through a static mesh decoder.
[0450] The displacement vector sub-bitstream decodes the displacement vector coefficients in the reverse order of encoding of the transmitting device, performs inverse quantization and inverse transformation, and then is inversely transformed from the local coordinate system to the Cartesian coordinate system to be restored to the final displacement vector. The mesh restoration unit calculates the vertex geometry information of the restored mesh by adding the displacement vector restored by the mesh restoration unit to the vertices generated through the subdivision process in the mesh refinement unit, thereby restoring the final geometry information.
[0451] The texture map sub-bitstream is restored to attribute data through the texture map decoder. After that, the map extraction unit extracts the texture map from the decoded attribute data. In addition, bit depth conversion and resolution conversion processes for data alignment are performed on the extracted data. Then, in order to perform the map reconstruction process, that is, the map reconstruction unit (or skipped texture map restoration unit (15022)) first parses sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI of V3C_AD. When sei_rbsp() is parsed, sei_message() is performed, and at this time, if the payloadType of sei_payload is 69, skipped_frame_indication (payloadSize) is performed for restoration of the texture map with skipped coding. The map reconstruction unit can parse the persistence_flag in the skipped_frame_indication to decide whether to apply the SEI message only for the currently decoded texture map frame, or to persist the SEI message until one of the conditions is met (such as before a new atlas sequence starts or until the bitstream ends). In addition, the map reconstruction unit can parse the num_original_gof_size to get the original size information of the current sequence (GOF), and infer the number of skipped texture maps from the difference with the number of currently decoded texture maps. The map reconstruction unit can determine whether the texture map coding of the current frame is skipped by parsing the texture_skip_flag for each frame while the current sequence is the original size (i.e., repeating as many times as the original size), and if the value of texture_skip_flag is 1, the map reconstruction unit can identify the direction that the skipped texture map referenced by parsing the ref_direction_idx.For example, if the value of ref_direction_idx is 0, it indicates that the skipped frame is derived by referencing the past frame, and if the value of ref_direction_idx is 1, it indicates that the skipped frame is derived by referencing the future frame. If the value of ref_direction_idx is 2, it indicates that the skipped frame is derived using a bidirectional texture map. If the value of ref_direction_idx is 3 or absent, it means that the texture map of the current frame is not skipped. In particular, in the case of referencing a bidirectional texture map, the texture map can be restored with the intermediate value of the past and future frames. In addition, the map reconstruction unit can calculate the POC (or index) difference between frames and perform texture map restoration with a weighted average according to the ratio of the reference texture map distances in order to restore a slightly higher quality texture map.
[0452] Afterwards, in the Atlas composition alignment section, buffering can be performed for a certain period of time according to the reference structure of each mesh element, and synchronization of the mesh elements can be performed so that the output can be matched to the POC. In addition, attribute dimension packing and chroma upsampling processes can be performed according to the chroma format. Once the attribute data is finally restored through the above processes, the final restoration mesh is generated together with the finally restored geometry information through the reconstruction process.
[0453] As mentioned above, among the mesh data components, texture maps account for the largest data proportion. The present disclosure can achieve a significant bit-saving effect by omitting the coding and transmission of such texture maps on a frame-by-frame basis without a significant difference in quality. In particular, the present disclosure proposes an encoding and decoding process that omits the coding of texture maps in a V-DMC codec when the current texture map and the reference texture map are similar, and a restoration method for the omitted texture maps. In particular, in a bidirectional reference texture map restoration method, each POC difference between the omitted texture map and the past / future reference texture maps is calculated, and a weighted average reference method according to the distance is proposed. This method can restore the omitted texture map to a texture value more similar to the original than a method that restores the omitted texture map with the average value of the bidirectional reference texture map. In addition, in order to apply the texture map omission method to the V-DMC technology, a method of transmitting syntax (i.e., information related to texture map restoration) through an SEI message based on a V3C bitstream is proposed. In this disclosure, a method for signaling information (e.g., information on whether texture map coding is omitted and reference direction information, etc.) required for restoration of a texture map whose coding is omitted at the transmitting side is proposed through NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI of V3C_AD is proposed. In this disclosure, by omitting texture map coding on a frame-by-frame basis through the proposed method and signaling information about the omitted texture map according to the V-DMC (V3C) bitstream structure, efficient compression can be performed, thereby obtaining a great effect in terms of bit savings and transmission speed.
[0454] So far, we have described a method for signaling information on whether to omit texture map coding (e.g., a flag) and reference information for deriving omitted texture maps through atlas data, especially SEI messages, and for restoring omitted texture maps at a receiving device based on this.
[0455] Hereinafter, as another embodiment of the present disclosure, a method for signaling information on whether texture map coding is omitted (e.g., a flag) and reference information for deriving an omitted texture map through a V3C unit header and restoring the omitted texture map in a receiving device based on the signal will be described. In this case, as one embodiment, the receiving device of the present disclosure restores the omitted texture map based on a multi-reference structure.
[0456] That is, in the receiving device of the present disclosure, when a frame with an omitted texture map is transmitted by applying a texture map coding omission technique in a V-DMC codec, the omitted texture map can be restored based on a multi-reference structure. In addition, as a method of signaling the related syntax for restoring the omitted texture map through a V3C unit header, the existing V3C_AVD (Attribute Video Data) can be used and / or V3C_TVD (Texture map Video Data) can be used.
[0457] As described in FIGS. 15 and 16, the transmitting device of the present disclosure can most effectively increase the compression efficiency of the V-DMC codec by increasing the compression efficiency of texture map data by omitting texture map coding and transmission of some frames. The present disclosure proposes a method for restoring the omitted texture map based on multiple references in a receiving device when the texture map coding omission technique is applied in the transmitting device. In addition, the present disclosure requires an efficient signaling method to transmit required information even when the texture map coding omission technique is applied. To this end, the present disclosure proposes a method for transmitting related syntaxes through a V3C unit header. In the current V-DMC structure, texture map information is signaled through V3C_AVD. However, since there is no extension exclusively for V-DMC, there is a problem that in order to apply the texture map coding omission technique, the V3C specification must be directly modified or a new V3C unit type must be required. To solve this problem, the present disclosure proposes a method of signaling through V3C_AVD of an existing V3C unit header and a method of creating a new V3C unit type, V3C_TVD, and signaling through it.
[0458] In this way, the present disclosure proposes a method for restoring frames of omitted texture maps based on a multi-reference structure, and a signaling method for a flag for whether texture map coding is omitted and reference direction information syntaxes for deriving omitted frames through a V3C unit header. In particular, the present disclosure proposes a V3C_TVD (Texture map Video Data) structure, which is a new V3C unit type for texture maps other than the V3C_AVD (Attribute Video Data) structure of the existing V3C unit header.
[0459] At this time, since the method of omitting the coding of the texture map in the transmitting device is the same and only the signaling method is different, a detailed description will be made by referring to the description of FIGS. 15 and 16 described above, and a brief description will be given here to avoid redundant description.
[0460] That is, in the V-DMC transmitter, the original mesh is generated as a base mesh through a mesh simplification unit (11011) and a mesh parameterization unit (11012). The generated base mesh undergoes quantization in a mesh quantization unit (11013), and then, if the current base mesh is of inter type, a motion vector is calculated from a previous reference restoration base mesh through a motion vector encoder (11015) and the motion vector is encoded. If the current base mesh is of intra type, encoding is performed through a static mesh encoder (11016) to generate a base mesh sub-bitstream.
[0461] And, the displacement vector calculation unit (11020) calculates a displacement vector which is the difference between the mesh data obtained by subdividing (11018) and fitting (11019) the simplified mesh through the mesh simplification unit (11011) and the mesh data (11017) restored from the previously encoded base mesh. At this time, in order to efficiently encode the calculated displacement vector, the displacement vector coordinate system conversion unit (11021) can convert the displacement vector coordinate system into a local coordinate system. The displacement vector encoder (11022) converts the displacement vector into a displacement vector coefficient and performs quantization and packing, and encodes the packed 2D image through a 2D video encoder such as H.264, HEVC, or VVC to generate a displacement vector sub-bitstream.
[0462] Meanwhile, in the texture map generation unit (11026), a process of generating a new texture map having color information corresponding to the texture coordinates of the mesh restored in the mesh restoration unit (11025) is performed. The texture map omission decision unit (11027) can determine whether to omit texture map coding of the current frame by calculating the difference in texture PSNR between the mesh of the current frame and the original mesh and the texture PSNR between the mesh of the reference frame and the original mesh using a method such as a point-based metric. When the texture map omission decision unit (11027) determines that the texture map coding of the current frame is omitted, the relevant syntaxes for texture map restoration can be signaled through the V3C unit header. Here, the signaling may be performed by the texture map omission decision unit (11027) or the texture map encoder (11028), or may be performed by a separate block or module. At this time, the transmitted V3C unit type can be transmitted using the existing V3C_AVD (Attribute Video Data, vuh_unit_type=4), and a new V3C unit type can be generated and transmitted through V3C_TVD (Texture map Video Data, vuh_unit_type=7). In the present disclosure, related syntaxes for performing texture map restoration may include information on whether texture map coding is skipped (i.e., flag information) (texture_skip_flag), information on the total number of output texture maps for deriving the number of skipped texture maps (vuh_output_texturemap_count), information on the maximum number of reference frames (vuh_ref_max_num), information on reference frame index (vuh_ref_frame_idx), and information on the weight of the reference frame (vuh_recon_weight). The texture map encoder (11028) may output H.Encoding is performed using a 2D video encoder such as 264, HEVC, or VVC to generate a texture sub-bitstream.
[0463] According to embodiments, a multiplexer (not shown) may multiplex the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream generated through the entire process into a single bitstream and then transmit the multiplexer to a receiving device. That is, a sub-bitstream may be generated for each component and transmitted as a single bitstream from the transmitting end through the multiplexer. Alternatively, the base mesh bitstream, the displacement vector bitstream, and the texture map bitstream may be encapsulated into a file / segment and transmitted to the receiving device. Here, the bitstream may be referred to as a sub-bitstream.
[0464] Although the description is omitted here, the sub-bitstreams (or bitstreams) generated through the entire process above may include an atlas sub-bitstream, a base mesh sub-bitstream, a geometry video (or displacement) sub-bitstream, an attribute video (or texture map) sub-bitstream, and a packed video sub-bitstream, and these sub-bitstreams may be finally generated as a single V3C bitstream (or V-DMC bitstream) through a multiplexer and transmitted to a receiving device.
[0465] In this case, a receiving device such as FIG. 17 or FIG. 18 separates each sub-bitstream from the V3C bitstream (V-DMC bitstream), performs decoding on each separated sub-bitstream, and also performs restoration of a texture map for which coding has been omitted.
[0466] The present disclosure will describe restoration of a texture map with decoding and coding of each sub-bitstream omitted, with reference to FIG. 18. The texture map restoration process will be described here, and the remainder will refer to the description of FIG. 18 described above.
[0467] That is, in the map reconstruction section of Fig. 18, if the texture map is coded with the texture map omitted, restoration is performed for the omitted texture map.
[0468] At this time, when texture map coding is omitted, the transmitting device performs omission according to the following principle, and the V-DMC encoder of the transmitting device can signal the relevant parameters.
[0469] That is, the PSNR is calculated using a point-based metric or the like for the mesh data of the original mesh and the current frame, and the PSNR is calculated in the same way for the mesh data of the original mesh and the reference frame. If the difference between the two PSNRs is below a certain threshold, the two frames are judged to be similar and texture map coding can be omitted.
[0470] And, when it is decided to skip texture map coding of the current frame, information on whether to skip texture map coding (vuh_texture_skip_flag), reference texture map frame index information (vuh_ref_frame_idx), and weight information (vuh_recon_weight) can be signaled. In addition, information on the maximum number of reference texture map frames (vuh_ref_max_num) can be signaled.
[0471] At this time, if there is only one reference texture map frame, such as in unidirectional prediction, the vuh_ref_max_num value can be signaled as 1. In addition, even in bidirectional prediction or unidirectional prediction, if there are two frames in the same direction, the vuh_ref_max_num value can be signaled as 2. In addition, even in bidirectional prediction or unidirectional prediction, if there are two or more frames in the same direction, the vuh_ref_max_num value can be signaled as 3 or more. However, information on the maximum number of reference texture map frames can be signaled according to the computational complexity of the encoder or according to the desired transmission speed or video quality.
[0472] The present disclosure can transmit reference frame index information (vuh_ref_frame_idx) for each frame in accordance with the size of the maximum number of reference texture map frames information (vuh_ref_max_num).
[0473] For example, if the size of the above vuh_ref_max_num is 3, the number of vuh_ref_frame_idx of frames from which texture maps are omitted may all be the same as 3. In this case, the vuh_ref_frame_idx values of frames from which texture maps are omitted may have different values. If the number of frames to be referenced is less than 3, they may have the same reference frame index value, or they may be assigned a specific value by agreement with the encoder and may not have the meaning of a reference frame index. Alternatively, the reference frame index information (vuh_ref_frame_idx) for each frame may be equal to or less than the size of vuh_ref_max_num.
[0474] In the present disclosure, assuming that the size of vuh_ref_max_num is 3, the number of vuh_ref_frame_idx values of frames in which texture map coding is omitted may be equal to or less than 3. The vuh_ref_frame_idx values may not be the same, such as 1 or 2, for each frame in which texture map coding is omitted.
[0475] Additionally, the present disclosure can signal original Group of Frames (GOF) size information (vuh_output_texturemap_count) to calculate how much texture map coding omission has been performed during the current sequence. In this case, the receiving device can derive the number of frames with omitted texture maps from the difference between the original GOF size information and the decoded frame.
[0476] The transmitting device of the present disclosure may directly transmit frame number information with omitted texture maps instead of transmitting GOF size information.
[0477] In one embodiment, the information described above is transmitted to the receiving device via the V3C unit header.
[0478] When signaling is performed as described above, the receiving device (e.g., the map reconstruction unit of FIG. 18 or the omitted texture map restoration unit of FIG. 17) can restore the omitted texture map according to the following principle.
[0479] 1) If the omitted texture map references a texture map of one frame.
[0480] This is when the value of vuh_ref_max_num is 1, the value of the received vuh_ref_frame_idx is 1, and the value of vuh_recon_weight is 1.
[0481] In this case, the texture map of the frame corresponding to vuh_ref_frame_idx can be copied as is to restore the texture map whose coding was omitted on the transmitting side.
[0482] 2) If the omitted texture map refers to texture maps of more than two frames and is restored with the average value.
[0483] This is when the value of vuh_ref_max_num is 2 or more, the value of vuh_ref_frame_idx is 2 or more, and the value of vuh_recon_weight is 1.
[0484] In this case, the texture map whose coding was omitted on the transmitting side can be restored as the average value of the texture maps of the frame corresponding to vuh_ref_frame_idx.
[0485] 3) If the omitted texture map refers to texture maps of more than two frames and is calculated and restored based on each weight.
[0486] This is when the value of vuh_ref_max_num is 2 or more, the value of vuh_ref_frame_idx is 2 or more, and the vuh_recon_weight value is assigned and transmitted respectively.
[0487] In this case, the texture map of the frame corresponding to vuh_ref_frame_idx can be multiplied by the weight vuh_recon_weight value and divided by the number of referenced frames to restore the texture map whose coding was omitted on the transmitting side as a weighted average value.
[0488] According to embodiments, the vuh_ref_frame_idx and vuh_recon_weight values may be the results of optimal values derived from the V-DMC encoder using a point-based metric or the like for texture map similarity within each sequence.
[0489] Additionally, vuh_ref_frame_idx can have values ranging from 1 to the total number of frames in the sequence, and a maximum number can be set for encoding / decoding efficiency. Furthermore, information about this maximum number can be signaled as vuh_ref_max_num.
[0490] In the present disclosure, the V-DMC bitstream may basically follow the bitstream structure defined in the V3C codec specification (ISO / IEC 23090-5) as shown in FIG. 21. In particular, the present disclosure may follow the existing V3C bitstream structure, but may not use some V3C units, or may partially follow a structure for V-DMC only, such as a V-DMC extension.
[0491] In Fig. 21, the V3C bitstream can be composed of V3C sample stream precision information and multiple sample stream V3C units as described above. In addition, each sample stream V3C unit is composed of V3C sample stream size information and a V3C unit. In addition, the V3C unit is again composed of a V3C unit header (V3C_unit_header) and a V3C unit payload (V3C_unit_payload).
[0492] The above V3C unit header includes type information (vuh_unit_type) indicating the type of data carried by the corresponding V3C unit payload. The V3C unit payload can carry one of V3C parameter set data, atlas data, base mesh data, displacement data, occupancy video data, geometry video data, attribute video data, common atlas data, or packed video data according to the type information (vuh_unit_type). For example, if the type information (vuh_unit_type) of the V3C unit header indicates a V3C parameter set (V3C_VPS), the V3C unit payload includes a V3C parameter set (v3c_parameter_set()) including overall encoding information of the bitstream, and if it indicates atlas data (V3C_AD), the V3C unit payload includes an atlas sub-bitstream (atlas_sub_bitstream()) carrying the atlas data. And, in one embodiment, if the type information (vuh_unit_type) of the V3C unit header indicates accumuncity video data (V3C_OVD), the V3C unit payload includes an accumuncity video sub-bitstream (video_sub_bitstream()) carrying the accumuncity video data, if it indicates geometry video data (V3C_GVD), it includes a geometry video sub-bitstream (video_sub_bitstream()) carrying the geometry video data, and if it indicates attribute video data (V3C_AVD), it includes an attribute video sub-bitstream (video_sub_bitstream()) carrying the attribute video data.In addition, if the type information (vuh_unit_type) of the V3C unit header indicates base mesh data (V3C_BMD), the V3C unit payload may include a base mesh sub-bitstream carrying base mesh data, and if it indicates displacement data (V3C_DD), the V3C unit payload may include a displacement sub-bitstream carrying displacement data (Displacements Data) (not shown).
[0493] Here, the V3C parameter set (V3C_VPS) may include parameter set information such as decoder configuration information related to mesh encoding / decoding and a sequence header. The atlas data (V3C_AD) may include additional information such as 2D mapping or texture mapping for a 3D object. In addition, the geometry video data (V3C_GVD) is displacement information compressed by a video codec. In the present disclosure, the geometry video data may be referred to as displacement video data. The attribute video data (V3C_AVD) is attribute or texture information compressed by a video codec. In addition, the base mesh data (V3C_BMD) is compressed base mesh information for mesh encoding / decoding. The displacement data (V3C_DD) is displacement information compressed by arithmetic coding. However, in the case of lossless encoding, displacement data (or displacement information) may not be included.
[0494] In one embodiment, the present disclosure may signal information for performing texture map restoration in a V3C unit header of a V3C unit carrying attribute video data (e.g., information on whether to skip texture map coding (i.e., flag information) (texture_skip_flag), information on the total number of output texture maps for deriving the number of skipped texture maps (vuh_output_texturemap_count), information on the maximum number of reference frames (vuh_ref_max_num), information on reference frame index (vuh_ref_frame_idx), and information on weight of reference frames (vuh_recon_weight), etc.).
[0495] In another embodiment, the present disclosure newly defines type information (vuh_unit_type) of a V3C unit header (V3C_TVD), and information for performing texture map restoration (e.g., information on whether to skip texture map coding (i.e., flag information) (texture_skip_flag), total output texture map count information for deriving the number of skipped texture maps (vuh_output_texturemap_count), maximum reference frame count information (vuh_ref_max_num), reference frame index information (vuh_ref_frame_idx), and reference frame weight information (vuh_recon_weight), etc.) can be signaled in the V3C unit header of the newly defined type information (vuh_unit_type). At this time, the V3C unit payload of the corresponding V3C unit can carry texture map information (or texture map data) compressed by a video codec.
[0496] Figure 28 shows examples of the types of V3C units assigned to the vuh_unit_type field.
[0497] Referring to FIG. 28, if the value of the vuh_unit_type field is 0, it indicates that the data included in the V3C unit payload of the corresponding V3C unit is V3C parameter set (V3C_VPS), if it is 1, it indicates that it is atlas data (V3C_AD), if it is 2, it indicates that it is accumulator video data (V3C_OVD), if it is 3, it indicates that it is geometry video data (V3C_GVD), if it is 4, it indicates that it is attribute video data (V3C_AVD), if it is 5, it indicates that it is packed video data (V3C_PVD), if it is 6, it indicates that it is common atlas data (V3C_CAD), and if it is 7, it indicates that it is texture map video data (V3C_TVD). Since the meaning, order, deletion, addition, etc. of the values assigned to the vuh_unit_type field in the present disclosure can be easily changed by those skilled in the art, the present disclosure will not be limited to the above embodiment.
[0498] In this way, in the present disclosure, texture map data can be transmitted via V3C_AVD and / or V3C_TVD. That is, in the former case, it directly follows the V3C codec specification (ISO / IEC 23090-5) and there is no extension for V-DMC only. Therefore, the present disclosure proposes a structure for adding syntax related to texture map restoration that is omitted from V3C_AVD in the V3C codec specification, as follows.
[0499] That is, the present disclosure creates a new V3C unit type called V3C_TVD to transmit texture map restoration related information (whether texture map coding is omitted, texture map reference information, etc.), and through this, texture map restoration related information can be signaled. In this way, the present disclosure creates V3C_TVD, which is a V3C unit type for V-DMC technology, without significantly modifying the existing V3C codec specification, and adds omitted texture map restoration related syntax (or texture map restoration related information) to the V3C unit header corresponding to V3C_TVD. In other words, whether texture map coding is omitted and reference information for deriving omitted frames can be included in the V3C unit header and transmitted. At this time, depending on the V3C unit type (vuh_unit_type), that is, the type information (vuh_unit_type) of the V3C unit header, for example, if the V3C unit type is V3C_AVD, it can be transmitted in the V3C_AVD unit header syntax, and if the V3C unit type is V3C_TVD, it can be transmitted in the V3C_TVD unit header syntax.
[0500] Fig. 29 shows an example of the syntax structure of a V3C unit header according to embodiments. In one embodiment, the V3C unit header (v3c_unit_header()) of Fig. 29 includes a vuh_unit_type field. The vuh_unit_type field indicates the type of the corresponding V3C unit. The vuh_unit_type field according to embodiments is also referred to as a v3c_unit_type field.
[0501] A V3C unit header according to embodiments may include a vuh_v3c_parameter_set_id field if the vuh_unit_type field indicates attribute video data (V3C_AVD), geometry video data (V3C_GVD), accumulative video data (V3C_OVD), atlas data (V3C_AD), common atlas data (V3C_CAD), or packed video data (V3C_PVD).
[0502] The above vuh_v3c_parameter_set_id field specifies the identifier (i.e., vuh_v3c_parameter_set_id) of the active V3C parameter set (V3C VPS).
[0503] The V3C unit header according to embodiments may further include a vuh_atlas_id field if the vuh_unit_type field indicates attribute video data (V3C_AVD), geometry video data (V3C_GVD), accumulator video data (V3C_OVD), atlas data (V3C_AD), or packed video data (V3C_PVD).
[0504] The above vuh_atlas_id field specifies the index of the atlas corresponding to the current V3C unit.
[0505] The V3C unit header according to embodiments may further include a vuh_attribute_index field, a vuh_attribute_partition_index field, a vuh_map_index field, a vuh_auxiliary_video_flag field, and a vuh_output_texturemap_count field when the vuh_unit_type field indicates attribute video data (V3C_AVD).
[0506] The above vuh_attribute_index field indicates the index of attribute video data carried as an attribute video data unit.
[0507] The above vuh_attribute_partition_index field indicates the index of the attribute dimension group carried as an attribute video data unit.
[0508] The above vuh_map_index field, if present, indicates the index of the current attribute stream.
[0509] If the value of the vuh_auxiliary_video_flag field is 1, it may indicate that the associated attribute video data unit contains only raw and / or EOM (Enhanced Occupancy Mode) coded points. As another example, if the value of the vuh_auxiliary_video_flag field is 0, it may indicate that the associated attribute video data unit may contain raw and / or EOM coded points. If the vuh_auxiliary_video_flag field does not exist, the value of the field may be inferred to be equal to 0. In some embodiments, raw and / or EOM coded points may also be referred to as PCM (Pulse Code Modulation) coded points.
[0510] The above vuh_output_texturemap_count field indicates the number of output texture maps.
[0511] A loop that repeats as many times as the value of the vuh_output_texturemap_count field may include a vuh_texture_skip_flag[i] field and a vuh_ref_max_num field.
[0512] The above vuh_texture_skip_flag[i] field is a flag indicating whether the texture map of the i-th frame is omitted. If the value of the vuh_texture_skip_flag[i] field is 1, it means that the i-th texture map is omitted, and if the value of the vuh_texture_skip_flag[i] field is 0, it means that the i-th texture map is not omitted.
[0513] The above vuh_ref_max_num field indicates the maximum number of reference frames for omitted texture map frames.
[0514] The loop that is added when the value of the above vuh_texture_skip_flag[i] field is 1 is repeated as many times as the value of the vuh_ref_max_num field, and may include the vuh_ref_frmae_idx[i][j] field and the vuh_recon_weight[i][j] field.
[0515] The above vuh_ref_frmae_idx[i][j] field indicates reference frame index information of the i-th texture map. At this time, up to j frames can be referenced according to the maximum number of reference frames.
[0516] The above vuh_recon_weight[i][j] field specifies weight information of the jth reference frame of the ith texture map.
[0517] Fig. 30 shows another example of the syntax structure of a V3C unit header according to embodiments. Fig. 30 is a syntax structure when the V3C unit type is V3C_TVD, that is, when information on whether texture map coding is omitted and reference information is transmitted through the V3C unit header of V3C_TVD.
[0518] A V3C unit header according to embodiments may include a vuh_v3c_parameter_set_id field if the vuh_unit_type field indicates attribute video data (V3C_AVD), geometry video data (V3C_GVD), accumulator video data (V3C_OVD), atlas data (V3C_AD), common atlas data (V3C_CAD), packed video data (V3C_PVD), or texture map video data (V3C_TVD).
[0519] The above vuh_v3c_parameter_set_id field specifies the identifier (i.e., vuh_v3c_parameter_set_id) of the active V3C parameter set (V3C VPS).
[0520] The V3C unit header according to embodiments may further include a vuh_atlas_id field if the vuh_unit_type field indicates attribute video data (V3C_AVD), geometry video data (V3C_GVD), accumulator video data (V3C_OVD), atlas data (V3C_AD), or packed video data (V3C_PVD).
[0521] The above vuh_atlas_id field specifies the index of the atlas corresponding to the current V3C unit.
[0522] The V3C unit header according to embodiments may further include a vuh_attribute_index field, a vuh_attribute_partition_index field, a vuh_map_index field, and a vuh_auxiliary_video_flag field when the vuh_unit_type field indicates attribute video data (V3C_AVD).
[0523] The above vuh_attribute_index field indicates the index of attribute video data carried as an attribute video data unit.
[0524] The above vuh_attribute_partition_index field indicates the index of the attribute dimension group carried as an attribute video data unit.
[0525] The above vuh_map_index field, if present, indicates the index of the current attribute stream.
[0526] If the value of the vuh_auxiliary_video_flag field is 1, it may indicate that the associated attribute video data unit contains only raw and / or EOM (Enhanced Occupancy Mode) coded points. As another example, if the value of the vuh_auxiliary_video_flag field is 0, it may indicate that the associated attribute video data unit may contain raw and / or EOM coded points. If the vuh_auxiliary_video_flag field does not exist, the value of the field may be inferred to be equal to 0. In some embodiments, raw and / or EOM coded points may also be referred to as PCM (Pulse Code Modulation) coded points.
[0527] A V3C unit header according to embodiments may further include a vuh_map_index field, a vuh_auxiliary_video_flag field, and a vuh_reserved_zero_12bits field when the vuh_unit_type field indicates geometry video data (V3C_GVD).
[0528] The above vuh_map_index field, if present, indicates the index of the current geometry stream.
[0529] If the value of the vuh_auxiliary_video_flag field is 1, it may indicate that the associated geometry video data unit contains only raw and / or EOM coded points. As another example, if the value of the vuh_auxiliary_video_flag field is 0, it may indicate that the associated geometry video data unit may contain raw and / or EOM coded points. If the vuh_auxiliary_video_flag field does not exist, the value of the field may be inferred to be equal to 0. In some embodiments, raw and / or EOM coded points may also be referred to as PCM (Pulse Code Modulation) coded points.
[0530] The above vuh_reserved_zero_12bits field is a reserved field for future use.
[0531] A V3C unit header according to embodiments may further include a vuh_output_texturemap_count field if the vuh_unit_type field indicates texture map video data (V3C_TVD).
[0532] The above vuh_output_texturemap_count field indicates the number of output texture maps.
[0533] A loop that repeats as many times as the value of the vuh_output_texturemap_count field may include a vuh_texture_skip_flag[i] field and a vuh_ref_max_num field.
[0534] The above vuh_texture_skip_flag[i] field is a flag indicating whether the texture map of the i-th frame is omitted. If the value of the vuh_texture_skip_flag[i] field is 1, it means that the i-th texture map is omitted, and if the value of the vuh_texture_skip_flag[i] field is 0, it means that the i-th texture map is not omitted.
[0535] The above vuh_ref_max_num field indicates the maximum number of reference frames for omitted texture map frames.
[0536] The loop that is added when the value of the above vuh_texture_skip_flag[i] field is 1 is repeated as many times as the value of the vuh_ref_max_num field, and may include the vuh_ref_frmae_idx[i][j] field and the vuh_recon_weight[i][j] field.
[0537] The above vuh_ref_frmae_idx[i][j] field indicates reference frame index information of the i-th texture map. At this time, up to j frames can be referenced according to the maximum number of reference frames.
[0538] The above vuh_recon_weight[i][j] field specifies weight information of the jth reference frame of the ith texture map.
[0539] The following is an example of implementing a method for restoring a texture map whose coding has been omitted on the transmitting side based on the multi-reference structure of the present disclosure in code (i.e., an algorithm). Below, skipCount represents the total number of frames whose coding has been omitted, numOutFrames represents the total number of output frames, numInFrames represents the total number of input frames, and texture_skip_flag[f] represents whether the texture map coding of the corresponding frame has been omitted. For example, if the value of texture_skip_flag[f] is 1, it means that the coding of the corresponding frame has been omitted. In addition, weight[i] represents a weight value assigned to each reference frame, and is used for calculating a weight-based average value when restoring a texture map. In other words, the algorithm below represents a case where an omitted texture map is restored by referencing texture maps of two or more frames and calculating based on each weight.
[0540] skipCount = numOutFrames - numInputFrames
[0541] for( f = 0; f < numOutFrames; f++ ) {
[0542] if (texture_skip_flag[f]) {
[0543] for( c=0; c < iNumComp; c++ ) {
[0544] for( y=0; y < iHeight[c]; y++ ) {
[0545] for( x=0; x < iWidth[c]; x++ ) {
[0546] for (i=0; i <skipCount[f]; i++){
[0547] outputFrames[f][c][y][x]
[0548] += weight[i] * inputFrames[ ref_idx[f][i] ][c][y][x]
[0549] total_weight += weight[i]
[0550] }
[0551] outputFrames[f][c][y][x]
[0552] = outputFrames[f][c][y][x] / total_weight
[0553] total_weight = 0
[0554] }
[0555] }
[0556] }
[0557] }
[0558] else {
[0559] outputFrames[ f ][ c ][ y ][ x ] = inputFrames[ f ][ c ][ y ][ x ]
[0560] }
[0561] }
[0562] The following is a more specific description of the operation of the receiving device of FIG. 18 when texture map restoration related information (whether texture map coding is omitted and texture map reference information, etc.) is transmitted from the transmitting device to the receiving device via the V3C unit header as described above.
[0563] That is, the receiving device of FIG. 18 receives the V-DMC (or V3C) bitstream transmitted from the transmitting device, separates the base mesh sub-bitstream (or bitstream), the displacement vector sub-bitstream (or bitstream), and the texture map sub-bitstream (or bitstream) through a demultiplexer, and then performs a process of decoding the separated base mesh bitstream, displacement vector bitstream, and texture map bitstream, respectively.
[0564] First, the base mesh sub-bitstream is decoded into a motion vector, which is the difference between base meshes, through a motion vector decoder in the case of an inter-frame, and then decoded into a base mesh by adding the decoded motion vector value to the reference base mesh. In the case of an intra-frame, it is decoded into a base mesh through a static mesh decoder.
[0565] The displacement vector sub-bitstream decodes the displacement vector coefficients in the reverse order of encoding of the transmitting device, performs inverse quantization and inverse transformation, and then is inversely transformed from the local coordinate system to the Cartesian coordinate system to be restored to the final displacement vector. The mesh restoration unit calculates the vertex geometry information of the restored mesh by adding the displacement vector restored by the mesh restoration unit to the vertices generated through the subdivision process in the mesh refinement unit, thereby restoring the final geometry information.
[0566] The texture map sub-bitstream is decoded into attribute data through a texture map decoder. The map extraction unit then extracts the texture map from the decoded attribute data. Furthermore, bit depth conversion and resolution conversion processes are performed on the extracted data for data alignment.
[0567] And, in order to perform the map reconstruction process, that is, the map reconstruction unit (or the omitted texture map restoration unit (15022)) first parses the v3c_unit_header() of the V3C sample stream. And, if the value of the type information (vuh_unit_type) of the V3C unit header is 4, it means that it is V3C_AVD, and syntax related to attribute data can be parsed. In addition, if the value of vuh_unit_type of the V3C unit header is 7, it means that it is V3C_TVD, and syntax related to texture map data can be parsed. In the present disclosure, the related information (i.e., syntaxes) for performing texture map restoration may be transmitted through the V3C unit header whose vuh_unit_type value is 4 and / or may be transmitted through the V3C unit header whose vuh_unit_type value is 7. If the related information (i.e., syntaxes) for performing texture map restoration is transmitted through a V3C unit header whose vuh_unit_type value is 4 (V3C_AVD), the related information for performing texture map restoration can be parsed by parsing v3c_unit_header() of FIG. 29. If the related information (i.e., syntaxes) for performing texture map restoration is transmitted through a V3C unit header whose vuh_unit_type value is 7 (V3C_TVD), the related information for performing texture map restoration can be parsed by parsing v3c_unit_header() of FIG. 30.In the present disclosure, related information (i.e., syntaxes) for performing texture map restoration may include information on whether to skip texture map coding (i.e., flag information) (texture_skip_flag), information on the total number of output texture maps for deriving the number of omitted texture maps (vuh_output_texturemap_count), information on the maximum number of reference frames (vuh_ref_max_num), information on reference frame index (vuh_ref_frame_idx), and information on the weight of the reference frame (vuh_recon_weight).
[0568] According to embodiments, the map reconstruction unit can know the number of original total output texture maps based on vuh_output_texturemap_count, and can derive the number of omitted texture maps through the difference with the number of texture maps currently input. In addition, it can know whether the current frame is a frame with omitted coding based on vuh_texture_skip_flag. If the value of vuh_texture_skip_flag is 1, it means that the texture map coding of the corresponding frame is omitted, and reference frame index information (vuh_ref_frame_idx), weight information of the reference frame (vuh_recon_weight), and the maximum number of reference frames (vuh_ref_max_num) can be parsed to restore the omitted texture map. Then, the texture map data value of the reference frame (vuh_ref_frame_idx) corresponding to the number of reference frames of the current frame is multiplied by the weight value (vuh_recon_weight) and divided by the number of omitted texture maps derived previously (i.e., frames with omitted texture map coding) to ultimately restore the omitted texture map. At this time, the value of vuh_ref_max_num may be 1, 2, or a specific value higher than that depending on the result calculated by the encoder or an agreement between codecs. In addition, if weighting is not applied to each reference frame, the vuh_ref_recon_weight value may be 1.
[0569] Once all texture maps, including the omitted texture maps, are restored through the above processes, a final restoration mesh is generated along with the previously restored final geometry information.
[0570] As described above, among mesh data components, texture maps account for the largest data proportion. The present disclosure can achieve a significant bit-saving effect by omitting coding and transmission of such texture maps on a frame-by-frame basis without a significant difference in quality. In particular, the present disclosure proposes a method of omitting coding of texture maps in a V-DMC codec when the current texture map and the reference texture map are similar, transmitting a bitstream including the remaining texture map, and restoring the omitted texture map at the transmitting end based on a weight by receiving the bitstream transmitted with the texture map omitted. In addition, the present disclosure proposes a method of signaling related syntax for restoring the omitted texture map (i.e., texture map restoration related information) through a V3C unit header based on a V3C bitstream. In particular, a method is proposed of signaling texture map restoration related information using an existing V3C unit type, V3C_AVD, and a method of signaling texture map restoration related information by creating a new V3C unit type called V3C_TVD. The present disclosure proposes a method that not only achieves significant bit savings by omitting texture maps on a frame-by-frame basis, but also restores the omitted texture maps based on weights, thereby restoring even higher-quality data. Furthermore, by proposing a method for transmitting texture map restoration-related information (i.e., syntax) that conforms to the V3C bitstream structure, significant improvements in efficient compression and transmission speed can be achieved.
[0571] Fig. 31 is a flowchart showing an example of a transmission method according to embodiments. The transmission method according to embodiments may include a step of encoding mesh data (S31011) and a step of transmitting a bitstream including the encoded mesh data (S31012). In Fig. 31, only the omission and signaling of the texture map will be described, and the description of Figs. 1 to 15 will be referred to for the remaining components. That is, as described in Figs. 15 and 16, in the step of encoding the mesh data (S31011), the texture map coding of some frames may be omitted. In one embodiment, in the step of encoding the mesh data (S31011), the difference in texture PSNR between the mesh of the current frame and the original mesh and the texture PSNR between the mesh of the reference frame and the original mesh may be calculated by a method such as a point-based metric, to determine whether to omit the texture map coding of the current frame.
[0572] In one embodiment, in the step of encoding the mesh data (S31011), if it is determined that the texture map coding of the current frame is omitted, texture map restoration related information (e.g., information on whether to omit texture map coding (texture_skip_flag), direction information of the reference texture map (refDirection_idx), index information of the reference texture map (ref_texture_idx), original sequence size information (num_original_gof_size), information on whether to continuously apply the SEI message during the current frame or sequence (persistence_flag), etc.) may be signaled through an SEI message. The information signaled through the SEI message may be transmitted through an SEI message (sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI) of a NAL sample stream in an atlas sub-bitstream (V3C_AD) of a V3C sample stream.
[0573] In another embodiment, in the step of encoding the mesh data (S31011), if it is determined that the texture map coding of the current frame is omitted, texture map restoration related information (e.g., texture map coding omission information (vuh_texture_skip_flag), total output texture map count information for deriving the number of omitted texture maps (vuh_output_texturemap_count), maximum reference frame count information (vuh_ref_max_num), reference frame index information (vuh_ref_frame_idx), reference frame weight information (vuh_recon_weight), etc.) may be signaled through the V3C unit header. At this time, the texture map restoration related information may be transmitted through the V3C unit header of the V3C unit type indicating attribute video data (i.e., V3C_AVD) or may be transmitted through the V3C unit header of the new V3C unit type (i.e., V3C_TVD indicating texture map video data).
[0574] In the step of transmitting the bitstream (S31012), the sub-bitstreams generated in the step of encoding the mesh data (S31011) (e.g., atlas sub-bitstream, base mesh sub-bitstream, geometry video (or displacement video) sub-bitstream, attribute video (or texture map video) sub-bitstream, packed video sub-bitstream) are multiplexed into one V3C bitstream (or V-DMC bitstream) and transmitted to a receiving device.
[0575] Fig. 32 is a flowchart illustrating an example of a receiving method according to embodiments. The receiving method according to embodiments may include a step (S32011) of receiving a bitstream containing mesh data and a step (S32012) of decoding the mesh data contained in the bitstream.
[0576] In the step of receiving the above bitstream (S32011), a V-DMC (or V3C) bitstream is received, and a base mesh sub-bitstream (or bitstream), a displacement vector sub-bitstream (or bitstream), and a texture map sub-bitstream (or bitstream) are separated from the received V-DMC (or V3C) bitstream. In Fig. 32, only the restoration and signaling of the omitted texture map will be described, and for the remaining components, refer to the descriptions of Figs. 1 to 18.
[0577] That is, in the step of decoding the above mesh data (S32012), the texture map sub-bitstream is restored to attribute data, and the texture map is extracted from the restored attribute data.
[0578] The present disclosure will explain, by dividing into different embodiments, when texture map restoration related information is received via an SEI message and when it is received via a V3C unit header, as follows.
[0579] In one embodiment, in the step (S32012) of decoding the mesh data, first, sei_rbsp() of NAL_PREFIX_ESEI or NAL_SUFFIX_ESEI of V3C_AD is parsed. When sei_rbsp() is parsed, sei_message() is performed, and at this time, if the payloadType of sei_payload is 69, skipped_frame_indication (payloadSize) for restoring a texture map with skipped coding is performed. In the present disclosure, related information (i.e., syntaxes) for performing texture map restoration may include information on whether to skip texture map coding (texture_skip_flag), direction information of a reference texture map (refDirection_idx), index information of a reference texture map (ref_texture_idx), original sequence size information (num_original_gof_size), and information on whether to continuously apply the SEI message during the current frame or sequence (persistence_flag).
[0580] The step (S32012) of decoding the above mesh data can determine whether to apply the SEI message only to the currently decoded texture map frame by parsing the persistence_flag in the skipped_frame_indication or to continuously apply the SEI message until one of certain conditions (such as before a new atlas sequence starts or until the bitstream ends) is satisfied. In addition, the step (S32012) of decoding the above mesh data can receive the original size information of the current sequence (GOF) by parsing the num_original_gof_size and infer the number of skipped texture maps as a difference from the number of currently decoded texture maps. The step (S32012) of decoding the above mesh data parses texture_skip_flag for each frame while the current sequence is in its original size (i.e., repeats as much as the original size) to determine whether the texture map coding of the current frame is omitted, and if the value of texture_skip_flag is 1, ref_direction_idx is parsed to identify information on which direction the omitted texture map referenced. For example, if the value of ref_direction_idx is 0, this indicates that the omitted frame is derived by referencing a past frame, and if the value of ref_direction_idx is 1, this indicates that the omitted frame is derived by referencing a future frame. If the value of ref_direction_idx is 2, this indicates that the omitted frame is derived using a bidirectional texture map. If the value of ref_direction_idx is 3 or absent, this means that the texture map of the current frame is not omitted. In particular, the step of decoding the above mesh data (S32012) can restore the texture map with the intermediate value of the past and future frames in the case of a reference to a bidirectional texture map.Additionally, in the step of decoding the mesh data (S32012), the POC (or index) difference between frames may be calculated for slightly higher quality texture map restoration, and texture map restoration may be performed using a weighted average based on the ratio of the reference texture map distance.
[0581] In another embodiment, in the step (S32012) of decoding the mesh data, first, v3c_unit_header() of the V3C sample stream is parsed. Then, if the value of the type information (vuh_unit_type) of the V3C unit header is 4, it means that it is V3C_AVD, and syntax related to attribute data can be parsed. In addition, if the value of vuh_unit_type of the V3C unit header is 7, it means that it is V3C_TVD, and syntax related to texture map data can be parsed. In the present disclosure, related information (i.e., syntaxes) for performing texture map restoration may be transmitted through the V3C unit header whose value of vuh_unit_type is 4 and / or may be transmitted through the V3C unit header whose value of vuh_unit_type is 7. If the related information (i.e., syntaxes) for performing texture map restoration is transmitted through a V3C unit header whose vuh_unit_type value is 4 (V3C_AVD), the related information for performing texture map restoration can be parsed by parsing v3c_unit_header() of FIG. 29. If the related information (i.e., syntaxes) for performing texture map restoration is transmitted through a V3C unit header whose vuh_unit_type value is 7 (V3C_TVD), the related information for performing texture map restoration can be parsed by parsing v3c_unit_header() of FIG. 30. In the present disclosure, related information (i.e., syntaxes) for performing texture map restoration may include information on whether to skip texture map coding (i.e., flag information) (texture_skip_flag), information on the total number of output texture maps for deriving the number of omitted texture maps (vuh_output_texturemap_count), information on the maximum number of reference frames (vuh_ref_max_num), information on reference frame index (vuh_ref_frame_idx), and information on the weight of the reference frame (vuh_recon_weight).
[0582] In the step (S32012) of decoding the above mesh data, the number of original total output texture maps can be known based on vuh_output_texturemap_count, and the number of omitted texture maps can be derived through the difference with the number of texture maps currently input. In addition, it can be known whether the current frame is a frame with omitted coding based on vuh_texture_skip_flag. If the value of vuh_texture_skip_flag is 1, it means that the texture map coding of the corresponding frame is omitted, and reference frame index information (vuh_ref_frame_idx), weight information of the reference frame (vuh_recon_weight), and the maximum number of reference frames (vuh_ref_max_num) can be parsed to restore the omitted texture map. Then, the texture map data value of the reference frame (vuh_ref_frame_idx) corresponding to the number of reference frames of the current frame is multiplied by the weight value (vuh_recon_weight) and divided by the number of omitted texture maps derived previously (i.e., frames with omitted texture map coding) to ultimately restore the omitted texture map. At this time, the value of vuh_ref_max_num may be 1, 2, or a specific value higher than that depending on the result calculated by the encoder or an agreement between codecs. In addition, if weighting is not applied to each reference frame, the vuh_ref_recon_weight value may be 1.
[0583] Once all texture maps, including the omitted texture maps, are restored through the above processes, a final restoration mesh is generated along with the previously restored final geometry information.
[0584] Each of the parts, modules, or units described above may be software, processors, or hardware parts that execute sequential execution processes stored in memory (or storage units). Each of the steps described in the embodiments described above may be performed by processors, software, or hardware parts. Each of the modules / blocks / units described in the embodiments described above may operate as a processor, software, or hardware. In addition, the methods presented in the embodiments may be implemented as code. This code may be written on a processor-readable storage medium and thus may be read by a processor provided by an apparatus.
[0585] Furthermore, throughout the specification, when a part is said to "include" a component, this does not exclude other components, unless otherwise specifically stated, but rather implies the inclusion of other components. Furthermore, terms such as "part" described in the specification mean a unit that processes at least one function or operation, which may be implemented using hardware, software, or a combination of hardware and software.
[0586] For convenience of explanation, this specification has been described separately in each drawing. However, it is also possible to design new embodiments by combining the embodiments described in each drawing. Furthermore, designing a computer-readable recording medium containing a program for executing the previously described embodiments, as required by those skilled in the art, is also within the scope of the embodiments.
[0587] The devices and methods according to the embodiments are not limited to the configurations and methods of the embodiments described above, but the embodiments may be configured by selectively combining all or part of each embodiment so that various modifications can be made.
[0588] Although preferred embodiments of the embodiments have been illustrated and described, the embodiments are not limited to the specific embodiments described above, and various modifications may be made by those skilled in the art to which the present disclosure pertains without departing from the spirit or scope of the embodiments claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the embodiments.
[0589] The various components of the devices of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. The various components of the embodiments may be implemented by a single chip, for example, a single hardware circuit. The components according to the embodiments may be implemented by separate chips. At least one of the components of the devices of the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may perform, or include instructions for performing, one or more of the operations / methods according to the embodiments. The executable instructions for performing the methods / operations of the devices of the embodiments may be stored in non-transitory CRMs or other computer program products configured to be executed by one or more processors, or may be stored in temporary CRMs or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiments may be used as a concept including not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. Additionally, it may include implementations in the form of carrier waves, such as transmissions via the Internet. Furthermore, processor-readable recording media may be distributed across network-connected computer systems, allowing processor-readable code to be stored and executed in a distributed manner.
[0590] In this document, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" means "at least one of A, B, and / or C". Additionally, "or" in this document is interpreted as "and / or". For example, "A or B" can mean 1) "A" only, 2) "B" only, or 3) "A and B". In other words, "or" in this document can mean "additionally or alternatively".
[0591] Various elements of the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various elements of the embodiments may be implemented on a single chip, such as a hardware circuit. In some embodiments, the embodiments may optionally be implemented on separate chips. In some embodiments, at least one of the elements of the embodiments may be implemented within one or more processors that include instructions for performing operations according to the embodiments.
[0592] Additionally, the operations according to the embodiments described in this document may be performed by a transceiver device including one or more memories and / or one or more processors according to the embodiments. One or more memories may store programs for processing / controlling the operations according to the embodiments, and one or more processors may control various operations described in this document. One or more processors may be referred to as a controller, etc. The operations according to the embodiments may be performed by firmware, software, and / or a combination thereof, and the firmware, software, and / or a combination thereof may be stored in a processor or a memory.
[0593] Terms such as "first" and "second" may be used to describe various components of the embodiments. However, the various components according to the embodiments should not be interpreted in a limited manner by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal may be referred to as a "second user input signal." Similarly, a second user input signal may be referred to as a "first user input signal." The use of these terms should be interpreted as not departing from the scope of the various embodiments. Although "first user input signal" and "second user input signal" are both user input signals, they do not necessarily mean the same user input signals unless the context clearly indicates otherwise.
[0594] The terminology used to describe the embodiments is for the purpose of describing particular embodiments and is not intended to be limiting of the embodiments. As used in the description of the embodiments and in the claims, the singular is intended to include the plural unless the context clearly dictates otherwise. The expressions “and / or” are used to mean all possible combinations of the terms. The expression “comprises” or “includes” describes the presence of features, numbers, steps, elements, and / or components, but does not mean that additional features, numbers, steps, elements, and / or components are not included. Conditional expressions such as “if” or “when” used to describe the embodiments are not intended to be limited to only optional cases. When a specific condition is satisfied, a related action is performed in response to a specific condition, or a related definition is intended to be interpreted.
[0595] As described above, the relevant contents have been described in the best form for carrying out the embodiments.
[0596] As described above, the embodiments may be applied, in whole or in part, to mesh data transmission and reception devices and systems. Those skilled in the art will appreciate that various modifications and variations may be made to the embodiments within the scope of the embodiments. The embodiments may include modifications and variations, and such modifications and variations do not depart from the scope of the claims and their equivalents.
Claims
1. A step of receiving a bitstream containing mesh data; and A step of decoding the above mesh data; comprising: How to decode.
2. In the first paragraph, the decoding step A base mesh processing step for restoring a base mesh from a base mesh bitstream; A displacement information processing step for restoring displacement information from a displacement vector bitstream; A texture map processing step for restoring texture maps from a texture map bitstream; a step of determining whether at least one texture map is omitted from the above texture map bitstream; and A decoding method comprising the step of restoring the at least one omitted texture map based on texture map restoration related information and at least one reference frame when it is determined that at least one omitted texture map exists.
3. In paragraph 2, A decoding method wherein the texture map restoration related information includes information for identifying whether at least one texture map is omitted and reference texture map information related to at least one reference frame.
4. In paragraph 2, A decoding method in which the above texture map restoration related information is carried through a SEI (Supplemental enhancement information) message.
5. In paragraph 2, A decoding method in which the above texture map bitstream is composed of one or more data units, each data unit is composed of a data unit header including type information and a data unit payload carrying texture map data, and information related to texture map restoration is carried through the data unit header.
6. Memory; and comprising at least one processor connected to said memory, At least one processor of the above: Receiving a bitstream containing mesh data; and Decode the above mesh data; configured to do so, Decoding device.
7. In the 6th paragraph, at least one processor A base mesh processing unit that restores the base mesh from the base mesh bitstream; A displacement information processing unit that restores displacement information from a displacement vector bitstream; A texture map processing unit that restores texture maps from a texture map bitstream; A decoding device including an omitted texture map restoration unit that determines whether at least one texture map is omitted from the texture map bitstream, and restores the at least one omitted texture map based on texture map restoration-related information and at least one reference frame if it is determined that there is the at least one omitted texture map.
8. In paragraph 7, A decoding device wherein the texture map restoration related information includes information for identifying whether at least one texture map is omitted and reference texture map information related to at least one reference frame.
9. In paragraph 7, A decoding device that carries the above texture map restoration related information through a SEI (Supplemental enhancement information) message.
10. In paragraph 7, A decoding device in which the above texture map bitstream is composed of one or more data units, each data unit comprising a data unit header including type information and a data unit payload carrying texture map data, and information related to texture map restoration is carried through the data unit header.
11. Step of encoding mesh data; and A step of transmitting a bitstream including the encoded mesh data; comprising: Encoding method.
12. In the 11th paragraph, the encoding step A step of generating a base mesh bitstream by encoding a base mesh generated by simplifying the original mesh; A step of generating a displacement vector bitstream by encoding displacement information generated based on the above base mesh; A step of restoring a mesh based on the encoded base mesh and the encoded displacement information; A step of generating texture maps based on the original mesh and the restored mesh, and determining whether to omit at least one texture map among the generated texture maps; A step of generating a texture map bitstream by encoding the remaining texture maps except for at least one texture map for which the omission is determined; and An encoding method comprising the step of generating texture map restoration related information for restoration of at least one texture map whose omission has been determined.
13. In paragraph 12, An encoding method wherein the texture map restoration related information includes information for identifying whether at least one texture map is omitted and reference texture map information related to at least one reference frame.
14. In paragraph 12, An encoding method in which the above texture map restoration related information is carried through a SEI (Supplemental enhancement information) message.
15. In paragraph 12, An encoding method in which the texture map bitstream is composed of one or more data units, each data unit is composed of a data unit header including type information and a data unit payload carrying texture map data, and information related to texture map restoration is carried through the data unit header.
16. A computer-readable storage medium storing a bitstream generated by the method according to Article 11.
17. Step of obtaining bitstream for image information; wherein the bitstream is generated based on a step of encoding mesh data and a step of transmitting a bitstream including the encoded mesh data; and A method comprising the step of transmitting data including the bitstream.
Citation Information
Patent Citations
Method and electronic device to perform operation according to expansion direction in response to expansion
KR1020230015025A
Encoder-assisted adaptive video frame interpolation
US20140376637A1
Apparatus, a method and a computer program for omnidirectional video
US20230059516A1
Point cloud data transmission device, point cloud data transmission method, point cloud data reception device, and point cloud data reception method
WO2021261865A1