3d data transmission device, 3d data transmission method, 3d data reception device, and 3d data reception method
By simplifying and subdividing 3D grid data, combined with 2D video encoding and decoding or zero-run coding technology, the problems of low efficiency and high encoding complexity of 3D grid data are solved, and efficient grid data compression and reconstruction are achieved.
Patent Information
- Application Number
- CN202380065176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-20
- Filing Date
- 2023-09-20
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to efficiently transmit and receive 3D grid data of a large number of points, resulting in low data transmission efficiency and high encoding/decoding complexity.
By simplifying and subdividing the original grid data, generating basic grid data and displacement information, encoding and decoding these data using 2D video encoding decoder or zero-run encoding, achieving efficient compression and reconstruction of grid data.
It improves the transmission efficiency and image quality of 3D grid data, reduces the complexity of encoding and decoding, and supports the transmission of grid content with different resolutions and image quality.
Smart Images

Figure CN119948874A_ABST
Abstract
Description
Technical Field
[0001] Embodiments provide a method for providing 3D content to provide users with various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and self-driving services. Background Art
[0002] Point cloud data or mesh data in 3D content is a collection of points in 3D space. However, due to the large number of points in 3D space, it is difficult to create point cloud data or mesh data.
[0003] In other words, a large throughput is required to transmit and receive 3D data (eg, point cloud or mesh data) having a considerable number of points. Summary of the invention
[0004] Technical issues
[0005] An object of the present disclosure is to provide an apparatus and method for efficiently transmitting and receiving mesh data to solve the above-mentioned problems.
[0006] Another object of the present disclosure is to provide an apparatus and method for solving delay and encoding / decoding complexity of mesh data.
[0007] Another object of the present disclosure is to provide an apparatus and method for compressing and reconstructing mesh data per resolution.
[0008] Another object of the present disclosure is to provide an apparatus and method for increasing compression efficiency of mesh data by applying a zero-run method when compressing or reconstructing mesh data per resolution.
[0009] The embodiments are not limited to the above-mentioned purposes, and the scope of the embodiments can be extended to other purposes that can be inferred by those skilled in the art based on the entire content of the present disclosure.
[0010] Technical Solution
[0011] To achieve these objectives and other advantages and in accordance with the purposes of the present invention, as embodied and broadly described herein, a method for transmitting three-dimensional (3D) data may include the steps of: encoding raw mesh data; and transmitting a bitstream containing the encoded mesh data and signaling information.
[0012] According to an embodiment, encoding of mesh data may include: generating basic mesh data by simplifying original mesh data and encoding the basic mesh data; generating additional vertices by subdividing the simplified mesh data one or more times and determining one or more levels that vary according to the number of subdivisions; reconstructing the encoded basic mesh data; generating displacement information based on the subdivided mesh data and the reconstructed basic mesh data; encoding the displacement information; reconstructing the encoded displacement information; reconstructing the mesh data based on the reconstructed basic mesh data and the reconstructed displacement information; regenerating a texture map based on a texture map of the original mesh data and the reconstructed mesh data; and encoding the regenerated texture map.
[0013] According to an embodiment, the encoding of the displacement information may include encoding the displacement information corresponding to at least one of the one or more levels using a 2D video codec.
[0014] According to an embodiment, the encoding of the displacement information may include encoding the displacement information corresponding to at least one of the one or more levels using zero-run encoding.
[0015] According to an embodiment, the encoding of the displacement information may include packing the displacement information in a plurality of frames corresponding to at least one of the one or more levels, and encoding the packed displacement information using zero-run encoding.
[0016] According to an embodiment, displacement information in a plurality of frames may be packed using an interleaving method or a serial method based on a mapping relationship of vertices between frames.
[0017] According to an embodiment, the signaling information may include information related to one or more levels or information related to packing of displacement information in a plurality of frames.
[0018] Depending on the implementation, the one or more levels may be at least one level of detail (LoD) or at least one scalable LoD (sLoD).
[0019] According to an embodiment, at least one sLoD may be configured based on at least one LoD, wherein a specific sLoD of the at least one sLoD may be mapped to one of the at least one LoD.
[0020] According to an embodiment, the number of levels in at least one sLoD may be different from the number of levels in at least one LoD.
[0021] According to an embodiment, the encoding of the texture map may include encoding the texture map corresponding to at least one of the one or more levels using a 2D video codec.
[0022] According to an embodiment, an apparatus for transmitting three-dimensional (3D) data may include: an encoder configured to encode original mesh data; and a transmitter configured to transmit a bitstream including the encoded mesh data and signaling information.
[0023] According to an embodiment, the encoder may include: a base mesh compressor configured to generate base mesh data by simplifying original mesh data and encode the base mesh data; a mesh subdivider configured to generate additional vertices by subdividing the simplified mesh data one or more times and determine one or more levels that vary according to the number of subdivisions; a base mesh reconstructor configured to reconstruct the encoded base mesh data; a displacement information generator configured to generate displacement information based on the subdivided mesh data and the reconstructed base mesh data; a displacement information encoder configured to encode the displacement information; a displacement information reconstructor configured to reconstruct the encoded displacement information; a mesh reconstructor configured to reconstruct the mesh data based on the reconstructed base mesh data and the reconstructed displacement information; a texture map generator configured to regenerate a texture map based on a texture map of the original mesh data and the reconstructed mesh data; and a texture map generator configured to encode the regenerated texture map.
[0024] According to an embodiment, a method for receiving three-dimensional (3D) data may include the following steps: receiving a bitstream including encoded mesh data and signaling information; decoding the encoded mesh data in the bitstream based on the signaling information; and rendering the decoded mesh data.
[0025] According to an embodiment, decoding of mesh data may include: reconstructing base mesh data from encoded mesh data; generating additional vertices by subdividing the reconstructed base mesh data one or more times and determining one or more levels based on signaling information and the number of subdivisions; decoding and reconstructing displacement information from the encoded mesh data; reconstructing mesh data based on the subdivided base mesh data and the reconstructed displacement information; decoding and reconstructing a texture map from the encoded mesh data; and performing rendering based on the reconstructed mesh data and the reconstructed texture map.
[0026] Beneficial Effects
[0027] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may provide high-quality 3D services.
[0028] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may implement various video codec schemes.
[0029] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may support general 3D content, for example, for autonomous driving services.
[0030] According to an embodiment, a 3D data transmission method, a 3D data transmission device, a 3D data receiving method, and a 3D data receiving device may perform scalable encoding and decoding on displacement information and / or texture maps included in mesh data, thereby allowing hierarchical encoding / decoding of mesh data, and selectively reconstructing and providing mesh content with optimal resolution and image quality based on a network and a receiving environment.
[0031] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may include one or more sub-bitstream sets having different resolutions and image qualities through scalable encoding and decoding, thereby providing a user with mesh contents having different resolutions and increasing compression efficiency of mesh data.
[0032] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may apply scalable encoding and decoding to displacement information and / or a texture map included in mesh data, thereby reconstructing the mesh data to a usable level based on the performance or display characteristics of a receiver and a network environment.
[0033] According to an embodiment, by appropriately distributing resources, by allowing a user to specify different degrees of accuracy and level of mesh data according to the importance and usage frequency of an object, and by reconstructing each object at a desired level when representing multiple mesh data objects, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus can efficiently utilize given resources.
[0034] According to an embodiment, a 3D data transmitting method, a 3D data transmitting apparatus, a 3D data receiving method, and a 3D data receiving apparatus may enable mesh data to be used in a wider range of network environments and applications, and may further expand the scope of utilization of mesh data by flexibly using resources of a receiver. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated into and constitute a part of this application. The accompanying drawings illustrate embodiments of the present disclosure and together with the description are used to illustrate the principles of the present disclosure. In order to better understand the various embodiments described below, reference should be made to the description of the following embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to represent the same or similar parts. In the drawings:
[0036] Figure 1A system for providing dynamic grid content according to an embodiment is shown;
[0037] Figure 2 A V-MESH compression method according to an embodiment is shown;
[0038] Figure 3 illustrates pre-processing in V-MESH compression according to an embodiment;
[0039] Figure 4 An edge midpoint subdivision method according to an embodiment is shown;
[0040] Figure 5 shows a displacement generation process according to an embodiment;
[0041] Figure 6 shows the intra-frame encoding process of V-grid data according to an embodiment;
[0042] Figure 7 illustrates an inter-frame encoding process of V-grid data according to an embodiment;
[0043] Figure 8 The lifting transformation process of displacement according to the embodiment is shown;
[0044] Fig. 9 shows a process of packing transform coefficients into a 2D image according to an embodiment;
[0045] Fig.10 illustrates the attribute transfer process in the V-MESH compression method according to an embodiment;
[0046] Fig.11 shows the intra-frame decoding process of V-grid data according to an embodiment;
[0047] Fig.12 shows an inter-frame decoding process of V-grid data according to an embodiment;
[0048] Fig.13 A mesh data transmission device according to an embodiment is shown;
[0049] Fig.14 A mesh data receiving device according to an embodiment is shown;
[0050] Fig.15 A grid data transmitting device according to an embodiment is shown;
[0051] Fig.16 is a diagram showing an example of zero-run encoding according to an embodiment;
[0052] Fig.17 is a flow chart illustrating an example method of zero-run encoding of transform coefficients according to an embodiment;
[0053] Fig.18 is a diagram showing an example of zero-run encoding in a normal coordinate system according to an embodiment;
[0054] Fig.19 is a flow chart showing an example of zero-run encoding of displacement vector transform coefficients of normal components in a normal coordinate system according to an embodiment;
[0055] Fig. 20 is a flow chart showing an example of zero-run encoding of displacement vector transform coefficients of tangential components in a normal coordinate system according to an embodiment;
[0056] Fig.21 An example of packing displacement vector transform coefficients of multiple frames in an interleaved manner according to an embodiment is shown;
[0057] Fig. 22 An example of packing displacement vector transform coefficients of multiple frames in a serial manner according to an embodiment is shown;
[0058] Fig.23 shows an example of zero-run encoding of displacement vector transform coefficients for two interleaved frames per LoD level according to an embodiment;
[0059] Fig.24 is a diagram showing an example of texture color mapping of a texture map generator according to an embodiment;
[0060] Fig.25 shows an example of compressing a texture map generated per level using a scalable video codec according to an embodiment;
[0061] Fig.26 is a diagram showing a mesh data receiving device according to an embodiment;
[0062] Fig. 27 (a) is a flowchart showing an example of zero-run decoding of displacement vector transform coefficients according to an embodiment;
[0063] Fig. 27 (b) is a flowchart showing an example of decoding of the absolute value of the displacement vector transform coefficient according to an embodiment;
[0064] Fig.28 (a) and Fig.28 (b) is a flowchart showing an example of zero-run decoding of a displacement vector transform coefficient of a normal component in a normal coordinate system according to an embodiment;
[0065] Fig.29 (a) and Fig.29(b) is a flowchart showing an example of zero-run decoding of a displacement vector transform coefficient of a tangential component in a normal coordinate system according to an embodiment;
[0066] Fig.30 is a flow chart illustrating entropy decoding of absolute values abs(dispn) of transform coefficients according to an embodiment;
[0067] Fig.31 is a flowchart illustrating an example method of decoding transform coefficients of multiple frames according to an embodiment;
[0068] Fig.32 is a flowchart illustrating another example method of decoding transform coefficients of a plurality of frames according to an embodiment;
[0069] Fig.33 is a diagram showing an example of scalable decoding of a texture map by a texture map decoder according to an embodiment;
[0070] Fig.34 shows an example syntax structure of LoD related information (LoD_Info()) in signaling information according to an embodiment;
[0071] Fig.35 shows an example syntax structure of displacement vector decoding related information (Decode_Disp()) in signaling information according to an embodiment;
[0072] Fig.36 shows an example syntax structure of information related to decoding of displacement vector transform coefficients (decode_displacement_coefficient()) in signaling information according to an embodiment;
[0073] Fig.37 An example syntax structure of packing related information for multiple frames (unpack_displacemenst_for_multiframe()) in signaling information according to an embodiment is shown;
[0074] Fig.38 is a flowchart illustrating an example sending method according to an embodiment; and
[0075] Fig.39 is a flow chart illustrating an example receiving method according to an embodiment. DETAILED DESCRIPTION
[0076] Reference will now be made in detail to preferred embodiments of the present disclosure, examples of which are shown in the accompanying drawings. The detailed description given below with reference to the accompanying drawings is intended to illustrate exemplary embodiments of the present disclosure, rather than to illustrate the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.
[0077] Although most of the terms used in the present disclosure are selected from general terms widely used in the art, some terms are arbitrarily selected by the applicant, and their meanings are explained in detail in the following description as needed. Therefore, the present disclosure should be understood based on the intended meanings of the terms rather than their simple names or meanings.
[0078] With recent advances in 3D data modeling and rendering technology, active research has been conducted on generating and processing 3D data across various fields, including virtual reality (VR), augmented reality (AR), autonomous driving, computer-aided design (CAD) / computer-aided manufacturing (CAM), and geographic information systems (GIS). 3D data can be represented as a point cloud or a mesh according to a representation format. The mesh consists of geographic information indicating the coordinates of each vertex or point, connection information indicating the connections between vertices, a texture map representing color information about the mesh surface as 2D image data, and texture coordinates indicating mapping information between the mesh surface and the texture map. In the present disclosure, a mesh is defined as a dynamic mesh when at least one element constituting the mesh changes over time, and is defined as a static mesh when it does not change.
[0079] Compared to 2D image data, dynamic mesh data involves significantly larger amounts of element data to represent the mesh. As a result, techniques have been developed for efficiently compressing large amounts of mesh data to store and transmit the data.
[0080] Figure 1 A system for providing dynamic grid content according to an embodiment is shown.
[0081] Figure 1 The system in the embodiment includes a sending device 100 and a receiving device 110. The sending device 100 may include a grid video acquisition unit (or part) 101, a grid video encoder 102, a file / segment encapsulator 103, and a sender 104. The receiving device 110 may include a receiver 111, a file / segment decapsulator 112, a grid video decoder 113, and a renderer 114. Figure 1The components in the embodiment may correspond to hardware, software, a processor and / or a combination thereof. In the following description, the mesh data transmitting device according to the embodiment may be interpreted as referring to the 3D data transmitting device or the transmitting device 100, or to the mesh video encoder (hereinafter referred to as the encoder) 102. The mesh data receiving device according to the embodiment may be interpreted as referring to the 3D data receiving device or the receiving device 110, or to the mesh video decoder (hereinafter referred to as the decoder) 113.
[0082] Figure 1 The system can perform video-based dynamic mesh compression and decompression.
[0083] With the advancement of 3D capture, modeling and rendering, users are allowed to access various forms of 3D content (e.g., AR, XR, metaverse and holograms) across multiple platforms and devices. 3D content is becoming more and more complex and realistic in its object representation to provide an immersive experience for users. However, for the generation and use of 3D models, a significant amount of data is required. Among various types of 3D content, 3D meshes are widely used for efficient data utilization and realistic object representation. An embodiment includes a series of processing steps in a system using mesh content.
[0084] First, the method of compressing dynamic mesh data starts with the video-based point cloud compression (V-PCC) standard technology for point cloud data. Point cloud data is data with color information in the coordinates (X, Y, Z) of the vertex (or point). In the present disclosure, the vertex coordinates (i.e., position information) are referred to as geometric information, and the color information about the vertex is referred to as attribute information. The geometric information and attribute information are collectively referred to as vertex information or point cloud data. Mesh data refers to vertex information including connection information between vertices. The content may be initially created in the form of mesh data. Alternatively, connection information may be added to the point cloud data, and the point cloud data may be transformed into mesh data.
[0085] Currently, the MPEG standard group defines two data types for dynamic mesh data: Category 1 mesh data having a texture map as color information and Category 2 mesh data having vertex colors as color information.
[0086] Standardization of grid encoding for Category 1 data is currently underway, with standardization of Category 2 data expected to follow. Figure 1 As shown, the overall process for providing grid content services may include acquisition, encoding, transmission, decoding, rendering and / or feedback processes.
[0087] In order to provide mesh content services, 3D data acquired by multiple cameras or special cameras can be processed into a mesh data type through a series of steps to generate a video. The generated mesh video can be sent through a series of operations, and the receiving side can process the received data back into a mesh video for rendering. Through this process, the mesh video can be provided to the user, allowing the user to interactively utilize the mesh content according to his or her intention.
[0088] like Figure 1 As shown, the grid compression system may include a sending device 100 and a receiving device 110. The sending device 100 may encode the grid video to output a bit stream, which may be transmitted to the receiving device 110 via a digital storage medium or a network in the form of a file or a stream (stream segment). The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD.
[0089] In the transmitting device 100, the encoder may be referred to as a grid video / image / picture / frame encoding device. In the receiving device 110, the decoder may be referred to as a grid video / image / picture / frame decoding device. The transmitter may be included in a grid video encoder and the receiver may be included in a grid video decoder. The renderer 114 may include a display, and the renderer and / or the display may be configured as a separate device or an external component. The transmitting device 100 and the receiving device 110 may also include a separate internal or external module / unit / component for the feedback process.
[0090] Mesh data uses multiple polygons to represent the surface of an object. Each polygon is defined by vertices in 3D space and connection information indicating how the vertices are connected. In addition, vertex attributes such as color and normal vectors may be included in the data. Mapping information that allows the mesh surface to be mapped onto a 2D plane may also be included in the mesh attributes. The mapping is typically described using a set of parametric coordinates associated with the mesh vertices (called UV coordinates or texture coordinates). The mesh contains a 2D attribute map that can be used to store high-resolution attribute information such as textures, normals, and displacements. Here, displacement can be used interchangeably with displacement information or displacement vectors.
[0091] The mesh video acquisition unit 101 may include processing the 3D object data acquired by a camera or the like into a mesh data type having the above-mentioned attributes through a series of operations, and generating a video composed of the mesh data. In the mesh video, the attributes of the mesh (e.g., vertices, polygons, connections between vertices, colors, and normals) may change over time. The mesh video having attributes and connection information that change over time is called a dynamic mesh video.
[0092] The grid video encoder 102 can encode the input grid video into one or more video streams. The video may contain multiple frames, and each frame may correspond to a still image / picture. In the present disclosure, the grid video may include a grid image / frame / picture. The term "grid video" can be used interchangeably with the grid image / frame / picture. The grid video encoder 102 can perform a video-based dynamic grid (V-Mesh) compression process. For compression and coding efficiency, the grid video encoder 102 can perform a series of processes such as prediction, transformation, quantization and entropy coding. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0093] The file / segment encapsulation module 103 can encapsulate the encoded grid video data and / or grid video related metadata in the form of a file or the like. The grid video related metadata can be received from a metadata processor. The metadata processing unit may be included in the grid video encoder 102, or may be configured as a separate component / module. The file / segment encapsulation module 103 can encapsulate the data into a file format such as ISOBMFF, or process it into a form such as a DASH segment. According to an embodiment, the file / segment encapsulator 103 may include grid video related metadata in a file format. For example, the grid video metadata may be included in boxes at various levels in the ISOBMFF file format, or as data on a separate track in a file. In some embodiments, the file / segment encapsulator 103 may encapsulate the grid video related metadata into a file.
[0094] The transmission processor may apply processing to the encapsulated mesh video data based on the file format for transmission. The transmission processor may be included in the transmitter 104 or implemented as a separate component / module. The transmission processor may process the mesh video data according to any transmission protocol. The processing for transmission may include transmission processing via a broadcast network and transmission processing via broadband. In some embodiments, the transmission processor may receive mesh video related metadata and mesh video data from a metadata processor and process it for transmission.
[0095] The transmitter 104 may transmit the encoded video / image information or data output in the form of a bit stream to the receiver 111 of the receiving device 110 via a digital storage medium or a network in the form of a file or stream. The digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitter 104 may include an element for generating a media file in a predetermined file format, and may include an element for transmission via a broadcast / communication network. The receiver 111 may extract the bit stream and transmit it to a decoding device.
[0096] The receiver 111 may receive the mesh video data transmitted by the mesh data transmitting device. The receiver 111 may receive the mesh video data via a broadcast network or a broadband network, or may receive the mesh video data via a digital storage medium, depending on a channel used for transmission.
[0097] The receiving processor may perform processing on the received mesh video data according to the transmission protocol. The receiving processor may be included in the receiver 111, or may be configured as a separate component / module. In order to correspond to the processing performed for transmission on the transmitting side, the receiving processor may perform the inverse process of the operation of the above-mentioned transmitting processor. The receiving processor may transmit the acquired mesh video data to the file / segment decapsulator 112 and transmit the acquired mesh video related metadata to the metadata parser. The mesh video related metadata acquired by the receiving processor may be in the form of a signaling table.
[0098] The file / segment decapsulator 112 can decapsulate the grid video data in the form of a file received from the receiving processor. The file / segment decapsulator 112 can decapsulate the file etc. according to ISOBMFF etc. to obtain a grid video bitstream or grid video related metadata (metadata bitstream). The obtained grid video bitstream can be transmitted to the grid video decoder 113, and the obtained grid video related metadata (metadata bitstream) can be transmitted to the metadata processor. The grid video bitstream may include metadata (metadata bitstream). The metadata processor may be included in the grid video decoder 113, or may be configured as a separate component / module. The grid video related metadata obtained by the file / segment decapsulator 112 may be in the form of a box or track in the file format. When necessary, the file / segment decapsulator 112 may receive the metadata required for decapsulation from the metadata processor. The grid video related metadata may be transmitted to the grid video decoder 113 for use in the grid video decoding process, or transmitted to the renderer 114 for use in the grid video rendering process.
[0099] The grid video decoder 113 may receive an input bitstream and perform an inverse operation corresponding to the operation of the grid video encoder 102 to decode the video / image. The decoded grid video / image may be displayed through a display of the renderer 114. The user may view all or part of the rendering result through a VR / AR display, a general display, etc.
[0100] The feedback process may include sending various types of feedback information that may be obtained during the rendering / display operation to a decoder on the sending side or the receiving side. The feedback process may provide interactivity when consuming the mesh video. In some embodiments, the feedback process may include sending head orientation information, viewport information indicating the area the user is currently viewing, etc. In some embodiments, the user may interact with objects implemented in the VR / AR / MR / autonomous driving environment. In this case, information related to the interaction may be transmitted to the sending side or the service provider during the feedback process. In some embodiments, the feedback process may be skipped.
[0101] The head orientation information may refer to information about the user's head position, angle, movement, etc. Based on this information, information about the area within the grid video that the user is currently viewing (ie, viewport information) may be calculated.
[0102] The viewport information may be information about the area of the grid video that the user is currently viewing. Based on this information, gaze analysis may be performed to determine how the user consumes the grid video, how long the user looks at a specific area of the grid video, etc. Gaze analysis may be performed on the receiving side, and the results may be transmitted to the sending side via a feedback channel. Devices such as VR / AR / MR displays may extract the viewport area based on the user's head position / orientation, the vertical or horizontal FOV supported by the device, etc.
[0103] In some embodiments, the feedback information may not only be transmitted to the transmitter, but also consumed on the receiving side. In other words, operations such as decoding and rendering may be performed on the receiving side based on the feedback information. For example, based on the head orientation information and / or the viewport information, only the grid video of the area currently being viewed by the user may be preferentially decoded and rendered.
[0104] The present disclosure relates to implementations of dynamic mesh video compression as described above. The methods / implementations disclosed herein can be applied to the Video-Based Dynamic Mesh Compression (V-Mesh) standard of the Moving Picture Experts Group (MPEG) or any next-generation video / image coding standard. Dynamic mesh video compression is a method for processing mesh connectivity information and attributes that vary over time. It can perform lossy and lossless compression for various applications such as real-time communications, storage, free viewpoint video, and AR / VR.
[0105] The dynamic mesh video compression method described below is based on the V-mesh method of MPEG.
[0106] In the present disclosure, a picture / frame may generally refer to a unit representing one image at a specific time.
[0107] A pixel or a picture element may refer to the smallest unit constituting a picture (or video). In addition, the term "sample" may be used as a term corresponding to a pixel. A sample may generally indicate a pixel or a pixel value. It may indicate only a pixel / pixel value of a luma component, or may indicate only a pixel / pixel value of a chroma component, or may indicate only a pixel / pixel value of a depth component.
[0108] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. In some cases, the term unit may be used interchangeably with terms such as block or area. Typically, an M×N block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0109] As mentioned above, Figure 1 The encoding process is performed as follows.
[0110] In other words, the video-based dynamic mesh compression (V-Mesh) compression method can provide a method for compressing dynamic mesh video data based on a 2D video codec (e.g., High Efficiency Video Surface (HEVC) and Versatile Video Coding (VVC)). In the V-Mesh compression process, the following data is received as input and compressed.
[0111] Input mesh: includes the 3D coordinates of the vertices that make up the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connections between the vertices that make up the surface. The mesh surface can be represented by triangles or other polygons, and the connection information between the vertices that make up the surface is stored according to a predetermined shape. The input mesh can be stored in the OBJ file format.
[0112] Attribute map (texture map is also used interchangeably below): contains information about the attributes of the mesh (color, normal, displacement, etc.), and stores the data in the form of a mapping of the mesh surface to a 2D image. The mapping that indicates which part of the mesh (surface or vertex) corresponds to each piece of data in the attribute map is based on the mapping information contained in the input mesh. Since the attribute map has data about each frame of the mesh video, it can also be called an attribute map video. The attribute map in the V-Mesh compression method mainly contains color information about the mesh and is stored in an image file format (PNG, BMP, etc.).
[0113] Material library file: Contains information about the material properties used in the mesh, especially information linking the input mesh to the corresponding property map. It is stored in the Wavefront Material Template Library (MTL) file format.
[0114] In the V-Mesh compression method, the following data and information may be generated through the compression process.
[0115] Base Mesh: Objects in the input mesh are represented using the minimum number of vertices determined according to user criteria by simplifying the input mesh through a preprocessing process.
[0116] Displacement: Used to represent the displacement information of the input mesh as similarly as possible to the base mesh, expressed in 3D coordinates.
[0117] Atlas information: Metadata required to reconstruct a mesh using base mesh, displacement and attribute map information. It can be used to generate and use sub-units of a mesh (sub-meshes, patches, etc.).
[0118] Reference Figures 2 to 7 Describes the method of encoding mesh position information (or vertex position information), see Figures 6 to 10 etc. describe a method of reconstructing grid position information to encode attribute information (attribute map).
[0119] Figure 2 A V-MESH compression method according to an embodiment is shown.
[0120] Figure 2 Show Figure 1 The encoding process may include a preprocessing process and an encoding process. Figure 2 As shown, Figure 1 The grid video encoder 102 may include a preprocessor 200 and an encoder 201. In addition, Figure 1 The sending device can be broadly called an encoder. Figure 1 The grid video encoder 102 may be referred to as an encoder. Figure 2 As shown, the V-Mesh compression method may include preprocessing 200 and encoding 201 . Figure 2 The preprocessor 200 may be located at Figure 2 The front end of the encoder 201. Figure 2 The pre-processor 200 and the encoder 201 may be referred to as a single encoder.
[0121] The preprocessor 200 may receive a static or dynamic mesh (M(i)) and / or an attribute map (A(i)). The preprocessor 200 may generate a base mesh m(i) and / or a displacement d(i) by preprocessing. The preprocessor 200 may receive feedback information from the encoder 201, and may generate a base mesh and / or a displacement based on the feedback information.
[0122] The encoder 201 may receive a base grid m(i), a displacement d(i), a static or dynamic grid M(i), and / or an attribute map A(i). In the present disclosure, at least one of the base grid m(i), the displacement d(i), the static or dynamic grid M(i), and / or the attribute map A(i) may be referred to herein as grid-related data. The encoder 201 may encode the grid-related data to generate a compressed bitstream.
[0123] Figure 3 Preprocessing in V-MESH compression according to an embodiment is shown.
[0124] Figure 3 Show Figure 2 The configuration and operation of the preprocessor. Figure 3 In the example, the input mesh may include a static or dynamic mesh M(i) and / or an attribute map A(i). The input mesh may also include 3D coordinates of vertices constituting the mesh, normal information about each vertex, mapping information for mapping the mesh surface to a 2D plane, and connection information between vertices constituting the surface.
[0125] Figure 3 200. The process of performing preprocessing on an input mesh is shown. Preprocessing 200 may include four operations: 1) Group of Frames (GoF) generation, 2) mesh simplification, 3) UV parameterization, and 4) fitting subdivision surface (300). According to an embodiment, GoF generation may be referred to as a GoF generation process or a GoF generator, mesh simplification may be referred to as a mesh simplification process or a mesh simplification part, UV parameterization may be referred to as a UV parameterization process or a UV parameterization part, and fitting subdivision surface may be referred to as a fitting subdivision surface process or a fitting subdivision surface part. The preprocessor 200 may generate a displacement and / or base mesh from the received input mesh and transmit it to the encoder 201. The preprocessor 200 may transmit GoF information related to GoF generation to the encoder 201.
[0126] Below, describe Figure 3 of each operation.
[0127] GoF generation: The process of generating a reference structure of mesh data. When the mesh of the previous frame and the current mesh have the same number of vertices, the same number of texture coordinates, the same vertex connection information, and the same texture coordinate connection information, the previous frame can be set as a reference frame. In other words, if only the vertex coordinate values are different between the current input mesh and the reference input mesh, the encoder 201 can perform inter-frame encoding. Otherwise, it performs intra-frame encoding for the frame.
[0128] Mesh Simplification: The process of simplifying an input mesh to create a simplified mesh, called a base mesh. Vertices to be removed from the original mesh can be selected based on user-defined criteria, and the selected vertices and triangles connected to the selected vertices can then be removed.
[0129] In the process of performing mesh simplification, the voxelized input mesh, the target triangle ratio (TTR) and the minimum triangle component (CCCount) information can be transmitted as input, and a simplified mesh can be obtained as output. In this process, connected triangle components that are smaller than the set minimum triangle component (CCCount) can be removed.
[0130] UV parameterization: The process of mapping a 3D surface to the texture domain of a simplified mesh. Parameterization can be performed using the UVAtlas tool. This process generates mapping information that indicates where each vertex of the simplified mesh can be mapped to on the 2D image. The mapping information is represented as texture coordinates and stored, and the final base mesh is generated through this process.
[0131] Fitting subdivision surface (300): A process of performing subdivision on a simplified mesh (i.e., a simplified mesh with texture coordinates). The displacement and base mesh generated by the process are output to the encoder 201. A user-defined method (e.g., edge midpoint method) may be applied as a subdivision method. The fitting process is performed so that the input mesh and the subdivided mesh become similar to each other. The mesh on which the fitting process is performed will be referred to herein as a fitted subdivided mesh.
[0132] Figure 4 An edge midpoint subdivision method according to an embodiment is shown.
[0133] Figure 4 Show reference Figure 3 Describes the edge midpoint subdivision method for fitting subdivision surfaces. Figure 4 , the original mesh containing four vertices is subdivided to create sub-meshes. The sub-meshes can be created by creating new vertices in the middle of the edges between the vertices. Then, a fitting process is performed to make the input mesh and the sub-mesh similar to each other, resulting in a fitted subdivided mesh.
[0134] Once the fitted subdivision mesh is generated, the displacement is calculated based on the result and the previously compressed and decoded base mesh (hereinafter referred to as the reconstructed base mesh). In other words, the reconstructed base mesh is subdivided in the same way as the fitted subdivision surface. The position difference between the result and each vertex in the fitted subdivision mesh is the displacement of each vertex. Since the displacement represents the position difference in 3D space, it is expressed as a value in the (x, y, z) space in the Cartesian coordinate system. According to the user input parameters, the (x, y, z) coordinate value can be converted to the (normal, tangent, bitangent) coordinate value in the local coordinate system.
[0135] Figure 5 A displacement generation process according to an embodiment is shown. Figure 5 The displacement generation process may be performed by the preprocessor 200 , or may be performed by the encoder 201 .
[0136] Figure 5 Detailed description for reference Figure 4 The fitting of the subdivision surface 300 is described to calculate the displacement.
[0137] The encoder and / or preprocessor according to the embodiment may include 1) a subdivider, 2) a local coordinate system calculator and 3) a displacement vector calculator. The subdivider may perform subdivision on the reconstructed base mesh to generate a subdivided reconstructed base mesh. Here, the reconstruction of the base mesh may be performed by the preprocessor 200, or may be performed by the encoder 201. The local coordinate system calculator may receive a fitted subdivided mesh and a subdivided reconstructed base mesh, and may transform a coordinate system associated with the mesh into a local coordinate system based on the received mesh. The local coordinate system calculation may be optional. The displacement calculator calculates the position difference between the fitted subdivided mesh and the subdivided reconstructed base mesh. For example, it may generate the position difference between vertices in two input meshes. The position difference between the vertices is the displacement.
[0138] The mesh data transmission method and device according to the embodiment may encode the mesh data as follows. Mesh data is a term including point cloud data. Point cloud data (which may be simply referred to as point cloud) according to the embodiment may refer to data including vertex coordinates (also referred to as geometric information) and color information (also referred to as attribute information). In addition, geometric images, attribute images, occupancy maps, and auxiliary information (also referred to as patch information) generated by patch generation and packaging based on vertex coordinates and color information may also be referred to as point cloud data. Therefore, point cloud data including connection information may be referred to as mesh data. The terms point cloud and mesh data may be used interchangeably herein.
[0139] According to an embodiment, the V-Mesh compression (reconstruction) method may include intra-frame coding ( Figure 6 ) and inter-frame coding ( Figure 7 ).
[0140] Based on the results generated by the GoF, intra-frame coding or inter-frame coding is performed. In intra-frame coding, the data to be compressed can be base grids, displacements, attribute maps, etc. In inter-frame coding, the data to be compressed can be displacements, attribute maps, and the motion field between the reference base grid and the current base grid.
[0141] Figure 6 The intra-frame encoding process in the V-MESH compression method according to the embodiment is shown. Figure 6 The various components of the intra-frame encoding process correspond to hardware, software, processor and / or a combination thereof.
[0142] Figure 6 The encoding process is described in detail Figure 1 That is, it represents the encoding of the grid video encoder 102. Figure 1The encoding is the configuration of the grid video encoder 102 during intra-frame encoding. Figure 6 The encoder may include a pre-processor 200 and / or an encoder 201 . Figure 6 The preprocessor 200 and the encoder 201 may correspond to Figure 3 A preprocessor 200 and an encoder 201 are provided.
[0143] The preprocessor 200 may receive an input mesh and perform the above-mentioned preprocessing, and may generate a base mesh and / or a fitted subdivided mesh through the preprocessing.
[0144] The quantizer 411 of the encoder 201 may quantize the base grid and / or the fitted subdivided grid. The static grid encoder 412 may encode the static grid (i.e., the quantized base grid) and generate a bitstream containing the encoded base grid (i.e., a compressed base grid bitstream). The static grid decoder 413 may decode the encoded static grid (i.e., the encoded base grid). The inverse quantizer 414 may inverse quantize the quantized static grid (i.e., the base grid) and output a reconstructed (restored) base grid. The displacement calculator 415 may generate displacements based on the reconstructed static grid (i.e., the base grid) and the fitted subdivided grid. According to an embodiment, the displacement calculator 415 subdivides the reconstructed base grid and then calculates the displacement, i.e., the position difference of each vertex between the subdivided base grid and the fitted subdivided grid. In other words, when the fitted subdivided grid is similar to the original grid, the displacement is a displacement vector that is the position difference between the vertices in the two grids. The forward linear lifter 416 may perform a lifting transform on the input displacement to generate a lifting coefficient (also called a transform coefficient). The quantizer 417 may quantize the lifting coefficient. The image packer 418 may pack the image based on the quantized lifting coefficients. The video encoder 419 may encode the packed image. That is, the quantized lifting coefficients are packed into frames as 2D images by the image packer 418, compressed by the video encoder 419, and output as a displacement bitstream (i.e., a compressed displacement bitstream).
[0145] The video decoder 420 decodes the compressed displacement bitstream. The image unpacker 421 may perform unpacking on the decoded displacement frame to output quantization lifting coefficients. The inverse quantizer 422 may inverse quantize the quantization lifting coefficients. The inverse linear lifting unit 423 applies inverse lifting to the inverse quantization lifting coefficients to generate a reconstructed displacement. The mesh reconstructor 424 restores the reconstructed and deformed mesh based on the reconstructed displacement output from the inverse linear lifting unit 423 and the reconstructed base mesh (also referred to as the subdivided reconstructed base mesh) output from the inverse quantizer 414. The reconstructed and deformed mesh is referred to herein as a reconstructed deformed mesh.
[0146] The attribute transfer 425 receives an input mesh and / or an input attribute map and regenerates the attribute map based on the reconstructed deformed mesh. The attribute map refers to a texture map corresponding to the attribute information in the mesh data component. In the present disclosure, the terms attribute map and texture map are used interchangeably. The push-pull fill unit 426 can fill data into the attribute map based on a push-pull method. The color space converter 427 can convert the space of the color component of the attribute map. For example, the attribute map can be converted from an RGB color space to a YUV color space. The video encoder 428 can encode the attribute map to output a compressed attribute bitstream.
[0147] The multiplexer 430 may multiplex the compressed base grid bitstream, the compressed displacement bitstream, and the compressed attribute bitstream to generate a compressed bitstream.
[0148] exist Figure 6 In the embodiment, the displacement calculator 415 may be included in the preprocessor 200. In addition, at least one of the quantizer 411, the static grid encoder 412, the static grid decoder 413, or the inverse quantizer 414 may be included in the preprocessor 200.
[0149] like Figure 6 As described in , the intra-frame encoding method includes base mesh encoding (also referred to as static mesh encoding). That is, when intra-frame encoding is performed on the current input mesh frame, the base mesh generated during the preprocessing of the preprocessor 200 may be quantized by the quantizer 411 and then encoded by the static mesh encoder 412 using the static mesh compression technology. For example, in the V-Mesh compression method, the Draco technology is applied to encode the base mesh, and vertex position information, mapping information (texture coordinates), vertex connection information, etc. related to the base mesh are compressed.
[0150] Figure 6 The encoder in compresses the base grid, displacements, and attributes in the frame to generate a bitstream, while Figure 7 The encoder in compresses the motion, displacement, and attributes between the current frame and the reference frame to generate a bitstream.
[0151] Figure 7 The inter-frame encoding process in the V-MESH compression method according to the embodiment is shown. Figure 7 The various components of the inter-frame encoding process correspond to hardware, software, a processor and / or a combination thereof.
[0152] Figure 7 The encoding process is described in detail. Figure 1 That is, it means when Figure 1 The encoding is the configuration of the encoder during inter-frame coding. Figure 7 The encoder may include a pre-processor 200 and / or an encoder 201 . Figure 7The preprocessor 200 and the encoder 201 may correspond to Figure 3 A preprocessor 200 and an encoder 201 are provided.
[0153] For Figure 6 The encoding operation corresponds to Figure 7 Components of the encoding operation, refer to Figure 6 That is, Figure 7 The operations of the quantizer 511, displacement calculator 515, wavelet transformer 516, quantizer 517, image packer 518, video encoder 519, video decoder 520, image unpacker 521, inverse quantizer 522 and inverse wavelet transformer 523, grid reconstructor 524, attribute transfer 525, push-pull fill 526, color space converter 527, video encoder 528 and multiplexer 530 are the same as those described above. Figure 6 The operations of the quantizer 411, static grid encoder 412, static grid decoder 413 and inverse quantizer 414, displacement calculator 415, forward linear lifting unit 416, quantizer 417, image packer 418, video encoder 419, video decoder 420, image unpacker 421, inverse quantizer 422, inverse linear lifting unit 423 and grid reconstruction 424, attribute transfer 425, push-pull filling 426, color space converter 427, video encoder 428 and multiplexer 430 are the same or similar, so the ... linear lifting unit 423 and grid reconstruction 424, attribute transfer 425, push-pull filling 426, color space converter 427, video encoder 428 and multiplexer 430 are the same or similar, so the operations of the quantizer 411, static grid encoder 412, static grid decoder 413 and inverse linear lifting unit 423 and grid reconstruction 424, attribute transfer 425, push-pull filling 426, color space converter 427, video encoder 428 and multiplexer 4 Figure 7 A detailed description is not given here to avoid redundancy.
[0154] exist Figure 7 In the embodiment of the present invention, for inter-frame based coding, the motion encoder 512 can obtain and encode the motion vector between the reconstructed quantized reference base grid and the quantized current base grid, and output the compressed motion bitstream. The motion encoder 512 can be called a motion vector encoder. The base grid reconstructor 513 can reconstruct the base grid based on the reconstructed quantized reference base grid and the encoded motion vector. The reconstructed base grid is inverse quantized by the inverse quantizer 514 and output to the displacement calculator 515.
[0155] exist Figure 7 In the embodiment, the displacement calculator 515 may be included in the preprocessor 200. In addition, at least one of the quantizer 511, the motion encoder 512, the base grid reconstructor 513, or the inverse quantizer 514 may be included in the preprocessor 200.
[0156] As reference Figure 7As described, the inter-frame coding method may include motion field coding (also known as motion vector coding). When the reference mesh and the current input mesh have a one-to-one vertex correspondence and only the position information about the vertices differs between them, inter-frame coding may be performed. When inter-frame coding is performed, the base mesh may not be compressed. Instead, the difference between the vertices of the reference base mesh and the current base mesh, that is, the motion field (or motion vector), may be calculated and encoded. The reference base mesh is the result of quantizing the decoded base mesh data and is determined by the reference frame index determined in the GoF generation. The motion field may be encoded as is. Alternatively, the predicted motion field may be calculated by averaging the motion fields of the reconstructed vertices among the vertices connected to the current vertex, and the residual motion field may be encoded as the difference between the value of the predicted motion field and the value of the motion field of the current vertex. The value of the residual motion field may be encoded using entropy coding. In addition to the motion field coding in the inter-frame coding, the process of encoding the displacement and attribute map is the same as the structure of the intra-frame coding method except for the base mesh coding.
[0157] Figure 8 A lifting transformation process of displacement according to an embodiment is shown.
[0158] Fig. 9 A process of packing transform coefficients (also called lifting coefficients) into a 2D image according to an embodiment is shown.
[0159] Figure 8 and Fig. 9 Shown separately Figure 6 and Figure 7 The process of transforming displacement and packing transform coefficients during the encoding process.
[0160] The encoding method according to an embodiment includes displacement encoding.
[0161] After base mesh encoding and / or motion field encoding, a reconstructed base mesh may be generated by reconstruction and inverse quantization, and displacement may be calculated between the subdivision result of the reconstructed base mesh and the fitted subdivision mesh generated by fitting the subdivision surface (see Figure 6 415 or Figure 7 515 in ). A data transformation process (eg, wavelet transform) may be applied to the displacement information for efficient encoding (see Figure 6 416 or Figure 7 516 in the above table).
[0162] Figure 8 Shown by Figure 6 Forward linear lifting unit 416 or Figure 7The wavelet transformer 516 of FIG. 510 uses a lifting transform to transform the displacement information. For example, a lifting transform based on a linear wavelet may be performed. The transform coefficients generated by the transform process are quantized by the quantizer 417 (or 517) and then packaged into a 2D image by the image packager 418 (or 518), such as Fig. 9 As shown. The transform coefficients can be organized into blocks, one block for every 256 (=16×16) units. The individual blocks can be packed in z-scan order. The number of rows in a block is fixed to 16, but the number of columns in a block can be determined by the number of vertices in the subdivided base grid. Within a block, the transform coefficients can be sorted and packed by Morton code. For packed images, displacement video can be generated per GoF. The displacement video can be encoded by the video encoder 419 (or 519) using a conventional video compression codec.
[0163] Reference Figure 8 , the base mesh (original) may include vertices and edges of LoD0. The first subdivided mesh generated by splitting (or subdividing) the base mesh includes vertices generated by further splitting (or subdividing) the edges of the base mesh. The first subdivided mesh includes vertices of LoD0 and vertices of LoD1. LoD1 includes subdivided vertices and vertices from the base mesh (LoD0). The first subdivided mesh may be split (or subdivided) to generate a second subdivided mesh. The second subdivided mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), LoD1 includes vertices further split (or subdivided) from LoD0, and LoD2 includes vertices further split (or subdivided) from LoD1. LoD is a level of detail indicating how detailed the mesh data content is. As the index of the level increases, the distance between vertices decreases and the level of detail increases. In other words, as the value of LoD decreases, the details of the mesh data content deteriorate. As the value of LoD decreases, the details of the mesh data content increase. LoD N includes the vertices included in LoD N-1. In the case where the mesh (or vertex) is further divided by subdivision, the mesh may be encoded based on a prediction and / or update method taking into account previous vertices v1 and v2 and subdivided vertex v. Instead of encoding the information of the current LoD N as is, a residual relative to the previous LoD N-1 may be generated. Therefore, the residual may be used to encode the mesh to reduce the size of the bitstream. The prediction process refers to the operation of predicting the current vertex v from previous vertices v1 and v2. Since adjacent subdivided meshes have similar data, this property may be utilized for efficient encoding. The current vertex position information is predicted from the residual of the previous vertex position information, and the previous vertex position information is updated by the residual. In the present disclosure, vertices and points may be used interchangeably. LoD may be defined in the subdivision of the base mesh. Depending on the embodiment, the subdivision of the base mesh may be performed by the preprocessor 200, or may be performed by a separate component / module.
[0164] Reference Fig. 9, the vertex has a transformation coefficient (also referred to as a lifting coefficient) generated by the lifting transformation. The transformation coefficient of the vertex related to the lifting transformation can be packed into the image by the image packer 418 (or 518) and then encoded by the video encoder 419 (or 519).
[0165] Fig.10 The attribute transfer process in the V-MESH compression method according to the embodiment is shown.
[0166] According to an embodiment, Fig.10 Show Figure 6 , Figure 7 Detailed operation of attribute transfer 425 (or 525) in the encoding of etc.
[0167] According to an embodiment, the encoding includes property graph encoding. According to an embodiment, the property graph encoding may be Figure 6 Video encoder 428 or Figure 7 The video encoder 528 executes.
[0168] According to an embodiment, in the present disclosure, the encoder compresses information about the input mesh through base mesh encoding (i.e., intra-frame encoding), motion field encoding (i.e., inter-frame encoding), and displacement encoding. The compressed input mesh in the encoding process is reconstructed through base mesh decoding (intra-frame), motion field decoding (inter-frame), and displacement video decoding, and as a result of the reconstruction, the reconstructed deformed mesh (hereinafter referred to as the reconstructed deformed mesh) is used to compress the input attribute map, such as Figure 6 and Figure 7 The reconstructed deformed mesh has position information about vertices, texture coordinates, and corresponding connectivity information, but no color information corresponding to the texture coordinates. Fig.10 As shown, in the V-Mesh compression method, a new attribute map having color information corresponding to the texture coordinates of the reconstructed deformed mesh is regenerated through the attribute transfer process of attribute transfer 425 (or 525).
[0169] According to an embodiment, attribute transfer 425 (or 525) first checks for each point P (u, v) in the 2D texture domain whether the corresponding vertex is within the texture triangle of the reconstructed deformed mesh. When the corresponding vertex is in the texture triangle T, the attribute transfer calculates the barycentric coordinates (α, β, γ) of P (u, v) according to the triangle T. Then, it calculates the 3D coordinates M (x, y, z) of P (u, v) based on the 3D vertex position and (α, β, γ) of the triangle T. The vertex coordinates M' (x', y', z') corresponding to the position closest to the calculated M (x, y, z) and the triangle T' containing the vertex are searched in the input mesh domain. Then, the barycentric coordinates (α', β', γ') of M' (x', y', z') in the triangle T' are calculated. The texture coordinates (u', v') are calculated based on the texture coordinates and (α', β', γ') corresponding to the three vertices of the triangle T', and the color information corresponding to the coordinates is searched in the input attribute map. The color information found in this manner is then assigned to the (u, v) pixel position in the new input attribute map. If P(u, v) does not belong to any triangle, a filling algorithm (e.g., the push-pull algorithm of push-pull filling 426 (or 526)) is used to fill the pixel at that position in the new input attribute map with the color value.
[0170] The new property graph generated by the property transfer 425 (or 525) is bundled into the GoF to construct a property graph video, which is compressed using the video codec of the video encoder 428 (or 528).
[0171] Available from Fig.10 See the reference relationship between the input mesh, input attribute map, reconstructed deformed mesh and reconstructed attribute map.
[0172] Figure 1 The decoding process can be performed Figure 1 Specifically, the decoding process is performed as disclosed below.
[0173] Fig.11 The intra-frame decoding process of the V-Mesh technology according to an embodiment is shown.
[0174] Fig.11 Show Figure 1 The configuration and operation of the grid video decoder 113 of the receiving device. In addition, Fig.11 It can be shown that by executing Figure 6 The mesh data is reconstructed by the inverse process of the intra-frame encoding process. Fig.11 The various components of the intra-frame decoding process correspond to hardware, software and / or a combination thereof.
[0175] First, the bitstream (i.e., compressed bitstream) received and input to the demultiplexer 611 of the intra decoder 610 may be separated into a mesh substream, a displacement substream, an attribute map substream, and a substream containing patch information about the mesh (e.g., V-PCC / V3C). The term V-PCC (video-based point cloud compression) used in the present disclosure may have the same meaning as V3C (visual volumetric video-based coding). The two terms may be used interchangeably. Therefore, in the present disclosure, the term V-PCC may be interpreted as V3C.
[0176] According to an embodiment, the grid substream may be input to and decoded by the static grid decoder 612 , the displacement substream may be input to and decoded by the video decoder 613 , and the property map substream may be input to and decoded by the video decoder 617 .
[0177] According to an embodiment, the mesh substream may be decoded by a decoder 612 of a static mesh codec (e.g., Google Draco) used in encoding to reconstruct connection information, vertex geometry information, vertex texture coordinates, etc. related to the decoding result of the reconstructed quantized base mesh (e.g., the reconstructed base mesh).
[0178] According to an embodiment, the displacement substream may be decoded into a displacement video by a decoder 613 of a video compression codec used in encoding. Then, image depacketization is performed by an image depacketizer 614, inverse quantization is performed by an inverse quantizer 615, and inverse transformation is performed by an inverse linear enhancement unit 616 to reconstruct displacement information about each vertex (i.e., reconstruct displacement).
[0179] According to an embodiment, the base grid reconstructed by the static grid decoder 612 is inversely quantized by the inverse quantizer 620 and output to the grid reconstructor 630. The grid reconstructor 630 reconstructs the reconstructed deformed grid (i.e., the decoded grid) based on the reconstructed displacement output from the inverse linear lifting unit 616 and the reconstructed base grid output from the inverse quantizer 620. In other words, the inversely quantized reconstructed base grid is combined with the reconstructed displacement information to generate a final decoded grid. In the present disclosure, the final decoded grid is referred to as a reconstructed deformed grid.
[0180] According to an embodiment, the property map substream is decoded by a decoder 617 corresponding to the video compression codec used in encoding, and then a color transformer 640 reconstructs a final property map (ie, a decoded property map) through color format conversion, color space conversion, etc.
[0181] According to an embodiment, the reconstructed decoded mesh and the decoded property map may be used as final mesh data available to a user at the receiving side.
[0182] Reference Fig.11, the received compressed bitstream includes patch information, a grid substream, a displacement substream and an attribute map substream. The term substream is interpreted to refer to a portion of a bitstream included in a bitstream. The bitstream contains patch information (data), grid information (data), displacement information (data) and attribute map information (data).
[0183] As mentioned above, Fig.11 The decoder performs intra-frame decoding as follows. The static grid decoder 612 decodes the grid substream to generate a reconstructed quantized base grid, and the inverse quantizer 620 inversely applies the quantization parameters of the quantizer to generate the reconstructed base grid. The video decoder 613 decodes the displacement substream, the image unpacker 614 unpacks the image of the decoded displacement video, and the inverse quantizer 615 inversely quantizes the quantized image. The inverse linear lifting unit 616 applies the lifting transform in the inverse process of the encoder to generate the reconstructed displacement. The grid reconstructor 630 generates a reconstructed deformed grid based on the reconstructed base grid and the reconstructed displacement. The video decoder 617 decodes the attribute map substream, and the color transformer 64 transforms the color format and / or space of the decoded attribute map to generate a decoded attribute map.
[0184] Fig.12 The inter-frame decoding process of the V-Mesh technology is shown.
[0185] Fig.12 Show Figure 1 Configuration and operation of the grid video decoder 113 of the receiving device. Fig.12 In the Figure 7 The mesh data is reconstructed by the inverse process of the inter-frame coding process. Fig.12 The various components of the intra-frame decoding process correspond to hardware, software and / or a combination thereof.
[0186] First, the bitstream received and input to the demultiplexer 711 of the intra decoder 710 can be separated into a motion substream (also called a motion vector substream), a displacement substream, a property map substream, and a substream containing patch information about the grid (e.g., V3C / V-PCC).
[0187] According to an embodiment, the motion substream may be input to and decoded by the motion decoder 712 , the displacement substream may be input to and decoded by the video decoder 713 , and the property map substream may be input to and decoded by the video decoder 717 .
[0188] According to an embodiment, the motion substream is decoded by the motion decoder 712 through entropy decoding and inverse prediction to reconstruct motion information (also called motion vector information). The base grid reconstructor 718 combines the reconstructed motion information with the pre-reconstructed and stored reference base grid to generate a reconstructed quantized base grid for the current frame. The inverse quantizer 720 applies inverse quantization to the reconstructed quantized base grid to generate a reconstructed base grid. The video decoder 713 decodes the displacement substream, the image unpacker 714 unpacks the image of the decoded displacement video, and the inverse quantizer 715 inverse quantizes the quantized image. The inverse linear lifting unit 716 applies the lifting transform in the inverse process of the encoder to generate a reconstructed displacement. The grid reconstructor 730 generates a reconstructed deformed grid (i.e., the final decoded grid) based on the reconstructed base grid and the reconstructed displacement.
[0189] According to an embodiment, the video decoder 717 decodes the attribute map substream in the same manner as intra-frame decoding, and the color converter 740 converts the color format and / or space of the decoded attribute map to generate a decoded attribute map. The decoded grid and the decoded attribute map can be used as final grid data available to the user at the receiving side.
[0190] Reference Fig.12 , the bitstream contains motion information (also called motion vectors), displacements, and attribute maps. Because inter-frame decoding is performed, Fig.12 The process also includes decoding inter-frame motion information. The reconstructed basic grid is generated by decoding the motion information and generating a reconstructed quantized basic grid for the motion information based on the reference basic grid. Fig.12 Zhongyu Fig.11 The same operation as in Fig.11 Description.
[0191] Fig.13 A mesh data transmitting device according to an embodiment is shown.
[0192] Fig.13 Corresponds to Figure 1 The transmitting device 100 or the grid video encoder 102, Figure 2 , Figure 6 or Figure 7 encoder (preprocessor and encoder) and / or corresponding sending encoding device. Fig.13 The various components correspond to hardware, software, processors and / or a combination thereof.
[0193] The operation process of using V-Mesh compression technology to compress and send dynamic mesh data at the sending end can be as follows: Fig.13 Configuration shown. Fig.13 The sending device can perform intra-frame coding (also known as intra-frame coding or intra-picture coding) and / or inter-frame coding (also known as inter-frame coding or inter-picture coding).
[0194] The preprocessor 811 receives the original mesh and generates a simplified mesh (or base mesh) and a fitting subdivision (or subdivision) mesh. Simplification can be performed based on a target number of vertices or a target number of polygons constituting the mesh. Parameterization can be performed on the simplified mesh to generate texture coordinates and texture connection information for each vertex. For example, parameterization is a process of mapping a 3D surface into a texture domain of a simplified mesh. When parameterization is performed using the UVAtlas tool, mapping information indicating where each vertex of the simplified mesh can be mapped on a 2D image is generated. The mapping information is represented as texture coordinates and stored, and the final base mesh is generated through this process. The mesh information can be quantized from a floating point form to a fixed point form. The result is a base mesh, which can be output to a motion vector encoder 813 or a static mesh encoder 814 through a switch unit 812. The preprocessor 811 can perform mesh subdivision on the base mesh to generate additional vertices. According to the subdivision method, vertex connection information including additional vertices, texture coordinates, and connection information about the texture coordinates can be generated. The preprocessor 811 can generate a fitting subdivision mesh by adjusting the vertex positions so that the subdivided mesh becomes similar to the original mesh.
[0195] According to an embodiment, when inter-coding is performed on a grid frame, the base grid is output to a motion vector encoder 813 through a switch unit 812. When intra-coding is performed on a grid frame, the base grid is output to a static grid encoder 814 through a switch unit 812. The motion vector encoder 813 may be referred to as a motion encoder.
[0196] For example, when intra-coding is performed on the mesh frame, the base mesh may be compressed by the static mesh encoder 814. In this case, connection information, vertex geometry information, vertex texture information, normal information, etc. related to the base mesh may be encoded. The base mesh bitstream generated by the encoding is sent to the multiplexer 823.
[0197] As another example, when inter-frame coding is performed on a grid frame, the motion vector encoder 813 may receive a base grid and a reference reconstructed base grid (or a reconstructed quantized reference base grid) as input, calculate a motion vector between the two grids, and encode its value. In addition, the motion vector encoder 813 may perform prediction based on connection information using a previously encoded / decoded motion vector as a predictor, and encode a residual motion vector obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by the encoding is sent to the multiplexer 823.
[0198] The base mesh reconstructor 815 may receive a base mesh encoded by the static mesh encoder 814 or a motion vector encoded by the motion vector encoder 813, and generate a reconstructed base mesh. For example, the base mesh reconstructor 815 may perform static mesh decoding on the base mesh encoded by the static mesh encoder 814 to reconstruct the base mesh. In this case, quantization may be applied before static mesh decoding, and inverse quantization may be applied after static mesh decoding. In another example, the base mesh reconstructor 815 may reconstruct the base mesh based on the reconstructed quantized reference base mesh and the motion vector encoded by the motion vector encoder 813. The reconstructed base mesh is output to the displacement calculator (or displacement vector calculator) 816 and the mesh reconstructor 820.
[0199] The displacement calculator 816 may perform mesh subdivision on the reconstructed base mesh. The displacement calculator 816 may calculate a displacement vector, which is a value of a vertex position difference between the subdivided reconstructed base mesh and the fitted subdivided (or subdivided) mesh generated by the preprocessor 811. In this case, as many displacement vectors as vertices in the subdivided mesh may be calculated. The displacement calculator 816 may transform the displacement vector calculated in the 3D Cartesian coordinate system to the local coordinate system based on the normal vector of each vertex.
[0200] The displacement vector video generator 817 may include a linear lifting part, a quantizer, and an image packer. That is, in the displacement vector video generator 817, the linear lifting unit may transform the displacement vector for efficient encoding. According to an embodiment, the transformation may be a lifting transformation, a wavelet transformation, etc. In addition, the quantizer may perform quantization on the transformed displacement vector value (i.e., the transform coefficient). In this case, different quantization parameters may be applied to the axes of the transform coefficients respectively. The quantization parameters may be derived by convention between the encoder / decoder. After transformation and quantization, the displacement vector information may be packed into a 2D image by the image packer. The displacement vector video generator 817 may generate a displacement vector video by grouping the packed 2D images of each frame. A displacement vector video may be generated for each group of frames (GoF) of an input grid.
[0201] The displacement vector video encoder 818 may encode the generated displacement vector video using a video compression codec. The generated displacement vector video bitstream is sent to a multiplexer 823.
[0202] The displacement vector reconstructor 819 may include a video decoder, an image unpacker, an inverse quantizer, and an inverse linear lifting part. That is, in the displacement vector reconstructor 819, the encoded displacement vector is decoded by the video decoder, the image unpacker performs image unpacking, the inverse quantizer performs inverse quantization, and the inverse linear lifting unit performs inverse transformation to reconstruct the displacement vector. The reconstructed displacement vector is output to the mesh reconstructor 820. The mesh reconstructor 820 reconstructs the deformed mesh based on the base mesh reconstructed by the base mesh reconstructor 815 and the displacement vector reconstructed by the displacement vector reconstructor 819. The reconstructed mesh (also referred to as the reconstructed deformed mesh) has reconstructed vertices, inter-vertex connection information, texture coordinates, and inter-texture coordinate connection information.
[0203] The texture map video generator 821 may regenerate the texture map based on the texture map (or attribute map) of the original mesh and the reconstructed deformed mesh output from the mesh reconstructor 820. According to an embodiment, the texture map video generator 821 may assign vertex-by-vertex color information in the texture map of the original mesh to the texture coordinates of the reconstructed deformed mesh. According to an embodiment, the texture map video generator 821 may generate a texture map video by grouping the texture maps regenerated at the frame level into GoFs.
[0204] The generated texture map video may be encoded by the texture map video encoder 822 using a video compression codec. The texture map video bitstream generated by the encoding is sent to the multiplexer 823.
[0205] The multiplexer 823 multiplexes the motion vector bitstream (for example, in the case of inter-frame coding), the base grid bitstream (for example, in the case of intra-frame coding), the displacement vector bitstream, and the texture map bitstream into a single bitstream. This single bitstream can be sent to the receiving side through the transmitter 824. Alternatively, for the motion vector bitstream, the base grid bitstream, the displacement vector bitstream, and the texture map bitstream, a file with one or more track data may be generated, or the bitstream may be encapsulated into fragments and sent to the receiving side through the transmitter 824.
[0206] Reference Fig.13, the transmitter (encoder) can encode the mesh in an intra-frame or inter-frame manner. According to intra-frame coding, the transmitting device can generate a base mesh, a displacement vector (or displacement) and a texture map (or attribute map). According to inter-frame coding, the transmitting device can generate a motion vector (or motion), a displacement vector (or displacement) and a texture map (or attribute map). The texture map obtained from the data input unit is generated and encoded based on the reconstructed mesh. The displacement is generated and encoded based on the vertex position difference between the base mesh and the segmented (or subdivided) mesh. More specifically, the displacement is the position difference between the fitted subdivided mesh and the subdivided reconstructed base mesh, that is, the vertex position difference between the two meshes. The base mesh is generated by simplifying the original mesh through preprocessing and encoding the simplified mesh. For motion, a motion vector is generated for the mesh of the current frame based on the reference base mesh in the previous frame.
[0207] Fig.14 A mesh data receiving device according to an embodiment is shown.
[0208] Fig.14 Corresponds to Figure 1 The receiving device 110 or the grid video decoder 113, Fig.11 or Fig.12 decoder and / or corresponding receiving decoding device. Fig.14 The various components correspond to hardware, software, processors and / or a combination thereof. Fig.14 The receiving (decoding) operation can be followed Fig.13 The reverse process of the corresponding process of the sending (encoding) operation.
[0209] The bit stream of the mesh data received by the receiver 910 is subjected to file / segment decapsulation and then demultiplexed into a compressed motion vector bit stream (e.g., inter-frame decoding) or a base mesh bit stream (e.g., intra-frame decoding), a displacement vector bit stream, and a texture map bit stream by the demultiplexer 911. For example, when the current mesh is inter-frame encoded, the motion vector bit stream is received, demultiplexed, and then output to the motion vector decoder 913 through the switch unit 912. In another example, when the current mesh is intra-frame encoded, the base mesh bit stream is received, demultiplexed, and then output to the static mesh decoder 914 through the switch unit 912. Here, the motion vector decoder 913 may be referred to as a motion decoder.
[0210] According to an embodiment, in the case where inter-frame coding is applied to the current grid based on the frame header information, the motion vector decoder 913 may decode the motion vector bitstream. According to an embodiment, the motion vector decoder 913 may use the previously decoded motion vector as a predictor and add it to the residual motion vector decoded from the bitstream to reconstruct the final motion vector.
[0211] According to an embodiment, when intra-coding is applied to the current mesh based on frame header information, the static mesh decoder 914 may decode the base mesh bitstream to reconstruct connection information, vertex geometry information, texture coordinates, normal information, etc. related to the base mesh.
[0212] According to an embodiment, the base grid reconstructor 915 may reconstruct the current base grid based on the decoded motion vector or the decoded base grid. For example, in the case where inter-frame coding is applied to the current grid, the base grid reconstructor 915 may add the decoded motion vector to the reference base grid and perform inverse quantization to generate a reconstructed base grid. In another example, in the case where intra-frame coding is applied to the current grid, the base grid reconstructor 915 may perform inverse quantization on the base grid decoded by the static grid decoder 914 to generate a reconstructed base grid.
[0213] According to an embodiment, the displacement vector video decoder 917 may decode the displacement vector bitstream into a video bitstream using a video codec.
[0214] According to an embodiment, the displacement vector reconstructor 918 extracts displacement vector transform coefficients from the decoded displacement vector video, and applies inverse quantization and inverse transform to the extracted displacement vector transform coefficients to reconstruct the displacement vector. To this end, the displacement vector reconstructor 918 may include an image unpacker, an inverse quantizer, and an inverse linear lifting section. If the reconstructed displacement vector is a value in a local coordinate system, an inverse transform to a Cartesian coordinate system may be performed.
[0215] The mesh reconstructor 916 may subdivide the reconstructed base mesh to generate additional vertices. By subdividing, vertex connection information including additional vertices, texture coordinates, and connection information about the texture coordinates may be generated. In this case, the mesh reconstructor 916 may combine the subdivided reconstructed base mesh with the reconstructed displacement vectors to generate a final reconstructed mesh (also referred to as a reconstructed deformed mesh).
[0216] According to an embodiment, the texture map video decoder 919 may decode the texture map bitstream into a video bitstream using a video codec to reconstruct the texture map. The reconstructed texture map has color information about each vertex in the reconstructed mesh, and the texture coordinates of each vertex can be used to obtain the color value of the vertex from the texture map.
[0217] According to an embodiment, the mesh reconstructed from the mesh reconstructor 916 and the texture map reconstructed from the texture map video decoder 919 are presented to the user through a rendering process in the mesh data renderer 920 .
[0218] Reference Fig.14, the receiving device (decoder) may decode the mesh in an intra-frame or inter-frame manner. According to intra-frame decoding, the receiving device may receive the base mesh, the displacement vector (or displacement), and the texture map (or attribute map), and render the mesh data based on the reconstructed mesh and the reconstructed texture map. According to inter-frame decoding, the receiving device may receive the motion vector (or motion), the displacement vector (or displacement), the texture map (or attribute map), and render the mesh data based on the reconstructed mesh and the reconstructed texture map.
[0219] The mesh data sending device and method according to the embodiment may preprocess the mesh data, encode the preprocessed mesh data, and send a bit stream containing the encoded mesh data. The point mesh data receiving device and method according to the embodiment may receive a bit stream containing mesh data and decode the mesh data. The mesh data sending / receiving method / device according to the embodiment may be referred to as the method / device according to the embodiment. The mesh data sending / receiving method / device may also be referred to as a 3D data sending / receiving method / device or a point cloud data sending / receiving method / device.
[0220] As described above, in the V-Mesh method, the displacement information generated during the encoding process is converted into a video that is compressed using an existing 2D video codec. In addition, the texture map (equivalent to the attribute map) of the input mesh data is also processed into a video that is compressed using an existing 2D video codec. In this case, the displacement video and the texture map video are generated and encoded based on a user-defined resolution and reconstructed at the same resolution on the receiving side. However, in the case where some video bitstream data is lost on the receiving side due to poor network conditions, it may not be possible to fully reconstruct the entire video data with the current V-Mesh method. In addition, if the available resource level on the receiving side is insufficient to reconstruct and utilize the sent mesh data, the received mesh data may not be used at all.
[0221] To solve these problems, the present disclosure proposes an apparatus and method for sending a bitstream with scalable coding of mesh data to allow the data to be received at a manageable level and reconstructed to the maximum possible level at the receiving side. In particular, in order to enable the V-Mesh method to perform scalable coding / decoding, the present disclosure proposes a scalable coding / decoding apparatus and method for displacement video and texture map video.
[0222] Currently, displacement information generated by the V-Mesh method is processed and encoded as video, but the compression of the video codec may be inefficient because the displacement information actually added to the video data has almost no temporal / spatial redundancy. Therefore, by applying a more efficient encoding method that reflects the characteristics of the displacement information, the compression performance of V-Mesh can be further improved. To this end, the present disclosure proposes a method for encoding displacement information in a manner other than a video compression method, and proposes a scalable encoding / decoding device and method based thereon.
[0223] In particular, the present disclosure proposes a scalable mesh encoding / decoding apparatus and method that is scalable by including one or more sub-bitstream sets with various resolutions and image qualities. To this end, the present disclosure proposes an apparatus and method that can reflect the scalability of displacement video and texture map video compressed using a traditional 2D video codec in a V-Mesh method. The present disclosure also proposes an apparatus and method that can provide not only scalability but also additional improvements in compression efficiency of displacement video.
[0224] Therefore, according to the present disclosure, hierarchical encoding / decoding of mesh data may be allowed, and mesh content with optimal resolution and image quality may be selectively reconstructed and provided based on a network and a receiving environment.
[0225] As described above, the present disclosure relates to a method for hierarchically segmenting and compressing displacement information and texture map information by a V-Mesh encoder of a transmitting device, and a device and method for hierarchically encoding displacement information by applying a coding method in a non-2D video form. The present disclosure also relates to a method for decoding displacement information and texture map information to a target layer by a V-Mesh decoder of a receiving device, and a device and method for decoding mesh data to a target layer using the same.
[0226] In the present disclosure, displacement video may be used interchangeably with displacement vector video or displacement vector transform coefficient video. In addition, displacement may be used interchangeably with displacement information or displacement vector.
[0227] Regarding the transmitting device of the present disclosure, this paper proposes a scalable coding method for V-Mesh, a scalable coding method for displacement video of V-Mesh, and a method for improving the compression performance of displacement information of V-Mesh. In addition, regarding the transmitting device, this paper proposes a scalable coding method for a new displacement information coding method and a scalable coding method for texture map video of V-Mesh.
[0228] Regarding the receiving device of the present disclosure, this paper proposes a scalable decoding method of V-Mesh, a scalable decoding method of displacement video for V-Mesh, and a method for improving the compression performance of V-Mesh displacement information. In addition, regarding the receiving device, this paper proposes a scalable encoding method for a new displacement information encoding method and a scalable encoding method for texture map video for V-Mesh.
[0229] In the present disclosure, as one of the elements constituting a mesh, geometric information (or geometry or geometric data) includes vertices (or points), edges, and polygons. Here, vertices define positions in 3D space, edges represent connections between vertices, and polygons formed by the combination of edges and vertices define the surface of the mesh. In other words, each vertex constituting the mesh represents a position in 3D space, for example, represented by X, Y, and Z coordinates. The polygon can be a triangle or a rectangle. In other words, the geometry forms the skeleton of the 3D model, defining the shape of the model that is visually represented when rendered.
[0230] Fig.15 A mesh data transmitting device according to an embodiment is shown. Fig.15 The sending device may be referred to as a scalable V-Mesh encoder.
[0231] Fig.15 Corresponds to Figure 1 The transmitting device 100 or the grid video encoder 102, Figure 2 , Figure 6 or Figure 7 Encoders (preprocessors and encoders), Fig.13 sending device and / or corresponding sending encoding device. Fig.15 The various components correspond to hardware, software, processors and / or their combination. Fig.15 In the example, you can change the execution order of blocks, omit some blocks, and add new blocks.
[0232] exist Fig.15 In a transmitting device, the V-Mesh encoder can encode the mesh using a progressive encoding method that simplifies the original mesh to create a base mesh and then gradually builds a complex mesh from the base mesh.
[0233] That is, Fig.15 The illustrated operation process of using V-Mesh compression technology on the sending side to compress and send dynamic mesh data. Fig.15 The sending device may support both the intra-frame coding (or intra-frame coding or intra-picture coding) process and / or the inter-frame coding (or inter-frame coding or inter-picture coding) process.
[0234] exist Fig.15In the embodiment of the present invention, a mesh simplification unit (or part) 11011 simplifies an input original mesh to generate a simplified base mesh. Mesh simplification may be performed based on the number of target vertices or the number of target polygons constituting the mesh. For example, a method such as simplification may be used to simplify the original mesh. Specifically, simplification may be a process of selecting vertices to be removed from the original mesh based on a specific reference point, and then removing the selected vertices and triangles connected to the selected vertices.
[0235] In other words, the mesh simplification unit 11011 may simplify the input mesh to a target number of vertices or a target number of faces. The simplification process may be performed using various methods such as triangle collapse and edge collapse.
[0236] According to an embodiment, the base mesh simplified by the mesh simplification unit 11011 is provided to the mesh parameterization unit 11012 and the mesh subdivider 11017 .
[0237] The mesh parameterization unit 11012 performs a process of mapping the 3D surface to the texture domain of the simplified mesh. In one embodiment, the mesh parameterization unit 11012 can use the UVAtlas tool to perform parameterization. In this process, mapping information indicating where each vertex of the simplified mesh can be mapped to on the 2D image is generated. The mapping information is represented as texture coordinates and stored, and the final base mesh is generated in this process. In other words, the mesh parameterization unit 11012 performs parameterization to generate texture coordinates (UV coordinates) and texture connection information for each vertex of the input mesh (i.e., the simplified mesh).
[0238] The final base mesh generated by the parameterization unit 11012 is input to the mesh quantizer 11013 to be quantized. According to an embodiment, the mesh quantizer 11013 may quantize mesh information (e.g., geometric information (x, y, z) and / or texture coordinates (u, v), normal information (nx, ny, nz), etc.) in floating point form into fixed point form. In some embodiments, quantization of specific components may be skipped.
[0239] The mesh subdivider 11017 subdivides the base mesh obtained by simplification of the mesh simplification unit 11011. The mesh subdivider 11017 may perform mesh subdivision on the base mesh to generate additional vertices. According to the subdivision method, vertex connection information, texture coordinates, and connection information about texture coordinates including the additional vertices may be generated. In this case, the geometry information connection information, texture coordinate connection information, and texture coordinates may be implicitly inferred and output according to the subdivision method. According to an embodiment, the mesh subdivider 11017 may perform subdivision using a method such as edge midpoint, Loop, or Catmul & Clark.
[0240] In addition, the mesh subdivision performed by the mesh subdivider 11017 may be performed n times by a user parameter or an agreement between an encoder (i.e., a transmitting device) / a decoder (i.e., a receiving device). According to an embodiment, when the vertices of the base mesh are defined as R0, the vertices newly generated by performing subdivision once are defined as R1, ..., the vertices generated by performing subdivision n times are defined as R n When the level of detail (LoD n ) can be defined as shown in Formula 1 below.
[0241] [Formula 1]
[0242] LoD n =R0∪R1∪,…,∪R n
[0243] Reference Figure 8 , for mesh subdivision, a base mesh (original) may include vertices and edges of LoD0. A first subdivided mesh generated by splitting (or subdividing) the base mesh includes vertices generated by further splitting (or subdividing) the edges of the base mesh. The first subdivided mesh includes vertices of LoD0 and vertices of LoD1. LoD1 includes subdivided vertices and vertices from the base mesh (LoD0). A second subdivided mesh may be generated by splitting (or subdividing) the first subdivided mesh. The second subdivided mesh includes LoD2. LoD2 includes base mesh vertices (LoD0), and LoD1 includes vertices further split (or subdivided) from LoD0 and vertices further split (or subdivided) from LoD1. In other words, LoD is a level of detail representing the degree of detail of the mesh data content. As the index of the level increases, the distance between vertices decreases, and the level of detail is generated. In other words, a smaller LoD value indicates that the mesh data content has less detail, while a larger LoD value indicates that the mesh data content has more detail. LoD N includes the same vertices included in the previous LoD (LoD N-1). In the present disclosure, vertices and points may be used interchangeably. LoD may be defined during subdivision of the base mesh.
[0244] According to an embodiment, the mesh fitting unit 11018 may perform fitting by adjusting the vertex positions so that the subdivided mesh from the mesh subdivider 11017 becomes similar to the original mesh, thereby generating a fitted subdivided mesh.
[0245] In the present disclosure, the combination of the mesh simplification unit 11011, the parameterization unit 11012, the mesh subdivider 11017, and the mesh fitting unit 11018 may be referred to as a preprocessor. According to an embodiment, the preprocessor may further include a displacement vector calculator 11019.
[0246] According to an embodiment, the base grid from the grid quantizer 11013 may be output to the motion vector encoder 11015 or the static grid encoder 11016 via the switching unit 11014 .
[0247] According to an embodiment, when inter-frame encoding is performed on a mesh frame, the base mesh is output to a motion vector encoder 11015 through a switching unit 11014. When intra-frame encoding is performed on a mesh frame, the base mesh is output to a static mesh encoder 11016 through a switching unit 11014. The motion vector encoder 11015 may be referred to as a motion encoder.
[0248] For example, when intra-coding is performed on the mesh frame, the base mesh may be compressed by the static mesh encoder 11016. In this case, connection information, vertex geometry information, vertex texture information, normal information, etc. related to the base mesh may be encoded. The base mesh bitstream generated by the encoding is sent to a multiplexer (not shown).
[0249] As another example, when inter-frame coding is performed on a grid frame, the motion vector encoder 11015 may receive a base grid and a reference reconstructed base grid (or a reconstructed quantized reference base grid) as input, calculate a motion vector between the two grids, and encode its value. In addition, the motion vector encoder 11015 may use a previously encoded / decoded motion vector as a predictor to perform prediction based on connection information, and perform entropy coding on a differential motion vector (also called a residual motion vector) obtained by subtracting the predicted motion vector from the current motion vector. The motion vector bitstream generated by the encoding is sent to a multiplexer (not shown) as a base grid bitstream. That is, in the case of intra-frame coding, a static grid bitstream is input to the multiplexer as a base grid bitstream. In the case of inter-frame coding, the motion vector bitstream is input to the multiplexer as a base grid bitstream.
[0250] According to an embodiment, the encoding may be performed by the motion vector encoder 11015 vertex by vertex or subframe by subframe.
[0251] exist Fig.15In the embodiment of the present invention, the base grid reconstructor 11020 may receive a base grid encoded by a static grid encoder 11016 or a motion vector encoded by a motion vector encoder 11015, and generate a reconstructed base grid. The base grid reconstructor 11020 reconstructs the base grid based on the encoding type (inter-frame encoding or intra-frame encoding) of the current grid. For example, the base grid reconstructor 11020 may reconstruct the base grid by performing static grid decoding on the base grid encoded by the static grid encoder 11016. In this case, quantization may be applied before static grid decoding, and inverse quantization may be applied after static grid decoding. In other words, when intra-frame encoding is performed, inverse quantization may be performed on the base grid quantized by the grid quantizer 11013 to reconstruct the current base grid. As another example, the base grid reconstructor 11020 may reconstruct the base grid based on the reconstructed quantized reference base grid and the motion vector encoded by the motion vector encoder 11015. In other words, when inter-frame encoding is performed, the reconstructed motion vector may be added to the reference reconstructed base grid to generate the current base grid. In this case, if the motion vector is not quantized, the motion vector reconstruction process can be skipped and the motion vector calculated by the motion vector encoder 11015 can be used to reconstruct the current base mesh. The reconstructed base mesh is output to the displacement vector calculator 11019 and the level-specific mesh reconstructor 11025.
[0252] According to an embodiment, the displacement vector calculator 11019 may perform mesh subdivision on the reconstructed base mesh. In addition, the displacement vector calculator 11019 may calculate a displacement vector, which is a value of a vertex position difference between the subdivided reconstructed base mesh and the fitted subdivided mesh generated by the mesh fitting unit 11018. In this case, the displacement vector may be calculated as many times as the number of vertices in the subdivided mesh. In other words, the displacement vector corresponding to the number of vertices in the subdivided mesh may be calculated by the displacement vector calculator 11019.
[0253] According to an embodiment, the displacement vector coordinate transformer 11021 may transform the vertex displacement vector calculated in the 3D Cartesian coordinate system (i.e., (x, y, z) space) into a local coordinate system (i.e., normal, tangential, bi-tangential coordinate system) based on the normal vector of each vertex. Here, the normal vector may be calculated for each subdivided vertex based on geometric information and connection information about neighboring vertices.
[0254] According to an embodiment, the displacement vector transformer 11022 may transform the displacement vector in the (x, y, z) or (n, t, bt) coordinate system. According to an embodiment, a lifting transform, a wavelet transform, etc. may be applied to the transformation. In the (n, t, bt) coordinate system, n represents the normal direction, t represents the tangent direction, and bt represents the bi-tangent direction.
[0255] For example, when performing a lifting transformation to predict the vertex R at the kth subdivision levelk When R t The sub - divided vertex displacement vectors of t can be used as predictors of the displacement vectors at the k - th subdivision level (where t < k or t <= k). In some embodiments, the prediction of the displacement vectors can be an average prediction of the n nearest points based on connection information among the vertices at a subdivision level lower than the current vertex or a distance - based weighted prediction. In some embodiments, the prediction can be based on the displacement vectors of the n vertices used to generate the current vertex in the mesh subdivision operation.
[0256] Then, when the lifting transform is performed, the residual signal generated by prediction can be used to update the displacement vectors of the vertices used in the prediction.
[0257] According to an embodiment, the displacement vector quantizer 11023 can quantize the displacement vector values transformed by the displacement vector transformer 11022, i.e., the transform coefficients. The transform coefficients can be quantized for each axis with different quantization parameters, and the quantization parameters or scaling parameters can be derived according to the protocol reached by the encoder / decoder to determine the quantization rate for each LoD level. According to an embodiment, the displacement vector reconstructor 11024 can perform displacement vector reconstruction by de - quantizing the quantized displacement vectors.
[0258] According to an embodiment, the level - specific mesh reconstructor 11025 can reconstruct the deformed mesh based on the displacement vectors reconstructed by the displacement vector reconstructor 11024 and the base mesh reconstructed by the base mesh reconstructor 11020. More specifically, the level - specific mesh reconstructor 11025 can subdivide the reconstructed base mesh generated by the base mesh reconstructor 11020 and add the reconstructed displacement vectors from the displacement vector reconstructor 11024 to generate a LoD (or sLoD) level - specific reconstructed mesh. According to an embodiment, the LoD (or sLoD) level - specific reconstructed mesh generated by the level - specific mesh reconstructor 11025 is provided to the texture map generator 11027.
[0259] That is, the level - specific mesh reconstructor 11025 can generate a LoD0 reconstructed mesh by adding the corresponding displacement vectors to the vertices of the reconstructed base mesh, generate a LoD1 reconstructed mesh by subdividing the reconstructed base mesh once and adding the displacement vectors to the subdivided mesh, and generate a LoD n reconstructed mesh by subdividing the reconstructed base mesh n times and adding the displacement vectors to the subdivided mesh. In this case, the level - specific mesh reconstructor 11025 can generate a reconstructed mesh for each LoD or can generate a reconstructed mesh for each scalable detail level (sLoD) that is a user - defined scalable level. The sLoD will be described in more detail below under "The First Embodiment of the Displacement Vector Encoding Method".
[0260] According to an embodiment, the displacement vector encoder 11026 may encode the displacement vector transform coefficient quantized by the displacement vector quantizer 11023 .
[0261] In the present disclosure, the displacement vector encoder 11026 may encode the displacement vector transform coefficients by performing one of the various methods described below for scalable encoding / decoding.
[0262] According to an embodiment, the displacement vector encoder 11026 may pack the displacement vector transform coefficients into a 2D image, and then may encode the image using a 2D video codec (eg, a video compression codec) or perform zero-run encoding to generate a displacement vector video bitstream.
[0263] According to an embodiment, the displacement vector video bitstream generated by the displacement vector encoder 11026 through encoding is sent to a multiplexer (not shown). According to an embodiment, regarding the method of selecting the displacement vector encoder 11026, a displacement vector encoder agreed upon by an encoder (i.e., a transmitting side) / decoder (i.e., a receiving side) may be used, or the encoder on the transmitting side may send the type of the displacement vector encoder selected by analyzing the characteristics of the displacement vector to the decoder on the receiving side.
[0264] Next, various encoding methods used by the displacement vector encoder 11026 for the displacement vector are described.
[0265] First Implementation Method of Displacement Vector Coding
[0266] The displacement vector transform coefficients quantized by the displacement vector quantizer 11023 may be packed into a 2D image and encoded by the displacement vector encoder 11026 using a 2D video codec.
[0267] In this case, for scalable coding, a 2D video may be generated for each LoD defined by the mesh subdivider 11017. That is, by grouping the transform coefficients corresponding to each LoD and packing them into each 2D image, a 2D video may be generated for each LoD and encoded using a scalable video codec. In this case, a video codec may be provided for each LoD. As an example, a 2D video may be generated based on the transform coefficients corresponding to LoD2 and encoded using the video codec of LoD2.
[0268] Alternatively, in the present disclosure, sLoDs as scalable levels may be defined, and displacement vector videos may be generated by packing images of respective sLoDs. The displacement vector videos may be encoded using a scalable video codec. In other words, by grouping transform coefficients corresponding to respective sLoDs and packing them into respective 2D images, 2D videos of respective sLoDs may be generated and encoded using a scalable video codec. In this case, a video codec may be provided for each sLoD. In the case of sLoD3, for example, a 2D video may be generated based on the transform coefficients corresponding to sLoD3 and encoded using the video codec of sLoD3.
[0269] In this case, the sLoD for scalable support may be defined separately from the LoD determined by the mesh subdivider 11017, and / or may be configured based on the LoD. In other words, a particular sLoD level may be mapped to the same LoD level, or may be mapped to a different LoD level. For example, sLoD1 may be mapped to LoD1, or may be mapped to LoD3. This means that the number of sLoD levels may or may not be equal to the number of LoD levels. For example, the number of sLoD levels may be 5, and the number of LoD levels may be 5 or 10. If the number of LoD levels is 10 (e.g., LoD0 to LoD9), and sLoD0, sLoD1, sLoD2, and sLoD3 are mapped to LoD0, LoD2, LoD5, and LoD10, respectively, the number of sLoD levels is 4.
[0270] In the present disclosure, sLoD may be defined by the mesh subdivider 11017 or the displacement vector encoder 11026. Alternatively, they may be defined by another separate block or software. In addition, defining sLoD for scalability is intended to increase versatility, and the LoDs may be grouped and simplified by defining sLoDs only for specific (or primary) LoD levels among the LoD levels.
[0271] According to an embodiment, the number of sLoD levels (e.g., num_scalable_LoD_minus1) and the corresponding LoD level values of each sLoD level may be sent, and whether to send sLoD may be determined by a 1-bit flag (e.g., sLoD_enable_flag). For example, sLoD_enable_flag set to 1 may indicate that the displacement vector transform coefficient is based on sLoD encoding, and sLoD_enable_flag set to 0 may indicate that the displacement vector transform coefficient is based on LoD encoding. In another embodiment, the LoD level values corresponding to each sLoD level may be predetermined by an agreement reached by the transmitting side / receiving side. Alternatively, only the sLoD level values may be tabulated and sent so that the LoD level mapped to the sLoD level can be extracted based on the table.
[0272] According to an embodiment, the number of sLoD levels (eg, num_scalable_LoD_minus1) and the LoD level value (eg, LoD_level_idx) corresponding to the sLoD level may be transmitted in signaling information.
[0273] According to an embodiment, the number of sLoD levels (eg, num_scalable_LoD_minus1) and the sLoD level value (eg, sLoD_idx) may be transmitted in signaling information.
[0274] According to an embodiment, the number of sLoD levels (eg, num_scalable_LoD_minus1), LoD level values corresponding to the sLoD levels (eg, LoD_level_idx), and the sLoD level values (eg, sLoD_idx) may be transmitted in signaling information.
[0275] According to an embodiment, when the LoD consists of N levels in the grid subdivider 11017 and (num_scalable_LoD_minus1)=k-1, the receiving device may implicitly derive the k-1th LoD_level_idx as N-1.
[0276] For example, when LoD has N levels and sLoD has k levels, the level value of LoD corresponding to the level value of sLoD may be defined as shown in the following Table 1. When sLoD is defined as in the case of Table 1, a displacement vector video corresponding to sLoD=1 means that a quantized displacement vector transform coefficient corresponding to LoD=3 is packed into a 2D image to generate a video.
[0277] [Table 1]
[0278] sLoD index Corresponding LoD level index 0 0 1 3 … … k-1 N-1
[0279] According to an embodiment, images of transform coefficients organized per LoD or sLoD may be temporally stacked within a group of frames (GoF) to configure a displacement vector video and encoded using a video codec. For example, in the case of sLoD1, transform coefficients corresponding to sLoD1 are grouped together and packed into a 2D image to generate a 2D frame. This means that a GoF consisting of one I frame, at least one B frame, and at least one P frame also consists of a 2D frame corresponding to sLoD1.
[0280] As described above, in the first embodiment of the method in which the displacement vector encoder 11026 encodes the displacement vector, a LoD-specific or sLoD-specific 2D video can be generated by grouping together the transform coefficients corresponding to each LoD or each sLoD and packing them into the corresponding 2D image, and can be encoded using a scalable video codec. In the case of LoD2, for example, a 2D video can be generated based on the transform coefficients corresponding to LoD2 and then encoded using the video codec of LoD2. In the case of sLoD3, for example, a 2D video can be generated based on the transform coefficients corresponding to sLoD3 and then encoded using the video codec of sLoD3.
[0281] Second Implementation Method of Displacement Vector Coding Method
[0282] According to an embodiment, the displacement vector transform coefficients quantized by the displacement vector quantizer 11023 may have low spatial / temporal redundancy, which may make compression using a 2D video codec inefficient. That is, as described above, the displacement vector calculator 11019 may perform mesh subdivision on the reconstructed base mesh and then calculate a displacement vector, which is the value of the vertex position difference between the subdivided reconstructed base mesh and the fitted subdivided mesh generated by the mesh fitting unit 11018. Therefore, the values of the displacement vectors are mostly zero or close to zero. In addition, these displacement vectors are quantized to increase the number of zeros.
[0283] The present disclosure exploits this feature to encode the displacement vector using another encoder (e.g., a zero-run encoder) instead of a 2D video codec. Similarly, zero-run encoding may be performed for a specific LoD level or only for a specific sLoD level. In the case of LoD2, for example, zero-run encoding may be performed on the (quantized) displacement vector transform coefficients corresponding to LoD2. In the case of sLoD3, for example, zero-run encoding may be performed on the (quantized) displacement vector transform coefficients corresponding to sLoD3. In this case, a zero-run encoder corresponding to the number of LoD levels and / or the number of sLoD levels may be required.
[0284] In the present disclosure, an arithmetic encoder may be used as a zero-run encoder for zero-run encoding.
[0285] In other words, in the present disclosure, zero-run coding instead of a 2D video codec may be applied to encode displacement vector transform coefficients to further improve compression efficiency and allow scalable coding.
[0286] According to an embodiment, the displacement vector coordinate transformer 11021 may transform the vertex displacement vector calculated in the 3D Cartesian coordinate system (ie, (x, y, z) space) into the local coordinate system (eg, (n, t, bt) space) based on the normal vector of each vertex.
[0287] According to an embodiment, regarding the displacement vector transformation coefficients (disp0, disp1, disp2) of the (x, y, z) or (n, t, bt) coordinate system, when the displacement vector transformation coefficients of each axis are all zero (ie, (0, 0, 0)), as Fig.16 As shown, the value of zeroRun can be increased and the transform coefficient encoder can be omitted. Fig.16 The axial components corresponding to disp0, disp1, and disp2 in may vary depending on the implementation.
[0288] Fig.16 is a diagram showing an example of zero-run encoding according to an embodiment.
[0289] exist Fig.16 In the example, when the (quantized) displacement vector transform coefficients (disp0, disp1, disp2) are (1, -1, 0), (-2, 0, 0), (0, 0, 0), (0, 0, 0), (0, 0, 0), (0, 0, 0), (-1, 0, 0), and N (0, 0, 0) items, the value of zeroRun of (1, -1, 0) is 0, because (1, -1, 0) is not followed by the displacement vector transform coefficient components of the three axes equal to 0, that is, (0, 0, 0); the value of zeroRun of (-2, 0, 0) is 3, because (-2, 0, 0) is followed by 3 (0, 0, 0) items; the value of zeroRun of (-1, 0, 0) is N, because (-1, 0, 0) is followed by N (0, 0, 0) items.
[0290] exist Fig.16 If one or more components of the transform coefficients (disp0, disp1, disp2) have values different from 0, then Fig.17 Transform coefficient encoding is performed as shown.
[0291] Fig.17 is a flow chart illustrating an example method of zero-run encoding of transform coefficients according to an embodiment.
[0292] exist Fig.17, the transform coefficients disp0, disp1, and disp2 may be the values of x, y, and z, or may be the values of n, t, and bt.
[0293] In other words, when one or more of the components of the three axes have values different from 0, Fig.17 Transform coefficient encoding is performed as shown.
[0294] In addition, Fig.17 In the present disclosure, three flags, namely, isK flag, isZero flag and isOne flag, can be used for zero-run coding. Zero-run coding is a method of compressing data by reducing the repetition of consecutive bit values. In the present disclosure, isN (isZero, isOne, ..., isK) is a flag used to determine whether the size (abs(disp)) of the transform coefficient is equal to N.
[0295] According to an embodiment, the isK flag (Is-Known flag) indicates what the value of the next bit is. The isK flag is set to 1 or 0. When set to 1, it indicates that the next bit value is a known value. For example, in the case of consecutive zeros, the isK flag is set to 1 and the value of the next bit is set to a known value of 0.
[0296] According to an embodiment, the isZero flag indicates the number of consecutive 0 bits. The isZero flag is set to 0 or a positive integer. When the isK flag is set to 1 and the value of the next bit is 0, the isZero flag is set to a value representing the number of consecutive 0s. This is used to compress sequences of 0s.
[0297] According to an embodiment, the isOne flag indicates the number of consecutive 1 bits. The isOne flag is also set to zero or a positive integer. When the isK flag is set to 1 and the value of the next bit is 1, the isOne flag is set to a value representing the number of consecutive 1s. This is used to compress a sequence of 1s.
[0298] Zero-run encoding can use the above-mentioned isK flag, isZero flag and isOne flag to effectively represent a bit pattern of consecutive 0s and 1s to compress data.
[0299] exist Fig.17 In the example, abs(disp) represents the absolute value of the displacement vector transform coefficient. The displacement vector transform coefficient can be a positive integer or 0, indicating the number of consecutive 0s or 1s indicated by the isZero flag or the isOne flag. In other words, abs(disp) represents the absolute value of the displacement vector transform coefficient (disp), and is therefore always an integer greater than or equal to 0.
[0300] exist Fig.17In the above, the transform coefficients can be encoded sequentially for each component (disp0, disp1, disp2).
[0301] According to an embodiment, when abs(disp) is 0, the isZero flag may be encoded as 1. When abs(disp) is 1, the isZero flag may be encoded as 0. In addition, when abs(disp) is 1, the isOne flag may be encoded as 1. When abs(disp) is 1, the isOne flag may be encoded as 0.
[0302] According to an embodiment, the isN flag is used to check whether the currently encoded coefficient is equal to N (N=0, 1, ..., K, where K is a positive integer). When the encoded coefficient is equal to N, the isN flag may be encoded and the operation may be terminated.
[0303] exist Fig.17 In , when all values up to the isK flag are 0, abs(disp)-(K+1) can finally be entropy coded.
[0304] When both disp0 and disp1 are 0, disp2, which is encoded last in sequence, can always be derived to be greater than or equal to 0. Therefore, disp2-1 can be encoded.
[0305] According to an embodiment, when the coordinate system of the displacement vector is the (n, t, bt) coordinate system, zero-run encoding may be performed on the normal component (n) and the tangential component (t, bt) respectively. In this case, the transform coefficients of the normal component and the tangential component may be encoded in parallel.
[0306] Fig.18 is a diagram showing an example of zero-run encoding in a (n, t, bt) coordinate system according to an embodiment.
[0307] According to an embodiment, when the coordinate system of the displacement vector is a normal coordinate system, zero-run encoding may be performed separately on the normal component (n) and the tangential component (t, bt) according to an agreement between the encoder (ie, the transmitting device) / decoder (ie, the receiving device), or a 1-bit flag (ie, separate_zero_run_flag) may be used to determine whether to perform zero-run encoding separately on the normal component (n) and the tangential component (t, bt). For example, separate_zero_run_flag set to 1 may indicate that the normal component and the tangential component are each subjected to zero-run encoding, while separate_zero_run_flag set to 0 may indicate that the normal component and the tangential component are simultaneously subjected to zero-run encoding. In this case, when separate_zero_run_flag is set to 1, the receiving device or the decoder may perform zero-run decoding on the normal component and the tangential component, respectively. When separate_zero_run_flag is set to 1, the receiving device or decoder may perform zero-run decoding on the normal component and the tangential component at the same time.
[0308] exist Fig.18 In , when the displacement vector transformation coefficient (n,t,bt) is (1,-1,0), (-2,0,0), (0,0,0), (0,0,0), (0,0,0), (0,0,0), (-1,0,0), and N (0,0,0) items, the normal is composed of (1), (-2), (0), (0), (0), (0), (0), (-1), and N (0) items. In this case, the value of zeroRun of (1) is 0 because there is no 0 after (1); the value of zeroRun of (-2) is 3 because there are three 0s after (-2); the value of zeroRunN of (-1) is N because there are N 0s after (-1).
[0309] In addition, when the displacement vector transform coefficients (n,t,bt) are (1,-1,0), (-2,0,0), (0,0,0), (0,0,0), (0,0,0), (0,0,0), (-1,0,0), and N (0,0,0) entries, the tangent and bitangent are composed of (-1,0), (,0,0), (0,0), (0,0), (0,0), (0,0), (0,0), (0,0), N (0,0) entries. In this case, (-1,0) is followed by 5+N (0,0) entries, so the value of zeroRunT is N+5.
[0310] Even in this case, zero-run encoding may be performed on a per-LoD or per-sLoD basis. Alternatively, zero-run encoding may be performed on the normal component and the tangential component individually up to a specific LoD level or a specific sLoD level, and then the normal component and the tangential component may be simultaneously zero-run encoded for the remaining LoD levels (or sLoD levels). For example, zero-run encoding may be performed on the normal component and the tangential component individually for LoD (or sLoD) 0 to k. Starting from LoD (or sLoD) k+1, zero-run encoding may be performed on the normal component and the tangential component simultaneously. Here, k may be determined by an agreement reached by the encoder / decoder, or may be determined by signaling.
[0311] Fig.19 is a flowchart illustrating an example of zero-run encoding of displacement vector transform coefficients of normal components in a (n, t, bt) coordinate system according to an embodiment.
[0312] Fig. 20 is a flowchart illustrating an example of zero-run encoding of displacement vector transform coefficients of tangential components in a (n, t, bt) coordinate system according to an embodiment.
[0313] Specifically, when zero-run encoding is performed on the normal component (n) and the tangential component (t, bt) separately, the displacement vector transformation coefficient disp of the normal component is n Can be as Fig.19 The figure is subjected to zero-run encoding, and the displacement vector transformation coefficient disp of the tangential component t and disp bt Can be as Fig. 20 The shown is subjected to zero-run encoding.
[0314] In addition Fig.19 and Fig. 20 In the example, the isK flag indicates what the value of the next bit is. The isK flag is set to 1 or 0. When set to 1, it indicates that the value of the next bit is a known value. For example, for consecutive 0s, the isK flag is set to 1, and the value of the next bit is set to a known value of 0.
[0315] The isZero flag indicates the number of consecutive 0 bits. The isZero flag is set to 0 or a positive integer. When the isK flag is set to 1 and the value of the next bit is 0, the isZero flag is set to a value indicating the number of consecutive 0s. This is used to compress sequences of 0s.
[0316] In addition, the isOne flag indicates the number of consecutive 1 bits. The isOne flag is also set to zero or a positive integer. When the isK flag is set to 1 and the value of the next bit is 1, the isOne flag is set to a value indicating the number of consecutive 1s. This is used to compress sequences of 1s.
[0317] exist Fig.19 and Fig. 20 In the abs(disp n ) indicates the absolute value of the displacement vector transformation coefficient of the normal component (n), abs(disp t ) indicates the absolute value of the displacement vector transformation coefficient of the tangential component (t), abs(disp bt ) indicates the absolute value of the displacement vector transform coefficient of the bi-tangential component (bt). In other words, the displacement vector transform coefficient of each component can be a positive integer or 0, indicating the number of consecutive 0s or 1s indicated by the isZero flag or the isOne flag. In other words, abs(disp n ) represents the displacement vector transformation coefficient (disp n ), and is therefore always an integer greater than or equal to 0.
[0318] According to an embodiment, when abs(disp n ) is 0, the isZero flag can be encoded as 1. When abs(disp n ) is 1, the isZero flag can be encoded as 0. In addition, when abs(disp n ) is 1, the isOne flag can be encoded as 1. n ) is 1, the isOne flag can be encoded as 0. The same rule applies to abs(disp t ) and abs(disp bt ).
[0319] According to an embodiment, the isN flag is used to check whether the currently encoded coefficient is equal to N (N=0, 1, ..., K, where K is a positive integer). When the encoded coefficient is equal to N, the isN flag may be encoded and the operation may be terminated.
[0320] exist Fig.19 and Fig. 20 In , when all values up to the isK flag are 0, abs(disp)-(K+1) can be finally entropy coded.
[0321] When disp t When it is 0, the last coded disp in order bt can always be deduced to be greater than or equal to 0. Therefore, disp bt -1 for encoding.
[0322] According to an embodiment, even when zero-run encoding is performed, the LoD determined by the grid subdivider 11017 or a separate level sLoD for scalable support may be defined, and zero-run encoding may be performed on the displacement vector transform coefficients per level (LoD level or sLoD level). This may allow scalable encoding and decoding of displacement vectors.
[0323] As described above, in the second embodiment of the method in which the displacement vector encoder 11026 encodes the displacement vector, a scalable zero-run encoder may be used in the zero-run encoding of the displacement vector transform coefficient of the (x, y, z) component or the (n, t, bt) component to perform zero-run encoding only on the displacement vector transform coefficient of a specific LoD level or a specific sLoD level (x, y, z) component or the (n, t, bt) component. In the case of LoD2, for example, zero-run encoding may be performed on the displacement vector transform coefficient of the (x, y, z) component or the (n, t, bt) component corresponding to LoD2. In the case of sLoD3, for example, zero-run encoding may be performed on the displacement vector transform coefficient of the (x, y, z) component or the (n, t, bt) component corresponding to sLoD3.
[0324] Third Implementation Method of Displacement Vector Coding Method
[0325] According to an embodiment, zero-run encoding may be performed by bundling displacement vector transform coefficients of a plurality of frames within a group of frames (GOF). In other words, two or more frames in a GOF may be grouped, and zero-run encoding may be performed based on each group of bundled displacement vector transform coefficients. For example, every two frames in a GoF may be grouped together, and then zero-run encoding may be performed based on each group of bundled displacement vector transform coefficients.
[0326] In the present disclosure, the packing method may vary according to the characteristics of the grouped frames.
[0327] Fig.21 An example of packing displacement vector transform coefficients of two frames in an interleaved manner according to an embodiment is shown.
[0328] Fig. 22 An example of packing displacement vector transform coefficients of two frames in a serial manner according to an embodiment is shown.
[0329] According to an embodiment, when the vertices in the grouped frames have a 1-to-1 mapping relationship, the displacement vector transformation coefficients with the same vertex index may be as follows: Fig.21 The transform coefficients are reconfigured by interleaving as shown, and zero-run coding can be performed on the reconfigured transform coefficients. For example, assume that the transform coefficients of frame t and the transform coefficients of frame t+1 are packed (or grouped). When the transform coefficients of frame t are C t,0 , C t,1 ,…,Ct,nt-1 And the transform coefficient of frame t+1 is C t+1,0 , C t+1,1 ,…,C t+1,nt-1 When, through interleaving, the input sequence of zero-run coding can be C t,0 , C t+1,0 , C t,1 , C t+1,1 ,…,C t,nt-1 and C t+1,nt-1 .
[0330] According to an embodiment, when the vertices in the grouped frames do not have a 1-to-1 mapping relationship, the transform coefficients may be sorted per LoD level using a method defined by an agreement between encoders / decoders based on the reconstructed vertex geometry information related to each frame, and the sorted transform coefficients may be combined by interleaving. Then, zero-run encoding may be performed. As a sorting method, for example, 3D Morton code may be used. In addition, as Fig. 22 As shown, the displacement vector transform coefficients may be reconfigured by ordering the transform coefficients in frame order rather than in an interleaved manner, and zero-run encoding may be performed on the reconfigured transform coefficients (e.g., in a serial manner). For example, assume that the transform coefficients of frame t and the transform coefficients of frame t+1 are packed (or grouped). When the transform coefficients of frame t are C t,0 , C t,1 ,…,C t,nt-1 And the transform coefficient of frame t+1 is C t+1,0 , C t+1,1 ,…,C t+1,nt-1 When the zero-run encoding input sequence is C t,0 , C t,1 ,…,C t,nt-1 , C t+1,0 , C t+1,1 , … and C t+1,nt-1 .
[0331] According to an embodiment, a method of packing displacement vector transform coefficients of a plurality of frames may be signaled to a decoder through signaling information (e.g., displacement_packing_method). That is, when the displacement vector transform coefficients of a plurality of frames are encoded in bundles, displacement_packing_method may specify a packing method of transform coefficients of each frame. For example, displacement_packing_method set to 0 indicates that the frames are packed in an interleaved manner, while displacement_packing_method set to 1 indicates that the frames are packed in a serial manner.
[0332] According to an embodiment, transform coefficients of a plurality of (or two or more) frames combined in an interleaved or serial manner may be subjected to zero-run encoding per LoD level (or sLoD level).
[0333] In this case, zero-run encoding can also be performed by bundling the displacement vector transform coefficients of the frames based on each LoD or each sLoD. For example, zero-run encoding can be performed by bundling the displacement vector transform coefficients corresponding to LoD1 in two frames in an interleaved manner or in a serial manner. As another example, zero-run encoding can be performed by bundling the displacement vector transform coefficients corresponding to sLoD2 in two frames in an interleaved manner or in a serial manner. Alternatively, zero-run encoding can be performed level by level until a specific LoD level or a specific sLoD level. For the remaining LoD levels (or sLoD levels) until the last LoD level (or sLoD level), zero-run encoding can be performed all at once. For example, for LoD (or sLoD) 0 to k, zero-run encoding can be performed based on each LoD (or sLoD) level. For LoD (or sLoD) level k+1 to the last LoD (or sLoD) level, zero-run encoding can be performed once.
[0334] Fig.23 An example of zero-run encoding of displacement vector transform coefficients for two interleaved frames per LoD level according to an embodiment is shown. Fig.23 middle, Less than
[0335] In other words, Fig.23 It is shown that zero-run coding is performed by interleaving the displacement vector transform coefficients corresponding to LoD0 in the two frames, zero-run coding is performed by interleaving the displacement vector transform coefficients corresponding to LoD1 in the two frames, and L-1 An example of performing zero-run encoding on the corresponding displacement vector transform coefficients.
[0336] In this case, the unit in which zero-run encoding is independently performed may be the LoD level, or may be the sLoD level defined for scalability.
[0337] On the other hand, the level-based reconstructed mesh (or reconstructed deformed mesh) from the level-specific mesh reconstructor 11025 has reconstructed vertices, vertex connection information, texture coordinates, and inter-texture coordinate connection information. That is, the level-specific mesh reconstructor 11025 may subdivide the reconstructed base mesh generated by the base mesh reconstructor 11020 and add the reconstructed displacement vectors from the displacement vector reconstructor 11024 to generate a reconstructed deformed mesh for each LoD (or sLoD) level.
[0338] According to an embodiment, the texture map generator 11027 may regenerate the texture map of the current mesh based on the texture map (also called attribute map) of the original mesh and the reconstructed mesh from the level-specific mesh reconstructor 11025 as input.
[0339] For example, color mapping to u,v coordinates in texture coordinate space may be performed by a texture color mapping process for all u,v coordinates.
[0340] Fig.24 is a diagram showing an example of texture color mapping of a texture map generator according to an embodiment. Fig.24 The various components in the EMBODIMENTS correspond to hardware, software, processors and / or a combination thereof.
[0341] According to an embodiment, the texture color mapping process is performed by a texture color mapper of the texture map generator 11027. According to an embodiment, the texture color mapper may include a determination unit 13011, a polygon geometry calculator 13012, a nearest geometry calculator 13013, and a color information distributor 13014.
[0342] According to an embodiment, for all input u,v coordinates (eg, texture coordinates of the original mesh), the determination unit 13011 determines whether the coordinates exist within a polygon in the texture coordinate space based on the texture connection information.
[0343] If it is determined that the coordinates exist in the polygon, the polygon geometry information calculator 13012 calculates the geometry information (x, y, z) about the polygon for the u, v coordinates existing in the polygon. The representative value of the polygon can be calculated as the geometry information about the polygon. The representative value can be an average value, a centroid, etc.
[0344] According to an embodiment, the nearest geometry information calculator 13013 calculates a polygon or a vertex of the original mesh that is closest to the geometry information about the polygon or the representative value of the geometry information about the polygon calculated by the polygon geometry information calculator 13012 .
[0345] According to an embodiment, the color information assigner 13014 assigns the texture map color of the current u,v position according to the color of the nearest polygon or vertex calculated by the nearest geometry information calculator 13013. For example, once the nearest polygon is calculated, the texture map color at the current u,v coordinate can be assigned based on the average value of the polygon.
[0346] According to an embodiment, when the texture map is regenerated by the texture color mapper as described above, the texture map generator 11027 may regenerate the texture map for each LoD (or sLoD) level based on the texture map (or attribute map) of the original mesh and the deformed base mesh (or reconstructed deformed mesh) reconstructed for each LoD (or sLoD) level by the level-specific mesh reconstructor 11025. According to an embodiment, the texture map generator 11027 may assign per-vertex color information of the texture map from the original mesh to the texture coordinates of the reconstructed base mesh (or reconstructed deformed mesh) for each LoD (or sLoD) level. According to an embodiment, the texture map generator 11027 may generate a texture map video by bundling the reconstructed texture map per GoF in each frame.
[0347] In the present disclosure, the texture map generator 11027 may generate a texture map of a grid for each LoD (or sLoD), and then the texture map encoder 11028 may perform scalable encoding on it per LoD (or sLoD) using a scalable 2D video codec.
[0348] According to an embodiment, a base mesh is defined as LoD0 and a mesh obtained by applying n-time mesh subdivision to the base mesh is defined as LoD n When the grid is Fig.24 As shown, the texture map generator 11027 can generate a LoD-specific texture map by performing a texture color mapping process on the mesh of each LoD. To this end, the subdivision-related information and the LoD-related information and / or the sLoD-related information defined in the mesh subdivider 11017 can be provided to the texture map encoder 11028. In addition, the LoD-specific information and / or the sLoD-specific information applied by the displacement vector encoder 11026 for displacement vector scalable encoding can be provided to the texture map encoder 11028.
[0349] According to an embodiment, the texture map generated for each LoD may be downsampled level by level and then may be encoded by the texture map encoder 11028 .
[0350] For example, when changing from LoD0 to LoD L-1 When L LoDs are arranged, the texture map of the nth LoD can be downsampled to size.
[0351] According to an embodiment, the sLoD level for scalable support may be defined separately from the LoD level determined by the mesh subdivider 11017, and texture map generation and downsampling may be performed on a per-sLoD level basis.
[0352] According to an embodiment, when determining the grid level based on each sLoD level, the grid level can be determined by downsampling to For example, the texture map of the nth sLoD level among the k sLoD levels can be downsampled to The size of is then encoded by the texture map encoder 11028.
[0353] That is, when downsampling is performed for each level, a downsampling method may be determined based on the existence and / or number of triangles in a T×T area currently being downsampled based on texture coordinates and texture connection information. Here, the level may be a LoD level or an sLoD level.
[0354] According to an embodiment, when there are multiple polygons (eg, triangles) in a T×T region of the nth level texture map, the downsampled color of the (n-1)th level texture map may be calculated based on an average or weighted sum of surface colors of the triangles.
[0355] exist Fig.25 , LoD2, LoD1, and LoD0 input to the scalable encoder are examples of texture maps downsampled by the texture map generator 1127 according to the LoD levels.
[0356] According to an embodiment, when there is one polygon in the T×T area of the nth level texture map, the down-sampled color of the (n-1)th level texture map may be calculated based on the surface color of the polygon.
[0357] According to an embodiment, when there is no corresponding coordinate in the polygon based on the texture coordinate connection information of all u,v coordinates, the texture map generated for each level may be filled based on the color value of the surrounding texture map. According to an embodiment, for example, push-pull filling may be used.
[0358] According to an embodiment, even when one texture map is generated instead of texture maps of each level, the texture map may be downsampled, wherein the downsampling size may be configured as an index and signaled to the receiving side.
[0359] According to an embodiment, the 2D scalable video codec in the texture map encoder 11028 may encode a texture map that has been downsampled on a per-level basis.
[0360] Fig.25 An example of compressing a texture map generated per level using a scalable video codec according to an embodiment is shown.
[0361] For example, a texture map of an nth LoD may be encoded (eg, inter-layer prediction) with reference to a texture map of a LoD lower than a current LoD.
[0362] Alternatively, the texture map generated for each level may be encoded using an independent video codec therefor. For example, when the number of LoD levels is 4, 4 independent video codecs may be used to encode the texture map of the corresponding LoD. As another example, when the number of sLoD levels is 2, 2 independent video codecs may be used to encode the texture map of the corresponding sLoD.
[0363] Therefore, the texture map video per LoD (or sLoD) level generated by the texture map generator 11027 may be encoded using a scalable 2D video codec in the texture map encoder 11028. The per LoD (or sLoD) level texture map substream (or texture map video bitstream) generated by the encoding is sent to a multiplexer (not shown).
[0364] According to an embodiment, the texture map video encoder may be a video encoder (e.g., VVC, HEVC, etc.), an encoder based on entropy coding, etc. As a method of selecting a texture map video encoder, a texture map encoder agreed upon by an encoder (i.e., a transmitting side) / decoder (i.e., a receiving side) may be used, or the type of the texture map encoder selected by the encoder on the transmitting side may be sent to the decoder on the receiving side.
[0365] In one embodiment, the LoD (or sLoD) level applied to the encoding of the displacement vector is the same as the LoD (or sLoD) level applied to the encoding of the texture map. For example, when the displacement vector corresponding to LoD3 is encoded using a 2D video codec, the texture map corresponding to LoD3 is encoded using a 2D video codec. As another example, when the displacement vector corresponding to sLoD1 is encoded using a 2D video codec method, the texture map corresponding to sLoD1 can be encoded using a 2D video codec method.
[0366] According to an embodiment, the multiplexer may multiplex the input base grid bitstream, displacement vector bitstream, and texture map bitstream into a single bitstream and send the single bitstream to the receiver via a transmitter (not shown). Alternatively, the base grid bitstream, displacement vector bitstream, and texture map bitstream may be encapsulated into a file / segment and sent to the receiver via a transmitter.
[0367] According to an embodiment, the bit stream multiplexed by the multiplexer can be sent via a network and / or stored on a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0368] Fig.26 is a diagram showing a mesh data receiving device according to an embodiment. In the present disclosure, Fig.26The receiving device may be referred to as a decoder.
[0369] Fig.26 Corresponds to Figure 1 The receiving device 110 or the grid video decoder 113, Fig.11 or Fig.12 The decoder, Fig.14 receiving device and / or corresponding receiving decoding device. Fig.26 The various components correspond to hardware, software, processors and / or a combination thereof. Fig.26 The receiving (decoding) operation can be followed by Fig.15 The sending (encoding) operation is the opposite of the corresponding process. Fig.26 In the example, the execution order of blocks can be changed, some blocks can be omitted, and some new blocks can be added.
[0370] According to an embodiment, the mesh data bitstream received by the receiver (not shown) is file / fragment decapsulated and then demultiplexed by a demultiplexer (not shown) into a base mesh bitstream, a displacement vector bitstream, and a texture map bitstream. In the case where inter-frame coding is applied to the current mesh, the base mesh bitstream may be a motion vector bitstream.
[0371] According to an embodiment, the basic grid bit stream is output to the motion vector decoder 15012 or the static grid decoder 15013 via the switching unit 15011.
[0372] For example, in the case where inter-frame coding is applied to the current grid, a base grid bit stream (i.e., a motion vector bit stream) is received, demultiplexed, and output to the motion vector decoder 15012 via the switching unit 15011. As another example, in the case where intra-frame coding is applied to the current grid, a base grid bit stream is received, demultiplexed, and output to the static grid decoder 15013 via the switching unit 15011. Here, the motion vector decoder 15012 may be referred to as a motion decoder.
[0373] According to an embodiment, the motion vector decoder 15012 may perform decoding on the motion vector bitstream on a vertex-by-vertex or sub-group-by-sub-group basis.
[0374] According to an embodiment, the motion vector decoder 15012 may use a previously decoded motion vector as a predictor and add it to a differential motion vector (i.e., a residual motion vector) decoded from a bitstream to reconstruct a final motion vector. That is, the motion vector decoder 15012 may decode a vertex or sub-block differential motion vector (or a residual motion vector) through a motion vector bitstream, and perform prediction based on connection information using a previously decoded motion vector as a predictor to decode the motion vector by adding the residual motion vector to the prediction.
[0375] According to an embodiment, the static mesh decoder 15013 may decode the base mesh bitstream to reconstruct connection information, vertex geometry information, texture coordinates (ie, attribute geometry information), normal information, etc. related to the base mesh.
[0376] According to an embodiment, the base grid reconstructor 15014 may reconstruct the current base grid based on the decoded motion vector or the decoded base grid. For example, in the case where inter-frame coding is applied to the current grid, the base grid reconstructor 15014 may add the decoded (or reconstructed) motion vector to the reference base grid and perform inverse quantization to generate a reconstructed base grid (i.e., the current base grid). As another example, in the case where intra-frame coding is applied to the current grid, the base grid reconstructor 15014 may perform inverse quantization on the base grid decoded (or reconstructed) by the static grid decoder 15012 to generate a reconstructed base grid (i.e., the current base grid).
[0377] According to an embodiment, the mesh subdivider 15015 may subdivide the base mesh to generate additional vertices. In the present disclosure, the geometry information connection information, the texture coordinate connection information and the texture coordinates may be implicitly inferred and generated according to the subdivision method.
[0378] Depending on the implementation, the mesh subdivider 15015 may perform subdivision using methods such as edge midpoint, Loop, Catmul & Clark, etc.
[0379] According to an embodiment, the mesh subdivider 15015 may perform mesh subdivision n times based on a user parameter or an agreement reached by the encoder / decoder. According to an embodiment, when the vertices of the base mesh are defined as R0, the vertices newly generated by performing subdivision once are defined as R1, ..., and the vertices generated by performing subdivision n times are defined as R n When LoD n It can be defined as shown in Formula 2 below.
[0380] [Formula 2]
[0381] LoD n =R0∪R1∪,…,∪R n
[0382] According to an embodiment, the displacement vector transform coefficient decoder 15017 may decode the demultiplexed displacement vector bitstream into a video bitstream using a video codec, or perform zero-run decoding on it.
[0383] According to an embodiment, the displacement vector transform coefficient decoder 15017 may decode and reconstruct the displacement vector through a reverse process of at least one of the first to third embodiments of the displacement vector encoding method on the transmitting side.
[0384] According to an embodiment, the displacement vector transform coefficient decoder 15017 may decode the displacement vector by decoding the received displacement vector transform coefficient of a specific LoD (or sLoD) level or the displacement vector transform coefficient of a specific LoD (or sLoD) level based on the processing capability of the receiver.
[0385] First Implementation Method of Displacement Vector Decoding
[0386] According to an embodiment, the displacement vector transform coefficient decoder 15017 of the receiving device decodes the displacement vector transform coefficient video encoded using the video codec of the transmitting side using the same video codec.
[0387] According to an embodiment, the displacement vector transform coefficient decoder 15017 may use a 2D scalable decoder to reconstruct a displacement vector transform coefficient video defined by a convention between encoders / decoders based on the LoD level determined by the grid subdivider 15015 or the LoD (i.e., sLoD) level configured per LoD (or sLoD) level defined for scalable coding or configured from received information. That is, 2D video decoding may be performed only for a displacement vector transform coefficient of a specific LoD (or sLoD) level with reference to signaling information (e.g., at least one of num_scalable_LoD_minus1, LoD_level_idx, sLoD_index, or sLoD_enable_flag).
[0388] According to an embodiment, the displacement vector transform coefficient decoder 15017 can decode the displacement vector transform coefficient video of each level from a lower level by parsing the sub-bitstream. The decoding can be performed by referring to an inter-layer reference of a video at a level lower than the current level when performing decoding. Here, the level can be a LoD level or an sLoD level.
[0389] According to an embodiment, the displacement vector transform coefficient decoder 15017 may perform a grid scalable function by decoding the displacement vector transform coefficient video up to a specific level of a level-specific substream.
[0390] According to an embodiment, the displacement vector transform coefficient decoder 15017 for 2D scalable video decoding may be a video decoder (eg, VVC, HEVC), an entropy coding-based decoder, or the like.
[0391] Second Implementation Method of Displacement Vector Decoding
[0392] According to an embodiment, the displacement vector transform coefficient decoder 15017 decodes the displacement vector transform coefficients (disp0, disp1 and disp2) in the (x, y, z) or (n, t, bt) coordinate system.
[0393] The displacement vector transform coefficients may be decoded using a zero-run decoder. According to an embodiment, zero-run decoding may be performed on all three transform coefficients (disp0, disp1, and disp2) at once. Alternatively, when the coordinate system of the displacement vector transform coefficients is a normal coordinate system, the normal component (n) and the tangential component (t, bt) may be zero-run decoded by the zero-run decoder, respectively.
[0394] In this case, zero-run decoding may be performed on all three transform coefficients of a specific LoD (or sLoD) level at once with reference to the signaling information. Alternatively, when the coordinate system of the displacement vector transform coefficient is a normal coordinate system, the normal component (n) and the tangential component (t, bt) of the specific LoD (or sLoD) level may each be zero-run decoded.
[0395] Fig. 27 (a) is a flowchart showing an example of zero-run decoding of displacement vector transform coefficients according to an embodiment.
[0396] Fig. 27 (b) is a flowchart showing an example of decoding the absolute value of the displacement vector transformation coefficient according to an embodiment.
[0397] in particular, Fig. 27 (a) and Fig. 27 (b) shows an example of a case where the displacement vector transform coefficient is decoded as a single zero run.
[0398] exist Fig. 27 (a) and Fig. 27 In (b), zero-run decoding is repeated as many times as the number of vertices (numVertex). In particular, the three transform coefficients (disp0, disp1, disp2) are subjected to zero-run decoding component by component.
[0399] According to an embodiment, in Fig. 27 In (a), the zeroRun decoder parses the value of zeroRun from the displacement vector bitstream, and the transform coefficient deriver derives the value of the transform coefficient of the vertex corresponding to the size of zeroRun as 0 according to the parsed value of zeroRun (i.e., when the value of zeroRun is not 0). When the parsed value of zeroRun is 0 or when the value of zeroRun is derived as 0, the transform coefficient decoder performs decoding on the absolute values of the three transform coefficients. In this case, for disp2, which is decoded last in order, decoding of the transform coefficient sign of disp2+1 may be performed.
[0400] When the coordinate system of the displacement vector is the (n, t, bt) coordinate system, zero-run decoding may be performed on the normal component (n) and the tangential component (t, bt), respectively.
[0401] Fig.28 (a) and Fig.28 (b) is a flowchart showing an example of zero-run decoding of a displacement vector transform coefficient of a normal component in a (n, t, bt) coordinate system according to an embodiment.
[0402] Fig.29 (a) and Fig.29 (b) is a flowchart showing an example of zero-run decoding of displacement vector transform coefficients of tangential components in the (n, t, bt) coordinate system according to an embodiment.
[0403] That is, when the normal component and the tangential component are decoded by the zero-run decoder, respectively, the implementation of the transform coefficient decoder is as follows: Fig.28 (a) and Fig.29 In particular, Fig.28 (b) is given by Fig.28 (a) An example of a detailed decoding process of a normal component displacement vector transform coefficient performed by a transform coefficient decoder, Fig.29 (b) is given by Fig.29 (a) is an example of a detailed decoding process of a displacement vector transform coefficient of a tangential component performed by a transform coefficient decoder.
[0404] According to an embodiment, a 1-bit flag (e.g., separate_zero_run_flag) included in the signaling information may determine whether to perform zero-run decoding on the normal component and the tangential component separately. According to an embodiment, separate_zero_run_flag serves as a flag for determining whether to perform separate zero-run decoding on the normal component and the tangential component separately. When separate_zero_run_flag is set to 1, zero-run decoding may be performed on the normal component and the tangential component separately. When separate_zero_run_flag is set to 0, zero-run decoding may be performed on the normal component and the tangential component at the same time.
[0405] According to an embodiment, for LoD (or sLoD) levels 0 to k, zero-run decoding may be performed separately on the normal component and the tangential component. For LoD (or sLoD) levels k+1 and above, the normal component and the tangential component may be decoded simultaneously (i.e., using one zero-run decoder).
[0406] In other words, through Fig.28 and Fig.29 The process in can decode the transform coefficients of the normal component and the tangential component in parallel.
[0407] Fig.30 is a graph showing the absolute value abs(dispn )’s entropy decoding flowchart.
[0408] Reference Fig.30 , when the value of the isZero flag decoded by the isZero decoder is 1, the absolute value abs(disp) of the displacement vector transform coefficient is set to 0. When the value of this flag is 0, the isOne flag is decoded by the isOne decoder. When the value of the isOne flag decoded by the isOne decoder is 1, abs(disp) is set to 1. When this value is 0, isK decoding is repeated until K. When the value of the isK flag is 1, abs(disp) is set to K. Otherwise, disp_minus_Kplus1 decoding is performed, and abs(disp) is set to disp_minus_Kplus1+K+1.
[0409] As described above, when the transmitting device performs zero-run decoding on the displacement vector transform coefficient per LoD or sLoD, the second embodiment of the method for the displacement vector transform coefficient decoder 15017 to decode the displacement vector may perform zero-run decoding per LoD or sLoD level. In this case, the displacement vector transform coefficient decoder 15017 may know whether sLoD is enabled from the sLoD_enable_flag flag included in the signaling information, and may derive the number of vertices corresponding to each level based on the level defined for the LoD or sLoD determined by the mesh subdivider 15015. In addition, when zero-run decoding is performed per LoD or sLoD level, the transform coefficient may be decoded per LoD or sLoD level based on the number of vertices derived per LoD or sLoD level. Then, the mesh scalability function may be performed by reconstructing the displacement vector transform coefficient up to a specific LoD or sLoD level among the LoD or sLoD level specific substreams.
[0410] Third Implementation Method of Displacement Vector Decoding
[0411] As described above, the displacement vector encoder on the sending side can perform zero-run encoding by bundling the displacement vector transform coefficients of multiple frames in the GOF. In particular, if the vertices in the grouped frames have a 1-to-1 mapping relationship, the displacement vector transform coefficients with the same vertex index can be packed in an interleaved manner. Otherwise, the displacement vector transform coefficients can be packed in a serial manner. Then, signaling information (e.g., displacement_packing_method) is used to signal the packing method. In addition, the displacement vector encoder on the sending side can perform zero-run encoding by bundling the displacement vector transform coefficients of the frame on a per-LoD or per-sLoD basis. Alternatively, the displacement vector encoder on the sending side can perform zero-run encoding for each LoD (or sLoD) level from LoD (or sLoD) 0 to k, and perform zero-run encoding once for LoD (or sLoD) level k+1 to the last LoD (or sLoD) level.
[0412] In this case, the displacement vector transform coefficient decoder 15017 on the receiving side may also perform zero-run decoding by bundling the displacement vector transform coefficients of multiple frames in the GOF.
[0413] The displacement vector encoder on the transmitting side is Fig.21 The transform coefficients of T(0,1,...,T-1) frames in the interleaved GoF are shown as follows or Fig. 22 In the case where the coefficients are serially packed to perform zero-run encoding as shown, the displacement vector transform coefficient decoder 15017 on the receiving side can use the same packing method to perform zero-run decoding on the transform coefficients of T (0, 1, ..., T-1) frames in GoF.
[0414] Fig.31 1 is a flowchart showing an example method for decoding transform coefficients of multiple frames according to an embodiment. That is, the figure shows an example of a decoding method performed by a displacement vector transform coefficient decoder on the receiving side when the transform coefficients of multiple frames have been packed and encoded by a displacement vector encoder on the transmitting side.
[0415] Reference Fig.31 , using the zeroRun decoder, transform coefficient deriver and transform coefficient decoder to The transform coefficient performs zeroRun decoding. For example, when the value of zeroRun parsed by the zeroRun decoder is not 0, the transform coefficient deriver derives the transform coefficient value of the vertex corresponding to the size of the parsed zeroRun as 0. When the value of the parsed zeroRun is 0 or when the value of zeroRun is deduced to 0, the transform coefficient decoder decodes the transform coefficient.
[0416] Then, after restoring all the transform coefficients, the transform coefficient divider may divide the divided coefficients according to a transform coefficient packing method (interleaving or serial packing) to restore the transform coefficients of each frame.
[0417] According to an embodiment, the displacement vector transform coefficient decoder 15017 may perform zero-run decoding per LoD (or sLoD) level by bundling transform coefficients of a plurality of frames per LoD (or sLoD) level (in an interleaved or serial manner).
[0418] Fig.32 1 is a flowchart illustrating another example method for decoding transform coefficients of multiple frames according to an embodiment. Specifically, the figure illustrates an example method for decoding and unpacking transform coefficients of multiple frames at a specific LoD (or sLoD) level when transform coefficients of multiple frames at the same LoD (or sLoD) level are packed and encoded by a displacement vector encoder at a transmitting side.
[0419] Reference Fig.32 , when the sum of the vertices of LoD n of multiple frames is When N LoDn Zero-run decoding is performed on transform coefficients in units of .
[0420] Then, a transform coefficient partitioner can be used to divide the transform coefficients of the nth LoD level of each frame into The zero-run decoding transform coefficients (N) of each level are divided frame by frame. LoDn ).here, Can be implicitly determined by the mesh subdivider 15015 based on the number of base meshes, where the level can be an sLoD level or a LoD level.
[0421] According to an embodiment, the level for scalable support (sLoD) may be defined separately from the LoD level determined by the grid subdivider 15015, and the displacement vector transform coefficients may be decoded on a per-level basis. Then, the grid scalability function may be performed by restoring the displacement vector transform coefficients up to a specific LoD (or sLoD) level in the LoD (or sLoD) level.
[0422] As described above, the displacement vector transform coefficient decoder 15017 decodes the transform coefficient by applying at least one of the first to third embodiments of the displacement vector decoding method. The decoding may be performed by applying a 2D video codec, or may be performed by applying a zero-run method. In addition, scalable decoding may be performed only on transform coefficients of a specific LoD level or sLoD level.
[0423] According to an embodiment, the transform coefficients decoded by the displacement vector transform coefficient decoder 15017 are provided to the displacement vector inverse transformer 15018.
[0424] The displacement vector inverse transformer 15018 performs an inverse transform corresponding to the transform performed by the encoder of the transmitting device. According to an embodiment, an inverse lifting transform, an inverse wavelet transform, etc. may be performed.
[0425] For example, assume that the encoder of the transmitting device performs a lifting transform. In this case, when predicting the vertex R at the k-th refinement level k the displacement vector of the subdivision vertex R t (t < k or t <= k) may be used as a predictor to perform the displacement vector at the k-th refinement level. According to an embodiment, when predicting the displacement vector, the average of n nearest points based on the connection information between vertices with a refinement level lower than the current vertex or a distance-based weighted average may be predicted. According to an embodiment, the prediction may be performed based on the displacement vectors of the n vertices used to generate the current vertex in the mesh subdivision step.
[0426] Then, when the displacement vector inverse transformer 15018 performs an inverse lifting transform, the parsed residual signal may be used to update the displacement vector of the vertex used for prediction by the encoder.
[0427] According to an embodiment, the displacement vector inverse quantizer 15019 may inverse-quantize the inverse-transformed displacement vector (i.e., the displacement vector transform coefficient).
[0428] According to an embodiment, the encoder of the transmitting device may quantize the transform coefficients with different quantization parameters for each axis, and may determine the quantization rate for each LoD level by deriving the quantization parameter or the scaling parameter according to the convention reached by the encoder / decoder.
[0429] According to an embodiment, when the reconstructed displacement vector has values in the normal coordinate system (n, t, bt), the displacement vector coordinate inverse transformer 15020 may inverse-transform it to the Cartesian coordinate system (x, y, z).
[0430] In other words, the encoder of the transmitting device may transform the vertex displacement vector calculated in the (x, y, z) space to the (normal, tangential, binormal) coordinate system (also referred to as the normal coordinate system) based on the normal vector of each vertex. Here, the normal vector may be calculated for each subdivision vertex based on the geometric information and the connection information about the neighboring vertices.
[0431] According to an embodiment, a unit including the vector inverse transformer 15018, the displacement vector inverse quantizer 15019, and the displacement vector coordinate inverse transformer 15020 may be referred to as a displacement vector reconstructor. That is, when the displacement vector transform coefficient decoded by the displacement vector transform coefficient decoder 15017 is processed by the displacement vector reconstructor, a final reconstructed displacement vector may be generated. In the case where the displacement vector transform coefficient decoder 15017 performs scalable decoding (e.g., 2D video decoding or zero-run decoding) only on the transform coefficient of a specific LoD level or sLoD level, the final reconstructed displacement vector is a displacement vector of the specific LoD level or sLoD level.
[0432] According to an embodiment, the received and demultiplexed texture map bitstream is input to the texture map decoder 15021. According to an embodiment, the texture map decoder 15021 may decode the texture map configured per a specific LoD level or sLoD level using a 2D scalable decoder. In the present disclosure, the LoD level is determined by the grid subdivider 15015, and the sLoD level may be defined by an agreement reached by the encoder / decoder based on the LoD level for scalable decoding, or may be configured by the received information. In other words, the texture map decoder 15021 may apply 2D scalable decoding to the texture map configured per LoD level or sLoD level to reconstruct the texture map. To this end, information related to the subdivision performed by the grid subdivider 15015, LoD related information and / or sLoD related information may be provided to the texture map decoder 15021 as signaling information. In addition, LoD related information and / or sLoD related information applied to the displacement vector scalable decoding performed by the displacement vector decoder 15017 may be provided to the texture map encoder 15021 as signaling information.
[0433] Fig.33 is a diagram showing an example of scalable decoding of a texture map by a texture map decoder according to an embodiment.
[0434] According to an embodiment, the texture map decoder 15021 can decode a level-specific texture map by parsing a sub-bitstream starting from a texture map of a lower LoD. Texture map decoding can be performed by referring to an inter-layer (level) reference of a texture map of a level lower than the current level.
[0435] Then, the grid scalable function may be performed by decoding the texture map up to a specific level among the level-specific texture map substreams. Here, the level may be a LoD level or an sLoD level. To this end, the texture map decoder 15021 may include as many scalable decoders as the LoD levels or sLoD levels. For example, the texture map decoder 15021 may perform decoding only on the texture map substream corresponding to the LoD0 level, may perform decoding only on the texture map substream corresponding to the LoD1 level, or may perform decoding only on the texture map substream corresponding to the LoD2 level.
[0436] According to an embodiment, the mesh reconstructor 15016 may generate a final reconstructed mesh by combining (i.e., adding) the reconstructed base mesh subdivided by the mesh subdivider 15015 and the displacement vector reconstructed by the displacement vector reconstructor. In this case, when the displacement vector reconstructed by the displacement vector reconstructor is a displacement vector up to a specific LoD level or sLoD level, the mesh reconstructed by the mesh reconstructor 15016 may also be a mesh up to a specific LoD level or sLoD level. In other words, the mesh reconstructor 15016 may reconstruct the mesh on a per LoD level or per sLoD level basis, and the mesh reconstructed on a per LoD level or per sLoD level basis may have reconstructed vertices, vertex connection information, texture coordinates, and texture coordinate connection information.
[0437] Therefore, when the displacement vector transform coefficient decoder 15017 performs scalable decoding to reconstruct the displacement vector of a specific level (e.g., LoD level or sLoD level), the mesh subdivider 15015 performs subdivision up to the level (e.g., LoD level or sLoD level). Then, by adding the reconstructed displacement vector of the level (e.g., LoD level or sLoD level) to the base mesh subdivided up to the level and applying the reconstructed texture map of the level to the added result, a final reconstructed mesh of the level can be constructed.
[0438] In this case, when the texture coordinates of the reconstructed mesh are not normalized values between 0 and 1, the mesh reconstructor 15016 may calculate the texture coordinates of the reconstructed mesh based on the W and H of the original texture map and the W of the texture map to be reconstructed. n and H n Normalization is performed to a value between 0 and 1, or scaling can be performed to (u,v), where 0≤u≤W n And 0≤v≤H n .
[0439] According to an embodiment, when using the reconstructed texture map (W n ×H n )When scaling texture coordinates in the range of 0≤u≤W, 0≤v≤H, the u,v coordinates can be calculated using the following formula 3.
[0440] [Formula 3]
[0441]
[0442] Therefore, the texture map decoder 15021 can decode the texture map bitstream of the specific LoD level or sLoD level using the video codec to reconstruct the texture map up to the specific LoD level or sLoD level. The reconstructed texture map has color information about each vertex included in the reconstructed mesh, and the color value of each vertex can be obtained from the texture map based on the texture coordinates of each vertex.
[0443] According to an embodiment, the reconstructed mesh from the mesh reconstructor 15016 and the reconstructed texture map from the texture map video decoder 15021 are presented to the user through a rendering process performed by a mesh data renderer (not shown). In one embodiment, when it is assumed that scalable decoding has been applied to the rendered reconstructed mesh and reconstructed texture map, the same LoD level or sLoD level may be applied.
[0444] As described above, signaling information is required to allow the encoder of the transmitting device to scalably encode the displacement vector and the texture map per LoD level or sLoD level, and to allow the decoder of the receiving device to scalably decode the displacement vector and the texture map encoded per LoD level or sLoD level. Signaling information is also required to allow the encoder of the transmitting device to perform 2D video encoding or zero-run encoding on the displacement vector per LoD level or sLoD level, and to allow the decoder of the receiving device to perform 2D video decoding or zero-run decoding on the displacement vector encoded per LoD level or sLoD level.
[0445] In the present disclosure, in a sending device, signaling information may be generated by an auxiliary information generator (not shown) in the sending device and provided to a corresponding block in the sending device and / or a receiving device, and an auxiliary information decoder (not shown) in the receiving device may parse the received signaling information and provide it to a corresponding block.
[0446] Figure 34 to Figure 37 An example syntax structure of signaling information related to the present disclosure is shown.
[0447] Fig.34 An example syntax structure of LoD related information (LoD_Info()) in signaling information according to an embodiment is shown.
[0448] According to an embodiment, the LoD related information (LoD_Info()) may include a subdivisionCount field and a sLoD_enable_flag field. In addition, according to a value of the sLoD_enable_flag field, the LoD related information (LoD_Info()) may further include a num_scalable_LoD_minus1 field.
[0449] The subdivisionCount field is an index indicating the number of times the mesh subdivider will perform subdivision. In other words, the subdivisionCount field indicates the number of LoD levels.
[0450] The sLoD_enable_flag field is a flag that determines whether to parse the level (sLoD) to perform scalable decoding on the receiving side. For example, the sLoD_enable_flag field set to 1 indicates that the encoder of the transmitting device has encoded the displacement vector transform coefficient based on sLoD, and the sLoD_enable_flag field set to 0 indicates that the encoder has encoded the coefficient based on LoD. Therefore, when the value of the sLoD_enable_flag field is 1, the decoder of the receiving device decodes the displacement vector transform coefficient based on sLoD, and when the value of the sLoD_enable_flag field is 0, the displacement vector transform coefficient is decoded based on LoD. For example, when the value of the sLoD_enable_flag field is 0, the mesh subdivision can perform encoding / decoding using the LoD level determined by subdivisionCount.
[0451] In the present disclosure, when a value of the sLoD_enable_flag field is 1, LoD related information (LoD_Info()) may further include a num_scalable_LoD_minus1 field.
[0452] The num_scalable_LoD_minus1 field indicates the number of sLoD levels minus 1. That is, the num_scalable_LoD_minus1 field is information for identifying the number of sLoD levels.
[0453] According to an embodiment, the LoD related information (LoD_Info()) may also include a loop that iterates as many times as the value of the num_scalable_LoD_minus1 field. In one embodiment, i may be initialized to 0 and incremented by 1 each time the loop is executed. The loop may iterate until the value of i is equal to the value of the num_scalable_LoD_minus1 field. The loop may include a LoD_level_idx[i] field and a LoD_level_idx[num_scalable_LoD_minus1]=subdivisionCount-1 field.
[0454] The LoD_level_idx[i] field indicates the index of the LoD level corresponding to the ith sLoD level. That is, the LoD_level_idx[i] field is the index of the LoD level corresponding to each sLoD level.
[0455] When the LoD consists of N levels in the mesh subdivider and the value of the num_scalable_LoD_minus1 field is k-1, the receiving device may implicitly derive the k-1th LoD_level_idx as N-1.
[0456] LoD_level_idx[num_scalable_LoD_minus1]=subdivisionCount-1 indicates that LoD_level_idx of the ith LoD level is set based on information about the number of LoD levels (subdivisionCount-1).
[0457] Fig.35 An example syntax structure of displacement vector decoding related information (Decode_Disp()) in signaling information according to an embodiment is shown.
[0458] According to an embodiment, the displacement vector decoding related information (Decode_Disp()) may include an isZero flag, an isOne flag, and an isK flag.
[0459] The isZero flag indicates the number of consecutive 0 bits. The isZero flag is set to 0 or a positive integer. When the isK flag is set to 1 and the value of the next bit is 0, the isZero flag is set to a value indicating the number of consecutive 0s. This is used to compress sequences of 0s.
[0460] In addition, the isOne flag indicates the number of consecutive 1 bits. The isOne flag is also set to zero or a positive integer. When the isK flag is set to 1 and the value of the next bit is 1, the isOne flag is set to a value indicating the number of consecutive 1s. This is used to compress sequences of 1s.
[0461] The isK flag (Is-Known flag) indicates what the value of the next bit is. The isK flag is set to 1 or 0. When set to 1, it indicates that the next bit value is a known value. For example, in the case of consecutive zeros, the isK flag is set to 1, and the value of the next bit is set to a known value of 0.
[0462] exist Fig.35 In the example, the group including isZero, isOne and isK is referred to as isN. That is, isN may include isZero, isOne, ... and isK. In other words, isN is a flag that determines whether the size (abs(disp)) of the transform coefficient is equal to N.
[0463] The size of the transform coefficient (abs(disp)) can be determined as follows.
[0464] abs(disp)==N? isN=1:isN=0
[0465] In other words, Fig.35 In the example above, when the value of the isZero flag is true (i.e., 0), the displacement transform vector coefficient (disp) is 0. Otherwise, the value of the isOne flag is checked. When the value of the isOne flag is true (i.e., 1), the displacement transform vector coefficient (disp) is 1. Additionally, when the value of the isOne flag is not true, the value of the isK flag is checked to see when it is true. When the value of the isK flag is true (i.e., K), the displacement transform vector coefficient (disp) is equal to K. When the value of the isK flag is not true, the displacement transform vector coefficient (disp) is determined by the value contained in Decode_Disp_Munus_Kplus1().
[0466] Fig.36 An example syntax structure of information (decode_displacement_coefficient()) related to decoding of a displacement vector transform coefficient in signaling information according to an embodiment is shown.
[0467] According to an embodiment, information related to decoding of a displacement vector transform coefficient (decode_displacement_coefficient()) may include a separate_zero_run_flag field.
[0468] When the coordinate system of the displacement vector is a normal coordinate system, the separate_zero_run_flag field may indicate whether to perform zero-run encoding on the normal component (n) and the tangential component (t, bt) separately. In other words, the separate_zero_run_flag field is a flag that determines whether to perform zero-run decoding on the normal component and the tangential component separately. For example, the separate_zero_run_flag field set to 1 indicates that the normal component and the tangential component are individually zero-run encoded, while the separate_zero_run_flag field set to 0 indicates that the normal component and the tangential component are simultaneously zero-run encoded. In this case, when the value of the separate_zero_run_flag field is 1, the receiving device or decoder may perform zero-run decoding on the normal component and the tangential component separately, and when the value of the separate_zero_run_flag field is 0, zero-run decoding may be performed on the normal component and the tangential component simultaneously.
[0469] Therefore, when the value of the separate_zero_run_flag field is 1, the decoded information about the displacement vector transform coefficient (decode_displacement_coefficient()) may include Decode_Normal_Coefficient() and Decode_Tangential_Coefficient(), and when the value of the separate_zero_run_flag field is 0, the decoded information about the displacement vector transform coefficient (decode_displacement_coefficient()) may include Decode_Coefficient().
[0470] Decode_Normal_Coefficient() may contain the displacement vector transform coefficients associated with the normal component, and Decode_Tangential_Coefficient() may contain the displacement vector transform coefficients associated with the tangential component.
[0471] Decode_Coefficient() may contain the displacement vector transformation coefficients associated with the normal and tangential components.
[0472] Fig.37 An example syntax structure of packing related information for multiple frames (unpack_displacemenst_for_multiframe()) in signaling information according to an embodiment is shown.
[0473] According to an embodiment, the packing related information of multiple frames (unpack_displacemenst_for_multiframe()) may include a displacement_packing_method field.
[0474] The displacement_packing_method field is an index indicating a method by which displacement vector transform coefficients of a plurality of frames should be packed when the displacement vector transform coefficients of the plurality of frames are encoded into a bundle. For example, a displacement_packing_method field set to 0 indicates that the frames are packed in an interleaved manner, while a displacement_packing_method field set to 1 indicates that the frames are packed in a serial manner.
[0475] Fig.38 is a flow chart showing an example transmission method according to an embodiment. The transmission method according to an embodiment may include: encoding mesh data (21011); and transmitting a bit stream including the encoded mesh data (21012).
[0476] According to an embodiment, the encoding (21011) of the mesh data may include defining LoD levels and / or sLoD levels when subdividing the base mesh, and performing displacement vector encoding and texture map encoding for each LoD level or sLoD level. The displacement vector encoding may be performed using a conventional 2D video codec, or may be performed using a zero-run encoder. For details of 2D video encoding or zero-run encoding of displacement vectors for each LoD level or sLoD level and 2D video encoding of texture maps for each LoD level or sLoD level (not described below to avoid redundancy), refer to Figures 15 to 25 In addition, for details of the signaling information of 2D video coding or zero-run coding of displacement vectors per LoD level or sLoD level and 2D video coding of texture maps per LoD level or sLoD level (not described below to avoid redundancy), refer to Figure 34 to Figure 37 Description.
[0477] Fig.39 is a flowchart illustrating an example receiving method according to an embodiment. The receiving method according to an embodiment may include: receiving a bit stream containing mesh data (22011); and decoding the mesh data contained in the bit stream (22012).
[0478] In the present disclosure, decoding (22012) of mesh data included in a bitstream may include performing displacement vector decoding and texture map decoding per LoD level or sLoD level based on signaling information of a multiplexed displacement vector bitstream and a texture map bitstream. Displacement vector decoding may be performed using a conventional 2D video codec or a zero-run decoder. For details of 2D video decoding or zero-run decoding of displacement vectors per LoD level or sLoD level and 2D video encoding of texture maps per LoD level or sLoD level (not described below to avoid redundancy), refer to Figures 15 to 33 In addition, for details of the signaling information of 2D video decoding or zero-run decoding of displacement vectors per LoD level or sLoD level and 2D video decoding of texture maps per LoD level or sLoD level (not described below to avoid redundancy), refer to Figure 34 to Figure 37 Description.
[0479] Therefore, since the existing V-Mesh compression method is designed to compress and reconstruct the displacement vector video and texture map video generated using a video codec during the encoding process based on the overall resolution, the received mesh data may not be used when it is difficult to receive and utilize the entire mesh data according to the conditions of the receiver. In order to solve this problem, a scalable mesh encoding method of a transmitting device and a scalable mesh decoding method of a receiving device are proposed, so that the mesh can be selectively reconstructed and utilized on the receiving side. In particular, the present disclosure proposes a method for scalable encoding (on the transmitting side) and scalable decoding (on the receiving side) of displacement vector video and texture map video, a method for encoding and decoding displacement vector transform coefficients with improved performance, and a method for performing scalable encoding and decoding based on displacement vector transform coefficients. Therefore, the present disclosure enables users to selectively reconstruct mesh data based on mesh subdivision levels according to the performance or display characteristics of the receiver and the network environment, up to the available level. In addition, when multiple mesh data objects are to be represented, users may be allowed to specify different levels of detail and levels of mesh data to be reconstructed according to the importance, frequency of use, and current status of the objects, and reconstruct each object according to the desired level. Thus, given resources can be appropriately allocated and efficiently used. Therefore, the present disclosure can enable mesh data to be used in a wider range of network environments and applications, and further expand the scope of utilization of mesh data by enabling resources of a receiver to be used more flexibly.
[0480] The above-mentioned various parts, modules or units can be software, processors or hardware parts that execute the continuous process stored in the memory (or storage unit). The various steps described in the above-mentioned embodiments can be performed by a processor, software or hardware part. The various modules / blocks / units described in the above-mentioned embodiments can be operated as a processor, software or hardware. In addition, the method presented in the embodiment can be executed as a code. The code can be written on a processor-readable storage medium, so it is read by a processor provided by the device.
[0481] In this specification, when a part "includes" or "contains" an element, unless otherwise mentioned, it means that the part also includes or contains another element. In addition, the term "...module (or unit)" disclosed in this specification means a unit for processing at least one function or operation, and can be implemented by hardware, software, or a combination of hardware and software.
[0482] Although the embodiments have been described with reference to the accompanying drawings for simplicity, new embodiments may be designed by combining the embodiments shown in the accompanying drawings. If a person skilled in the art designs a computer-readable recording medium having a program for executing the embodiments mentioned in the above description, it may fall within the scope of the attached claims and their equivalents.
[0483] The apparatus and method may not be limited to the configuration and method of the above-described embodiments. The above-described embodiments may be configured by selectively combining with each other entirely or partially, so that various modifications can be made.
[0484] Although preferred embodiments of the embodiments have been shown and described, the embodiments are not limited to the specific embodiments described above, and various modifications may be made by a person skilled in the art without departing from the spirit of the embodiments claimed for protection in the claims, and these modifications should not be understood in isolation from the technical ideas or vision of the embodiments.
[0485] Various elements of the device of the embodiment can be implemented by hardware, software, firmware or a combination thereof. Various elements in the embodiment can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiment, the components according to the embodiment can be implemented as separate chips respectively. According to the embodiment, at least one or more components of the device according to the embodiment may include one or more processors capable of executing one or more programs. The one or more programs may execute any one or more of the operations / methods according to the embodiment, or include instructions for executing them. The executable instructions for executing the method / operation of the device according to the embodiment may be stored in a non-transitory CRM or other computer program products configured to be executed by one or more processors, or may be stored in a temporary CRM or other computer program products configured to be executed by one or more processors. In addition, the memory according to the embodiment can be used as a concept that not only covers volatile memory (e.g., RAM) but also covers non-volatile memory, flash memory and PROM. In addition, it can also be implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, the processor-readable recording medium can be distributed in a computer system connected via a network so that the processor-readable code can be stored and executed in a distributed manner.
[0486] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". "A / B / C" may mean "at least one of A, B, and / or C". In addition, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "in addition or alternatively".
[0487] The various elements of the embodiments may be implemented by hardware, software, firmware or a combination thereof. The various elements of the embodiments may be performed by a single chip (e.g., a single hardware circuit). Depending on the embodiment, the elements may be selectively performed by separate chips, respectively. Depending on the embodiment, at least one element of the embodiments may be performed in one or more processors including instructions for performing operations according to the embodiments.
[0488] The operation according to the embodiment described in this specification may be performed by a sending / receiving device including one or more memories and / or one or more processors according to the embodiment. One or more memories may store programs for processing / controlling the operation according to the embodiment, and one or more processors may control the various operations described in this specification. One or more processors may be referred to as controllers, etc. In an embodiment, the operation may be performed by firmware, software, and / or a combination thereof. Firmware, software, and / or a combination thereof may be stored in a processor or memory.
[0489] The various elements of the embodiment can be described using terms such as the first and second. However, the various components according to the embodiment should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, the first user input signal can be referred to as the second user input signal. Similarly, the second user input signal can be referred to as the first user input signal. The use of these terms should be interpreted as not departing from the scope of various embodiments. Both the first user input signal and the second user input signal are user input signals, but do not mean the same user input signal unless the context clearly stipulates otherwise. The terms used to describe the embodiment are only used to describe a specific embodiment and are not intended to limit the embodiment. As used in the description of the embodiment and the claims, unless the context clearly stipulates otherwise, the singular form "one", "one" and "the" include plural indicators. The expression "and / or" is used to include all possible combinations of items. Terms such as "including" or "having" are intended to indicate the presence of figures, numbers, steps, elements and / or components, and should be understood as not excluding the possibility of additional presence of figures, numbers, steps, elements and / or components.
[0490] As used herein, conditional expressions such as "if" and "when" are not limited to optional situations and are intended to be interpreted as performing relevant operations or interpreting relevant definitions according to specific conditions when specific conditions are met. Implementation methods may include changes / modifications within the scope of the claims and their equivalents. It will be apparent to those skilled in the art that various modifications and variations may be made in the present disclosure without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is intended to cover modifications and variations of the present disclosure as long as they fall within the scope of the appended claims and their equivalents.
[0491] Mode of the present disclosure
[0492] As described above, the relevant contents have been described in the best mode for carrying out the embodiment.
[0493] Industrial Applicability
[0494] As described above, the embodiments may be applied in whole or in part to 3D data transmission / reception devices and systems. It will be apparent to those skilled in the art that various changes or modifications may be made to the embodiments within the scope of the embodiments. Therefore, the embodiments are intended to cover modifications and variations as long as they fall within the scope of the appended claims and their equivalents.
Claims
1. A method for transmitting three-dimensional 3D data, the method comprising the following steps: Encode the raw grid data; as well as A bitstream containing the encoded mesh data and signaling information is sent.
2. The method according to claim 1, wherein: The encoding of the grid data comprises the following steps: Generating basic mesh data by simplifying the original mesh data, and encoding the basic mesh data; generating additional vertices by subdividing the simplified mesh data one or more times, and determining one or more levels that vary according to the number of subdivisions; Reconstruct the encoded basic mesh data; Generating displacement information based on the subdivided mesh data and the reconstructed basic mesh data; encoding the displacement information; Reconstruct the encoded displacement information; reconstructing the mesh data based on the reconstructed basic mesh data and the reconstructed displacement information; regenerating a texture map based on the texture map of the original mesh data and the reconstructed mesh data; and Encode the regenerated texture map.
3. The method according to claim 2, wherein: The encoding of the displacement information comprises the following steps: The displacement information corresponding to at least one of the one or more levels is encoded using a 2D video codec.
4. The method according to claim 2, wherein: The encoding of the displacement information comprises the following steps: The displacement information corresponding to at least one of the one or more levels is encoded using zero-run encoding.
5. The method according to claim 2, wherein: The encoding of the displacement information comprises the following steps: packing the displacement information in a plurality of frames corresponding to at least one of the one or more levels; and The packed displacement information is encoded using run-length zero encoding.
6. The method according to claim 5, wherein: The displacement information in the multiple frames is packed using an interleaving method or a serial method based on the mapping relationship of the vertices between the multiple frames.
7. The method according to claim 5, wherein: The signaling information includes information related to the one or more levels or information related to packing of the displacement information in the plurality of frames.
8. The method according to claim 2, wherein: The one or more levels are at least one level of detail LoD or at least one scalable LoD sLoD.
9. The method according to claim 8, wherein: The at least one sLoD is configured based on the at least one LoD, Wherein, a specific sLoD among the at least one sLoD is mapped to one of the at least one LoD.
10. The method according to claim 9, wherein: The number of levels in the at least one sLoD is different from the number of levels in the at least one LoD.
11. The method according to claim 2, wherein: The encoding of the texture map comprises the following steps: The texture map corresponding to at least one of the one or more levels is encoded using a 2D video codec.
12. A device for transmitting three-dimensional (3D) data, the device comprising: an encoder configured to encode raw mesh data; as well as A transmitter is configured to transmit a bit stream containing encoded mesh data and signaling information.
13. The device according to claim 12, wherein: The encoder comprises: a base mesh compressor configured to generate base mesh data by simplifying the original mesh data and to encode the base mesh data; a mesh subdivider configured to generate additional vertices by subdividing the simplified mesh data one or more times, and determine one or more levels that vary according to the number of subdividings; a base mesh reconstructor configured to reconstruct the encoded base mesh data; a displacement information generator configured to generate displacement information based on the subdivided mesh data and the reconstructed base mesh data; a displacement information encoder configured to encode the displacement information; a displacement information reconstructor configured to reconstruct the encoded displacement information; a mesh reconstructor configured to reconstruct mesh data based on the reconstructed base mesh data and the reconstructed displacement information; a texture map generator configured to regenerate a texture map based on the texture map of the original mesh data and the reconstructed mesh data; and A texture map generator is configured to encode the regenerated texture map.
14. A method for receiving three-dimensional (3D) data, the method comprising the following steps: receiving a bit stream comprising encoded mesh data and signaling information; Decoding the encoded mesh data in the bitstream based on the signaling information; as well as Renders the decoded mesh data.
15. The method according to claim 14, wherein: The decoding of the grid data comprises the following steps: reconstructing base mesh data from the encoded mesh data; generating additional vertices by subdividing the reconstructed base mesh data one or more times, and determining one or more levels based on the signaling information and the number of subdivisions; decoding and reconstructing displacement information from the encoded mesh data; Reconstructing mesh data based on the subdivided basic mesh data and the reconstructed displacement information; decoding and reconstructing a texture map from the encoded mesh data; and Rendering is performed based on the reconstructed mesh data and the reconstructed texture map.