Point cloud data sending device, point cloud data sending method, point cloud data receiving device and point cloud data receiving method

Through geometry-based point cloud compression technology, spatial partition encoding of point cloud data is solved, and the problems of low point cloud data transmission efficiency and insufficient compression performance in the existing technology are achieved, and efficient point cloud data processing and transmission are achieved.

CN114097229BActive Publication Date: 2025-05-13LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080048834.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-03
Filing Date
2020-05-28
Publication Date
2025-05-13
Estimated Expiration
2040-05-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively process and transmit large amounts of point cloud data, resulting in high latency and encoding/decoding complexity and insufficient point cloud compression performance.

Method used

Geometry-based point cloud compression (G-PCC) technology is used to divide point cloud data into slices or tiles through spatial partitioning, and encode them to generate a bitstream containing geometric slices, attribute slices and signaling information.

Benefits of technology

It improves the transmission efficiency and decoding performance of point cloud data, reduces the delay and encoding complexity, and enhances the performance of point cloud compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114097229B_ABST
    Figure CN114097229B_ABST
Patent Text Reader

Abstract

The point cloud data transmitting method according to the embodiment may include acquiring point cloud data, encoding the point cloud data, and transmitting a bit stream including the encoded point cloud data and signaling information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments relate to methods and apparatus for processing point cloud content. Background Art

[0002] Point cloud content is content represented by a point cloud, which is a collection of points belonging to a coordinate system representing a three-dimensional space. Point cloud content can express media configured in three dimensions and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), XR (extended reality), and self-driving. However, tens of thousands to hundreds of thousands of point data are required to represent point cloud content. Therefore, a method for efficiently processing large amounts of point data is needed. Summary of the invention

[0003] Technical issues

[0004] The purpose of the present disclosure designed to solve the above-mentioned problems is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device and a point cloud data receiving method for effectively sending and receiving point clouds.

[0005] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device and a point cloud data receiving method for solving time delay and encoding / decoding complexity.

[0006] Another object of the present disclosure is to provide a point cloud data sending device, a point cloud data sending method, a point cloud data receiving device and a point cloud data receiving method, which can improve the point cloud compression performance by improving the encoding of the attributes of geometric point cloud compression (G-PCC).

[0007] The objects of the present disclosure are not limited to the aforementioned objects, and other objects of the present disclosure not mentioned above will become clear to those of ordinary skill in the art after reviewing the following description.

[0008] Technical Solution

[0009] To achieve these objectives and other advantages and in accordance with an embodiment, a method of transmitting point cloud data may include acquiring point cloud data, encoding the point cloud data, and transmitting a bit stream including the encoded point cloud data and signaling information.

[0010] In an embodiment, encoding includes: spatially partitioning the geometric information and attribute information of the point cloud data into slices or tiles, each of the tiles including one or more slices, and encoding the geometric information and attribute information of the point cloud data based on the slices or tiles.

[0011] In an embodiment, a bitstream consists of geometry-based point cloud compression (G-PCC) units, each of the G-PCC units having a header and a payload, the header including type information for identifying the payload included in the payload, the payload including at least one of a geometry slice, an attribute slice, or signaling information, and the signaling information is one of a sequence parameter set including sequence level information, a geometry parameter set including geometry-related information, an attribute parameter set including attribute-related information, and a tile parameter set including tile-related information.

[0012] In an embodiment, a geometry slice includes a geometry slice header and geometry slice data, the geometry slice data including a slice-based geometry bitstream, and the geometry slice header includes at least identification information for identifying a geometry parameter set referenced by the geometry bitstream or information related to the slice and / or tile to which the geometry bitstream belongs.

[0013] In an embodiment, an attribute slice includes an attribute slice header and attribute slice data, the attribute slice data includes an attribute bitstream based on the slice, and the attribute slice header includes at least identification information for identifying an attribute parameter set referenced by the attribute bitstream or identification information for identifying a geometric slice related to the attribute bitstream.

[0014] According to an embodiment, an apparatus for sending point cloud data may include: an acquirer configured to acquire point cloud data, an encoder configured to encode the point cloud data, and a transmitter configured to send a bit stream including the encoded point cloud data and signaling information.

[0015] In one embodiment, the encoder includes a spatial partitioner configured to spatially partition the geometric information and attribute information of the point cloud data into slices or tiles, each of the tiles including one or more slices; and a video encoder configured to encode the geometric information and attribute information of the point cloud data based on the slices or tiles.

[0016] In an embodiment, a bitstream consists of geometry-based point cloud compression (G-PCC) units, each of the G-PCC units having a header and a payload, the header including type information for identifying data included in the payload, the payload including at least one of a geometry slice, an attribute slice, or signaling information, the signaling information being one of a sequence parameter set including sequence level information, a geometry parameter set including geometry-related information, an attribute parameter set including attribute-related information, and a tile parameter set including tile-related information.

[0017] In an embodiment, a geometry slice includes a geometry slice header and geometry slice data, the geometry slice data including a slice-based geometry bitstream, and the geometry slice header includes at least identification information for identifying a geometry parameter set referenced by the geometry bitstream or information related to the slice and / or tile to which the geometry bitstream belongs.

[0018] In an embodiment, an attribute slice includes an attribute slice header and attribute slice data, the attribute slice data includes a slice-based attribute bitstream, and the attribute slice header includes at least identification information for identifying an attribute parameter set referenced by the attribute bitstream or identification information for identifying a geometric slice related to the attribute bitstream.

[0019] According to an embodiment, a method of receiving point cloud data may include: receiving a bit stream including point cloud data and signaling information, decoding the point cloud data based on the signaling information, and rendering the point cloud data.

[0020] In an embodiment, decoding includes decoding geometric information and attribute information of point cloud data based on slices or tiles including one or more slices based on signaling information, and spatially combining the slice-based or tile-based decoded geometric information and attribute information based on the signaling information.

[0021] In an embodiment, a bitstream consists of geometry-based point cloud compression (G-PCC) units, each of the G-PCC units having a header and a payload, the header including type information for identifying data included in the payload, the payload including at least one of a geometry slice, an attribute slice, or signaling information, the signaling information being one of a sequence parameter set including sequence level information, a geometry parameter set including geometry-related information, an attribute parameter set including attribute-related information, and a tile parameter set including tile-related information.

[0022] In an embodiment, a geometry slice includes a geometry slice header and geometry slice data, the geometry slice data including a slice-based geometry bitstream, and the geometry slice header includes at least identification information for identifying a geometry parameter set referenced by the geometry bitstream or information related to the slice and / or tile to which the geometry bitstream belongs.

[0023] In an embodiment, the attribute slice includes an attribute slice header and attribute slice data, the attribute slice data includes a slice-based attribute bitstream, and the attribute slice header includes at least identification information for identifying an attribute parameter set referenced by the attribute bitstream or identification information for identifying a geometric slice related to the attribute bitstream.

[0024] According to an embodiment, an apparatus for receiving point cloud data may include: a receiver configured to receive a bit stream including point cloud data and signaling information; a decoder configured to decode the point cloud data based on the signaling information; and a renderer configured to render the point cloud data.

[0025] In an embodiment, the decoder includes a video decoder configured to decode geometric information and attribute information of point cloud data based on a slice or a tile including one or more slices based on signaling information, and a post-processor configured to spatially combine the decoded geometric information and attribute information based on a slice or a tile based on the signaling information.

[0026] In an embodiment, a bitstream is composed of geometry-based point cloud compression (G-PCC) units, each of the G-PCC units having a header and a payload, the header including type information for identifying data included in the payload, the payload including at least one of a geometry slice, an attribute slice, or signaling information, the signaling information being one of a sequence parameter set including sequence level information, a geometry parameter set including geometry-related information, an attribute parameter set including attribute-related information, and a tile parameter set including tile-related information.

[0027] In an embodiment, a geometry slice includes a geometry slice header and geometry slice data, the geometry slice data includes a slice-based geometry bitstream, and the geometry slice header includes at least identification information for identifying a geometry parameter set referenced by the geometry bitstream or information related to the slice and / or tile to which the geometry bitstream belongs.

[0028] In an embodiment, the attribute slice includes an attribute slice header and attribute slice data, the attribute slice data includes a slice-based attribute bitstream, and the attribute slice header includes at least identification information for identifying an attribute parameter set referenced by the attribute bitstream or identification information for identifying a geometric slice related to the attribute bitstream.

[0029] Beneficial Effects

[0030] The point cloud data transmitting method, the point cloud data transmitting device, the point cloud data receiving method, and the point cloud data receiving device according to the embodiments can provide a point cloud service of good quality.

[0031] According to the point cloud data sending method, point cloud data sending device, point cloud data receiving method and point cloud data receiving device of the embodiments, various video encoding and decoding methods can be implemented.

[0032] The point cloud data transmitting method, the point cloud data transmitting device, the point cloud data receiving method, and the point cloud data receiving device according to the embodiments may provide general point cloud content such as a self-driving service.

[0033] The point cloud data transmitting method, point cloud data transmitting device, point cloud data receiving method and point cloud data receiving device according to the embodiments can perform spatially adaptive partitioning of point cloud data for independently encoding and decoding point cloud data, thereby improving parallel processing and providing scalability.

[0034] According to the embodiment, the point cloud data sending method, point cloud data sending device, point cloud data receiving method and point cloud data receiving device can perform encoding and decoding by spatially partitioning the point cloud data in units of tiles and / or slices, and sending necessary data for it with signals, thereby improving the encoding and decoding performance of the point cloud.

[0035] According to the embodiment, the point cloud data sending method, the point cloud data sending device, the point cloud data receiving method and the point cloud data receiving device can configure the G-PCC bit stream with a G-PCC unit including geometry, attributes and signaling information, thereby improving encoding and decoding performance.

[0036] According to the point cloud data sending method, point cloud data sending device, point cloud data receiving method and point cloud data receiving device of the embodiment, the G-PCC unit can be encapsulated into a G-PCC access unit to send and receive the G-PCC access unit, thereby supporting efficient access to the point cloud based on the G-PCC access unit, especially the G-PCC bit stream. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings, which are included to provide a further understanding of the disclosure and are incorporated in and constitute a part of this application, illustrate embodiments of the disclosure and together with the description serve to explain the principle of the disclosure.

[0038] Figure 1 An exemplary point cloud content providing system according to an embodiment is illustrated.

[0039] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.

[0040] Figure 3 An exemplary process of capturing point cloud video according to an embodiment is illustrated.

[0041] Figure 4 An exemplary block diagram of a point cloud video encoder according to an embodiment is illustrated.

[0042] Figure 5 Illustrated is an example of voxels in 3D space according to an embodiment.

[0043] Figure 6 Illustrated are examples of an octree and occupancy codes according to an embodiment.

[0044] Figure 7 An example of a neighbor node pattern according to an embodiment is illustrated.

[0045] Figure 8 Illustrated is an example of point configuration for point cloud content of each LOD according to an embodiment.

[0046] Figure 9 Illustrated is an example of point configuration for point cloud content of each LOD according to an embodiment.

[0047] Figure 10 An example of a block diagram of a point cloud video decoder according to an embodiment is illustrated.

[0048] Figure 11 An example of a point cloud video decoder according to an embodiment is illustrated.

[0049] Figure 12 A configuration of point cloud video encoding for a transmitting device according to an embodiment is illustrated.

[0050] Figure 13 The diagram illustrates a configuration for point cloud video decoding of a receiving device according to an embodiment.

[0051] Figure 14 Illustrated is an architecture for storing and streaming G-PCC based point cloud data according to an embodiment.

[0052] Figure 15 An example of storage and transmission of point cloud data according to an embodiment is illustrated.

[0053] Figure 16 An example of a receiving device according to an embodiment is illustrated.

[0054] Figure 17 An exemplary structure operatively connected to a method / apparatus for transmitting and receiving point cloud data according to an embodiment is illustrated.

[0055] Figure 18 An example of a point cloud transmitting device according to an embodiment is illustrated.

[0056] Figure 19 (a) to Figure 19 (c) illustrates an embodiment of partitioning a bounding box into one or more tiles.

[0057] Figure 20 An example of a point cloud receiving device according to an embodiment is illustrated.

[0058] Figure 21 An exemplary bitstream structure for transmitting / receiving point cloud data according to an embodiment is illustrated.

[0059] Figure 22 (a) and Figure 22(b) is a diagram illustrating an example of a bit stream structure of point cloud data and a connection relationship between elements in the bit stream according to an embodiment.

[0060] Figure 23 An embodiment of a syntax structure of a sequence parameter set according to the present disclosure is shown.

[0061] Figure 24 is a table showing examples of attribute types assigned to the attribute_label_four_bytes field according to an embodiment.

[0062] Figure 25 An embodiment of a syntax structure of a tile parameter set according to the present disclosure is shown.

[0063] Figure 26 An embodiment of a syntax structure of a geometry parameter set according to the present disclosure is shown.

[0064] Figure 27 An embodiment of a syntax structure of a property parameter set according to the present disclosure is shown.

[0065] Figure 28 is a table showing an example of attribute coding types assigned to the attr_coding_type field according to an embodiment.

[0066] Figure 29 An embodiment of a syntax structure of geometry_slice_bitstream() according to the present disclosure is shown.

[0067] Figure 30 An embodiment of a syntax structure of a geometry slice header according to the present disclosure is shown.

[0068] Figure 31 An embodiment of a syntax structure of geometry slice data according to the present disclosure is shown.

[0069] Figure 32 An embodiment of a syntax structure of attribute_slice_bitstream() according to the present disclosure is shown.

[0070] Figure 33 An embodiment of the syntax structure of the attribute slice header according to the present disclosure is shown.

[0071] Figure 34 An embodiment of a syntax structure of attribute slice data according to the present disclosure is shown.

[0072] Figure 35 An example of a G-PCC bitstream structure according to an embodiment is shown.

[0073] Figure 36An embodiment of a syntax structure of metadata_slice_bitstream() according to the present disclosure is shown.

[0074] Figure 37 An embodiment of the syntax structure of the metadata slice header according to the present disclosure is shown.

[0075] Figure 38 An embodiment of the syntax structure of metadata slice data according to the present disclosure is shown.

[0076] Figure 39 An exemplary syntax structure of each G-PCC unit according to an embodiment is shown.

[0077] Figure 40 An exemplary syntax structure of a G-PCC unit header according to an embodiment is shown.

[0078] Figure 41 An example of a G-PCC unit type assigned to the gpcc_unit_type field according to an embodiment is shown.

[0079] Figure 42 An exemplary syntax structure of a G-PCC unit payload according to an embodiment is shown.

[0080] Figure 43 An exemplary structure of a G-PCC access unit according to an embodiment is shown.

[0081] Figure 44 An exemplary syntax structure of a G-PCC access unit header according to an embodiment is shown.

[0082] Figure 45 An exemplary syntax structure of a G-PCC access unit payload according to an embodiment is shown.

[0083] Figure 46 is a flowchart illustrating a method of transmitting point cloud data according to an embodiment.

[0084] Figure 47 is a flowchart illustrating a method of receiving point cloud data according to an embodiment. Specific embodiments

[0085] Now, a description will be given in detail according to the exemplary embodiments disclosed herein with reference to the accompanying drawings. For the sake of brief description with reference to the accompanying drawings, the same reference numerals may be provided for the same or equivalent components, and their descriptions will not be repeated. It should be noted that the following examples are only used to embody the present disclosure and do not limit the scope of the present disclosure. What can be easily inferred from the detailed description and examples of the present disclosure by experts in the technical field to which the present invention belongs will be interpreted as being within the scope of the present disclosure.

[0086] The detailed description in this specification should be interpreted in all aspects as illustrative rather than restrictive. The scope of the present disclosure should be determined by the appended claims and their legal equivalents, and all changes coming within the meaning and equivalency range of the appended claims are intended to be embraced herein.

[0087] Now, reference will be made in detail to the preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The detailed description given below with reference to the accompanying drawings is intended to explain exemplary embodiments of the present disclosure, rather than to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. Although most of the terms used in this specification have been selected from common terms widely used in the art, the applicant has arbitrarily selected some terms, and their meanings will be explained in detail as needed in the following description. Therefore, the present disclosure should be understood based on the original meaning of the terms rather than their simple names or meanings. In addition, the following drawings and detailed descriptions should not be interpreted as being limited to the specifically described embodiments, but should be interpreted as including equivalents or substitutes of the embodiments described in the drawings and detailed descriptions.

[0088] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown.

[0089] Figure 1 The point cloud content providing system illustrated in FIG. 1 may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 may perform wired or wireless communication to transmit and receive point cloud data.

[0090] The point cloud data sending device 10000 according to an embodiment can protect and process point cloud video (or point cloud content), and send the point cloud video (or point cloud content). According to an embodiment, the sending device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device and / or a server. According to an embodiment, the sending device 10000 may include a device configured to communicate with a base station and / or other wireless devices using a radio access technology (e.g., 5G new RAT (NR), long term evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server.

[0091] According to the embodiment, the sending device 10000 includes a point cloud video acquisition unit 10001, a point cloud video encoder 10002 and / or a transmitter (or communication module) 10003.

[0092] The point cloud video acquisition unit 10001 according to the embodiment acquires the point cloud video through a processing process such as capturing, synthesizing or generating. The point cloud video is a point cloud content represented by a point cloud as a collection of points in a 3D space, and can be referred to as point cloud video data. The point cloud video according to the embodiment may include one or more frames. A frame represents a still image / picture. Therefore, the point cloud video may include a point cloud image / frame / picture, and may be referred to as a point cloud image, frame or picture.

[0093] The point cloud video encoder 10002 according to the embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 may encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next generation coding. The point cloud compression coding according to the embodiment is not limited to the above-mentioned embodiment. The point cloud video encoder 10002 may output a bit stream containing encoded point cloud video data. The bit stream may contain not only the encoded point cloud video data, but also signaling information related to the coding of the point cloud video data.

[0094] The transmitter 10003 according to the embodiment transmits a bitstream containing encoded point cloud video data. The bitstream according to the embodiment is encapsulated in a file or segment (e.g., a streaming segment) and transmitted through various networks such as a broadcast network and / or a broadband network. Although not shown in the figure, the sending device 10000 may include an encapsulator (or encapsulation module) configured to perform an encapsulation operation. According to an embodiment, the encapsulator may be included in the transmitter 10003. According to an embodiment, the file or segment may be sent to the receiving device 10004 through a network, or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). The transmitter 10003 according to the embodiment can perform wired / wireless communication with the receiving device 10004 (or receiver 10005) through a network such as 4G, 5G, 6G, etc. In addition, the transmitter can perform necessary data processing operations according to the network system (e.g., 4G, 5G or 6G communication network system). The sending device 10000 can send encapsulated data in an on-demand manner.

[0095] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006 and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include a device configured to communicate with a base station and / or other wireless devices using a radio access technology (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), a robot, a vehicle, an AR / VR / XR device, a portable device, a home appliance, an Internet of Things (IoT) device, and an AI device / server.

[0096] The receiver 10005 according to the embodiment receives a bitstream containing point cloud video data or a file / segment in which a bitstream is encapsulated from a network or storage medium. The receiver 10005 can perform necessary data processing according to a network system (e.g., a communication network system such as 4G, 5G, 6G, etc.). The receiver 10005 according to the embodiment can decapsulate the received file / segment and output a bitstream. According to the embodiment, the receiver 10005 may include a decapsulator (or decapsulation module) configured to perform a decapsulation operation. The decapsulator can be implemented as an element (or component) separate from the receiver 10005.

[0097] The point cloud video decoder 10006 decodes the bitstream containing the point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to the method of encoding the point cloud video data (for example, in the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression coding, which is the reverse process of point cloud compression. Point cloud decompression coding includes G-PCC coding.

[0098] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 may output the point cloud content by rendering not only the point cloud video data but also the audio data. According to an embodiment, the renderer 10007 may include a display configured to display the point cloud content. According to an embodiment, the display may be implemented as a separate device or component rather than being included in the renderer 10007.

[0099] The arrow indicated by the dotted line in the figure represents the transmission path of the feedback information obtained by the receiving device 10004. Feedback information is information used to reflect the interactivity with the user consuming the point cloud content, and includes information about the user (e.g., header orientation information, viewport information, etc.). In particular, when the point cloud content is the content of a service that requires interaction with the user (e.g., self-driving service, etc.), the feedback information may be provided to the content sender (e.g., sending device 10000) and / or the service provider. Depending on the embodiment, the feedback information may be used in the receiving device 10004 and the sending device 10000, or may not be provided.

[0100] The header orientation information according to an embodiment is information about the user's header position, orientation, angle, motion, etc. The receiving device 10004 according to an embodiment may calculate the viewport information based on the header orientation information. The viewport information may be information about the area of ​​the point cloud video that the user is watching. The viewpoint is the point through which the user is watching the point cloud video, and may refer to the center point of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size and shape of the area may be determined by the field of view (FOV). Therefore, in addition to the header orientation information, the receiving device 10004 may also extract the viewport information based on the vertical or horizontal FOV supported by the device. In addition, the receiving device 10004 performs gaze analysis, etc., to check the way the user consumes the point cloud, the area where the user gazes in the point cloud video, the gaze time, etc. According to an embodiment, the receiving device 10004 may send feedback information including the gaze analysis result to the sending device 10000. The feedback information according to the embodiment may be obtained in the rendering and / or display process. The feedback information according to the embodiment may be protected by one or more sensors included in the receiving device 10004. According to an embodiment, the feedback information may be protected by the renderer 10007 or a separate external element (or device, component, etc.). Figure 1 The dotted line in represents the process of sending feedback information protected by the renderer 10007. The point cloud content providing system can process (encode / decode) the point cloud data based on the feedback information. Therefore, the point cloud video decoder 10006 can perform a decoding operation based on the feedback information. The receiving device 10004 can send the feedback information to the sending device 10000. The sending device 10000 (or the point cloud video encoder 10002) can perform an encoding operation based on the feedback information. Therefore, the point cloud content providing system can efficiently process necessary data (for example, point cloud data corresponding to the user header position) based on the feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide point cloud content to the user.

[0101] According to an embodiment, the sending device 10000 may be referred to as an encoder, a sending device, a transmitter, a sending system, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, a receiving system, etc.

[0102] (Through a series of processes of obtaining / encoding / sending / decoding / rendering) in accordance with the embodiment Figure 1 The point cloud data processed in the point cloud content providing system may be referred to as point cloud content data or point cloud video data. According to an embodiment, point cloud content data may be used as a concept covering metadata or signaling information related to point cloud data.

[0103] Figure 1The elements of the point cloud content providing system illustrated in the figure may be implemented by hardware, software, a processor and / or a combination thereof.

[0104] Figure 2 is a block diagram illustrating a point cloud content providing operation according to an embodiment.

[0105] Figure 2 The block diagram shows Figure 1 The operation of the point cloud content providing system described in . As described above, the point cloud content providing system can process point cloud data based on point cloud compression compilation (e.g., G-PCC).

[0106] According to the embodiment, the point cloud content providing system (e.g., the point cloud sending device 10000 or the point cloud video acquisition unit 10001) can acquire a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system for representing a 3D space. According to the embodiment, the point cloud video may include a Ply (polygon file format or Stanford triangle format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. The geometry includes the position of the point. The position of each point can be represented by a parameter (e.g., the value of the X, Y and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y and Z axes). Attributes include attributes of the point (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc. of each point). A point has one or more attributes. For example, a point may have an attribute as a color or two attributes as color and reflectivity. According to an embodiment, a geometric structure may be referred to as a position, geometric information, geometric data, etc., and an attribute may be referred to as an attribute, attribute information, attribute data, etc. The point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquiring unit 10001) may protect the point cloud data from information related to the acquisition process of the point cloud video (e.g., depth information, color information, etc.).

[0107] According to an embodiment, a point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) may encode point cloud data (20001). The point cloud content providing system may encode point cloud data based on point cloud compression compilation. As described above, point cloud data may include geometric structures and attributes of points. Therefore, the point cloud content providing system may perform geometric encoding for encoding the geometric structure and output a geometric bitstream. The point cloud content providing system may perform attribute encoding for encoding the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system may perform attribute encoding based on geometric encoding. According to an embodiment, the geometric bitstream and the attribute bitstream may be multiplexed and output as one bitstream. According to an embodiment, the bitstream may also include signaling information related to geometric encoding and attribute encoding.

[0108] According to the embodiment, the point cloud content providing system (eg, the transmitting device 10000 or the transmitter 10003) may transmit the encoded point cloud data (20002). Figure 1 As shown in the figure, the encoded point cloud data can be represented by a geometry bitstream and an attribute bitstream. In addition, the encoded point cloud data can be sent in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and attribute encoding). The point cloud content providing system can encapsulate a bitstream carrying the encoded point cloud data and send the bitstream in the form of a file or a segment.

[0109] According to the embodiment, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can receive a bit stream containing encoded point cloud data. In addition, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bit stream.

[0110] The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10006) can decode the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) sent in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10006) can decode the point cloud video data based on the signaling information related to the encoding of the point cloud video data contained in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10006) can decode the geometry bitstream to reconstruct the position (geometry) of the point. The point cloud content providing system can reconstruct the attributes of the point by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10006) can reconstruct the point cloud video based on the position according to the reconstructed geometry and the decoded attributes.

[0111] According to an embodiment, a point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can render decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometric structure and attributes decoded by the decoding process. The points in the point cloud content can be rendered as vertices with a certain thickness, a cube with a certain minimum size centered on the corresponding vertex position, or a circle centered on the corresponding vertex position. All or part of the rendered point cloud content is provided to the user through a display (e.g., a VR / AR display, a common display, etc.).

[0112] The point cloud content providing system (e.g., receiving device 10004) according to the embodiment may protect the feedback information (20005). The point cloud content providing system may encode and / or decode the point cloud data based on the feedback information. Figure 1 The feedback information and operations described are the same, so a detailed description thereof is omitted.

[0113] Figure 3 Illustrated is an exemplary process for capturing point cloud video according to an embodiment.

[0114] Figure 3 Graphic reference Figures 1 to 2 An exemplary point cloud video capture process for a point cloud content providing system is described.

[0115] Point cloud content includes point cloud videos (images and / or videos) representing objects and / or environments located in various 3D spaces (e.g., 3D spaces representing real environments, 3D spaces representing virtual environments, etc.). Therefore, a point cloud content providing system according to an embodiment may use one or more cameras (e.g., an infrared camera capable of protecting depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), a projector (e.g., an infrared pattern projector for protecting depth information), LiDRA, etc. to capture point cloud videos. A point cloud content providing system according to an embodiment may extract the shape of a geometric structure composed of points in a 3D space from the depth information, and extract the attributes of each point from the color information to protect the point cloud data. Images and / or videos according to an embodiment may be captured based on at least one of an inward-facing technology and an outward-facing technology.

[0116] Figure 3The left portion of the figure illustrates an inward-facing technique. An inward-facing technique refers to a technique of capturing an image of a central object with one or more cameras (or camera sensors) arranged around the central object. The inward-facing technique can be used to generate point cloud content that provides a 360-degree image of a key object to a user (e.g., VR / AR content that provides a 360-degree image of an object (e.g., a key object such as a character, player, object, or actor) to a user).

[0117] Figure 3 The right side of the figure illustrates an outward-facing technique. Outward-facing techniques refer to techniques that capture the environment of a central object rather than an image of the central object using one or more cameras (or camera sensors) arranged around the central object. Point cloud content for providing a surrounding environment as it appears from the user's perspective (e.g., content representing the external environment that can be provided to a user of a self-driving vehicle) can be generated using outward-facing techniques.

[0118] As shown in the figure, point cloud content can be generated based on the capture operation of one or more cameras. In this case, the coordinate system is different in each camera, so the point cloud content providing system can calibrate one or more cameras to set the global coordinate system before the capture operation. In addition, the point cloud content providing system can generate point cloud content by synthesizing any image and / or video with the image and / or video captured by the above-mentioned capture technology. The point cloud content providing system may not perform when it generates point cloud content representing a virtual space. Figure 3 The point cloud content providing system according to the embodiment may perform post-processing on the captured image and / or video. In other words, the point cloud content providing system may remove unnecessary areas (e.g., background), identify the space to which the captured image and / or video is connected, and perform an operation of filling the space hole when there is a space hole.

[0119] The point cloud content providing system can generate a piece of point cloud content by performing coordinate transformation on the points of the point cloud video secured from each camera. The point cloud content providing system can perform coordinate transformation on the points based on the coordinates of each camera position. Therefore, the point cloud content providing system can generate content representing a wide range, or can generate point cloud content with high density of points.

[0120] Figure 4 An exemplary point cloud video encoder is illustrated according to an embodiment.

[0121] Figure 4 Show Figure 1An example of a point cloud video encoder 10002. The point cloud video encoder reconstructs and encodes point cloud data (e.g., the location and / or attributes of points) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) according to network conditions or applications. When the total size of the point cloud content is large (e.g., for 30fps, giving 60Gbps of point cloud content), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on the maximum target bit rate to provide the point cloud content according to the network environment, etc.

[0122] As reference Figures 1 to 2 As described above, the point cloud video encoder can perform geometry encoding and attribute encoding. Geometry encoding is performed before attribute encoding.

[0123] The point cloud video encoder according to an embodiment includes a coordinate transformer (transformation coefficient) 40000, a quantizer (quantization and removal of points (voxelization)) 40001, an octree analyzer (analysis of octree) 40002, and a surface approximation analyzer (analysis of surface approximation) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometry reconstructor (reconstruction of geometry) 40005, a color transformer (transformation of color) 40006, an attribute transformer (transformation of attributes) 40007, a RAHT transformer (RAHT) 40008, an LOD generator (generation of LOD) 40009, a lifting transformer (lifting) 40010, a coefficient quantizer (quantization of coefficients) 40011 and / or an arithmetic encoder (arithmetic coding) 40012.

[0124] The coordinate transformer 40000, the quantizer 40001, the octree analyzer 40002, the surface approximation analyzer 40003, the arithmetic encoder 40004 and the geometry reconstructor 40005 may perform geometry coding. The geometry coding according to the embodiment may include octree geometry coding, direct coding, trisoup geometry coding and entropy coding. Direct coding and trisoup geometry coding are applied selectively or in combination. The geometry coding is not limited to the above examples.

[0125] As shown in the figure, the coordinate transformer 40000 according to the embodiment receives the position and transforms it into coordinates. For example, the position can be transformed into position information in a three-dimensional space (e.g., a three-dimensional space represented by an XYZ coordinate system). The position information in the three-dimensional space according to the embodiment can be referred to as geometric information.

[0126] The quantizer 40001 according to the embodiment quantizes the geometric information. For example, the quantizer 40001 can quantize the points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 40001 performs the following quantization operation: the difference between the position value of each point and the minimum position value is multiplied by a preset quantization scaling value, and then the nearest integer value is found by rounding the value obtained by the multiplication. Therefore, one or more points can have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized point. Voxelization means the smallest unit representing position information in 3D space. The point of the point cloud content (or 3D point cloud video) according to the embodiment can be included in one or more voxels. The term voxel, which is a compound word of volume and pixel, refers to a 3D cubic space generated when the 3D space is divided into units (unit = 1.0) based on the axis representing the 3D space (e.g., X axis, Y axis, and Z axis). Quantizer 40001 can match a group of points in 3D space with voxels. According to an embodiment, a voxel may include only one point. According to an embodiment, a voxel may include one or more points. In order to represent a voxel as a point, the position of the center point of the voxel can be set based on the position of one or more points included in the voxel. In this case, the properties of all positions included in a voxel can be combined and assigned to the voxel.

[0127] The octree analyzer 40002 according to an embodiment performs octree geometry compilation (or octree compilation) to present voxels in an octree structure. The octree structure represents points matched with voxels based on the octree structure.

[0128] The surface approximation analyzer 40003 according to the embodiment may analyze and approximate the octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region including a plurality of points to efficiently provide an octree and voxelization.

[0129] The arithmetic encoder 40004 according to an embodiment performs entropy coding on the octree and / or the approximate octree. For example, the coding scheme includes arithmetic coding. As a result of the coding, a geometry bitstream is generated.

[0130] The color converter 40006, the attribute converter 40007, the RAHT converter 40008, the LOD generator 40009, the lifting converter 40010, the coefficient quantizer 40011 and / or the arithmetic encoder 40012 perform attribute coding. As described above, a point may have one or more attributes. The attribute coding according to the embodiment is also applied to the attributes of a point. However, when the attribute (e.g., color) includes one or more elements, the attribute coding is applied independently to each element. The attribute coding according to the embodiment includes color transform coding, attribute transform coding, regional adaptive hierarchical transform (RAHT) coding, hierarchical nearest neighbor prediction (prediction transform) coding based on difference, and hierarchical nearest neighbor prediction coding based on difference with an update / lifting step (lifting transform). According to the point cloud content, the above-mentioned RAHT coding, prediction transform coding and lifting transform coding can be selectively used, or a combination of one or more coding schemes can be used. The attribute coding according to the embodiment is not limited to the above examples.

[0131] The color converter 40006 according to the embodiment performs color conversion compilation of the color value (or texture) included in the conversion attribute. For example, the color converter 40006 can convert the format of the color information (e.g., from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment can be optionally applied according to the color value included in the attribute.

[0132] The geometry reconstructor 40005 according to the embodiment reconstructs (decompresses) the octree and / or the approximate octree. The geometry reconstructor 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of the points. The reconstructed octree / voxel can be referred to as the reconstructed geometry (restored geometry).

[0133] The attribute transformer 40007 according to the embodiment performs attribute transformation to transform the attribute based on the position and / or the reconstructed geometry for which geometric coding is not performed. As described above, since the attribute depends on the geometry, the attribute transformer 40007 can transform the attribute based on the reconstructed geometry information. For example, based on the position value of the point included in the voxel, the attribute transformer 40007 can transform the attribute of the point at the position. As described above, when the position of the voxel center is set based on the position of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of the one or more points. When triplet geometry coding is performed, the attribute transformer 40007 can transform the attribute based on the triplet geometry coding.

[0134] The attribute transformer 40007 can perform attribute transformation by calculating the average of the attributes or attribute values ​​(e.g., the color or reflectivity of each point) of neighboring points within a specific position / radius from the position (or position value) of the center of each voxel. The attribute transformer 40007 can apply weights according to the distance from the center to each point when calculating the average. Therefore, each voxel has a position and a calculated attribute (or attribute value).

[0135] The attribute converter 40007 can search for neighbor points within a specific position / radius from the center of each voxel based on a KD tree or a Morton code. The KD tree is a binary search tree and supports the ability to manage point data structures based on position so that the nearest neighbor search (NNS) can be performed quickly. The Morton code is generated by presenting the coordinates (e.g., (x, y, z)) representing the 3D positions of all points as bit values ​​and mixing the bits. For example, when the coordinates representing the position of the point are (5, 9, 1), the bit values ​​of the coordinates are (0101, 1001, 0001). Mixing the bit values ​​according to the bit index in the order of z, y, and x produces 010001000111. The value is represented as a decimal number 1095. That is, the Morton code value of the point with coordinates (5, 9, 1) is 1095. The attribute converter 40007 can sort the points based on the Morton code value and perform NNS by depth-first traversal processing. After the attribute transformation operation, when NNS is needed in another transformation process for attribute coding, KD tree or Morton code is used.

[0136] As shown in the figure, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.

[0137] The RAHT transformer 40008 according to an embodiment performs RAHT coding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT transformer 40008 may predict attribute information of a higher-level node in the octree based on attribute information associated with a lower-level node in the octree.

[0138] The LOD generator 40009 according to an embodiment generates a level of detail (LOD). The LOD according to an embodiment is the detail level of the point cloud content. As the LOD value decreases, it indicates that the detail level of the point cloud content decreases. As the LOD value increases, it indicates that the detail of the point cloud content increases. Points can be classified by LOD.

[0139] The lifting transformer 40010 according to an embodiment performs lifting transformation compilation that transforms the attributes of the point cloud based on weights. As described above, the lifting transformation compilation can be optionally applied.

[0140] The coefficient quantizer 40011 according to an embodiment quantizes the attribute of the attribute encoding based on the coefficient.

[0141] The arithmetic encoder 40012 according to an embodiment encodes quantized properties based on arithmetic coding.

[0142] Although not shown in this figure, Figure 4 The elements of the point cloud video encoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud content providing device. The one or more processors may perform the above Figure 4 At least one of the operations and / or functions of the elements of the point cloud video encoder. In addition, one or more processors can operate or execute a set of software programs and / or instructions to perform Figure 4 The operation and / or functionality of the elements of the point cloud video encoder. According to an embodiment, one or more memories may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state storage devices).

[0143] Figure 5 An example of a voxel according to an embodiment is shown.

[0144] Figure 5 , which is a voxel located in a 3D space represented by a coordinate system consisting of three axes, namely, an X-axis, a Y-axis, and a Z-axis. Figure 4 As described, the point cloud video encoder (eg, quantizer 40001) may perform voxelization. A voxel refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on axes representing the 3D space (eg, X-axis, Y-axis, and Z-axis). Figure 5 An example of a voxel generated by an octree structure is shown in which a voxel is generated by two poles (0, 0, 0) and (2 d ,2 d ,2 d ) is recursively subdivided. A voxel consists of at least one point. The spatial coordinates of the voxel can be estimated based on the positional relationship with the voxel group. As mentioned above, the voxel has properties like the pixels of a 2D image / video (such as color or reflectivity). The details of the voxel are similar to those of the reference Figure 4 The details described are the same, so their description is omitted.

[0145] Figure 6 An example of an octree and an occupancy code according to an embodiment is shown.

[0146] As reference Figures 1 to 4As described, the point cloud content providing system (point cloud video encoder 10002) or the octree analyzer 40002 of the point cloud video encoder performs octree geometry coding (or octree coding) based on the octree structure to efficiently manage the area and / or position of voxels.

[0147] Figure 6 The upper part of FIG. 1 shows an octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes (eg, X-axis, Y-axis, and Z-axis) of the coordinate system. The octree structure is formed by recursively subdividing the two poles (0, 0, 0) and (2 d ,2 d ,2 d ) is created by defining the cube axis-aligned bounding box. Here, 2 d can be set to the value of the minimum bounding box that forms all points around the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined in Equation 1. In Equation 1, (x int n ,y int n ,z int n ) represents the position (or position value) of the quantization point.

[0148] Equation 1

[0149]

[0150] like Figure 6 As shown in the middle of the upper part, the entire 3D space can be divided into eight spaces according to the partition. Each divided space is represented by a cube with six faces. Figure 6 As shown in the upper right side of , each of the eight spaces is divided again based on the axes of the coordinate system (e.g., X-axis, Y-axis, and Z-axis). Therefore, each space is divided into eight smaller spaces. The divided smaller spaces are also represented by cubes with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.

[0151] Figure 6 The lower part of shows the octree occupancy code. The occupancy code of the octree is generated to indicate whether each of the eight divided spaces generated by dividing one space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the divided space, and the child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when at least one point is contained in the space corresponding to the child node, the node is assigned a value of 1. When no point is contained in the space corresponding to the child node (the space is empty), the node is assigned a value of 0. Since Figure 6The occupancy code shown in is 00100001, so it indicates that the space corresponding to the third child node and the eighth child node among the eight child nodes respectively contains at least one point. As shown in the figure, each of the third child node and the eighth child node has 8 child nodes, and the child nodes are represented by 8-bit occupancy codes. The figure shows that the occupancy code of the third child node is 10000111, and the occupancy code of the eighth child node is 01001111. The point cloud video encoder (e.g., arithmetic encoder 40004) according to the embodiment can perform entropy coding on the occupancy code. In order to improve compression efficiency, the point cloud video encoder can perform intra-frame / inter-frame coding on the occupancy code. The receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to the embodiment reconstructs the occupancy code based on the occupancy code.

[0152] The point cloud video encoder (e.g., octree analyzer 40002) according to an embodiment may perform voxelization and octree compilation to store the positions of points. However, points are not always evenly distributed in 3D space, so there are specific areas where there are fewer points. Therefore, it is inefficient to perform voxelization on the entire 3D space. For example, when a specific area contains fewer points, there is no need to perform voxelization in the specific area.

[0153] Therefore, for the above-mentioned specific area (or nodes other than the leaf nodes of the octree), the point cloud video encoder according to the embodiment can skip voxelization and perform direct coding to directly encode the positions of the points included in the specific area. The coordinates of the direct coding points according to the embodiment are called direct coding mode (DCM). The point cloud video encoder according to the embodiment can also perform triplet geometry coding based on the surface model to reconstruct the positions of the points in the specific area (or node) based on voxels. Triplet geometry coding is a geometry coding that represents an object as a series of triangular meshes. Therefore, the point cloud video decoder can generate a point cloud from the mesh surface. Triplet geometry coding and direct coding according to the embodiment can be selectively performed. In addition, triplet geometry coding and direct coding according to the embodiment can be performed in combination with octree geometry coding (or octree coding).

[0154] In order to perform direct coding, the option of using direct mode to apply direct coding should be enabled. The node to which direct coding is to be applied is not a leaf node, and there should be fewer points than a threshold within a specific node. In addition, the total number of points to which direct coding is to be applied should not exceed a preset threshold. When the above conditions are met, the point cloud video encoder (or arithmetic encoder 40004) according to an embodiment can perform entropy coding on the position (or position value) of the point.

[0155] A point cloud video encoder (e.g., surface approximation analyzer 40003) according to an embodiment may determine a specific level of an octree (a level less than the depth d of the octree), and may perform triplet geometry encoding using a surface model starting from that level to reconstruct the position of a point in the region of a node based on voxels (triplet mode). A point cloud video encoder according to an embodiment may specify a level at which triplet geometry encoding will be applied. For example, when a specific level is equal to the depth of the octree, the point cloud video encoder does not operate in triplet mode. In other words, a point cloud video encoder according to an embodiment may operate in triplet mode only when the specified level is less than the depth value of the octree. A 3D cubic area of ​​a node at a specified level according to an embodiment is referred to as a block. A block may include one or more voxels. A block or voxel may correspond to a patch. The geometric structure is represented as a surface within each block. A surface according to an embodiment may intersect at most once with each edge of a block.

[0156] A block has 12 edges, so there are at least 12 intersections in a block. Each intersection is called a vertex (or top point). When there is at least one occupied voxel adjacent to the edge among all blocks sharing the edge, the vertex existing along the edge is detected. The occupied voxel according to an embodiment refers to the voxel containing the point. The position of the vertex detected along the edge is the average position of the edge of all voxels adjacent to the edge among all blocks sharing the edge.

[0157] Once the vertex is detected, the point cloud video encoder according to an embodiment can perform entropy coding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the vertex position value (relative position value within the edge). When triplet geometry coding is applied, the point cloud video encoder according to an embodiment (e.g., geometry reconstructor 40005) can generate a restored geometry (reconstructed geometry) by performing triangle reconstruction, upsampling, and voxelization.

[0158] Vertices at the edges of a block determine a surface that passes through the block. The surface according to an embodiment is a non-planar polygon. In a triangle reconstruction process, a surface represented by a triangle is reconstructed based on the starting point of the edge, the direction vector of the edge, and the position value of the vertex. According to Equation 2, the triangle reconstruction process is performed by the following operations: i) calculating the centroid value of each vertex, ii) subtracting the center value from each vertex value, and iii) estimating the sum of the squares of the values ​​obtained by the subtraction.

[0159] Equation 2

[0160] C1) ② ③

[0161] Then, the minimum value of the sum is estimated, and the projection process is performed according to the axis with the minimum value. For example, when the element x is the smallest, each vertex is projected onto the x-axis relative to the center of the block and onto the (y, z) plane. When the value obtained by projecting onto the (y, z) plane is (ai, bi), the value of θ is estimated by atan2(bi, ai), and the vertices are sorted according to the value of θ. Table 1 below shows the vertex combination for creating a triangle according to the number of vertices. The vertices are sorted from 1 to n. Table 1 below shows that for four vertices, two triangles can be constructed according to the combination of vertices. The first triangle can be composed of vertices 1, 2, and 3 among the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 among the sorted vertices.

[0162] [Table 1] Triangles formed from vertices ordered 1,…,n

[0163] [Table 1]

[0164]

[0165] An upsampling process is performed to add points in the middle along the sides of the triangle and perform voxelization. The added points are generated based on the upsampling factor and the width of the block. The added points are called refinement vertices. The point cloud video encoder according to an embodiment can voxelize the refinement vertices. In addition, the point cloud video encoder can perform attribute encoding based on the voxelized position (or position value).

[0166] Figure 7 An example of a neighbor node pattern according to an embodiment is illustrated.

[0167] In order to improve the compression efficiency of the point cloud video, the point cloud video encoder according to an embodiment may perform entropy coding based on context adaptive arithmetic coding.

[0168] As reference Figures 1 to 6 Described, Figure 1 Point cloud content providing system or point cloud video encoder 10002 or Figure 4 The point cloud video encoder or arithmetic encoder 40004 can immediately perform entropy coding on the occupancy code. In addition, the point cloud content providing system or the point cloud video encoder can perform entropy coding (intra-frame coding) based on the occupancy code of the current node and the occupancy of the neighboring nodes, or perform entropy coding (inter-frame coding) based on the occupancy code of the previous frame. The frame according to the embodiment represents a collection of point cloud videos generated simultaneously. The compression efficiency of the intra-frame coding / inter-frame coding according to the embodiment may depend on the number of neighboring nodes referenced. When the number of bits increases, the operation becomes complicated, but the coding can be biased to one side, which can increase the compression efficiency. For example, when a 3-bit context is given, 2 bits need to be used. 3= 8 ways to perform compilation. The division into parts for compilation affects the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be met.

[0169] Figure 7 The diagram illustrates a process of obtaining an occupancy pattern based on the occupancy of neighbor nodes. The point cloud video encoder according to an embodiment determines the occupancy of neighbor nodes of each node of the octree and obtains the value of the neighbor pattern. The neighbor node pattern is used to infer the occupancy pattern of the node. Figure 7 The upper part of the diagram shows a cube corresponding to the node (the cube in the middle) and six cubes (neighboring nodes) that share at least one face with the cube. The nodes shown in the diagram are nodes at the same depth. The numbers shown in the diagram represent the weights (1, 2, 4, 8, 16, and 32) associated with the six nodes, respectively. The weights are assigned in sequence according to the positions of the neighboring nodes.

[0170] Figure 7 The lower part shows the neighbor node mode value. The neighbor node mode value is the sum of the values ​​multiplied by the weights of the occupied neighbor nodes (neighbor nodes with points). Therefore, the neighbor node mode value is 0 to 63. When the neighbor node mode value is 0, it indicates that there is no node with a point (unoccupied node) among the neighbor nodes of the node. When the neighbor node mode value is 63, it indicates that all neighbor nodes are occupied nodes. As shown in the figure, since the neighbor nodes assigned weights 1, 2, 4 and 8 are occupied nodes, the neighbor node mode value is 15, which is the sum of 1, 2, 4 and 8. The point cloud video encoder can perform coding according to the neighbor node mode value (for example, when the neighbor node mode value is 63, 64 types of coding can be performed). According to an embodiment, the point cloud video encoder can reduce the coding complexity by changing the neighbor node mode value (for example, based on a table through which 64 is changed to 10 or 6).

[0171] Figure 8 An example of point configuration in each LOD according to an embodiment is illustrated.

[0172] As reference Figures 1 to 7 As described, the encoded geometry is reconstructed (decompressed) before performing attribute encoding. When direct compilation is applied, the geometry reconstruction operation may include changing the placement of directly encoded points (e.g., placing directly encoded points in front of the point cloud data). When triplet geometry encoding is applied, the geometry reconstruction process is performed by triangle reconstruction, upsampling, and voxelization. Since the attributes depend on the geometry, attribute encoding is performed based on the reconstructed geometry.

[0173] The point cloud video encoder (e.g., LOD generator 40009) can classify (reorganize) the points by LOD. The figure shows the point cloud content corresponding to the LOD. The leftmost picture in the figure represents the original point cloud content. The second picture from the left side of the figure represents the distribution of points in the lowest LOD, and the rightmost picture in the figure represents the distribution of points in the highest LOD. That is, the points in the lowest LOD are sparsely distributed, and the points in the highest LOD are densely distributed. That is, as the LOD rises in the direction indicated by the arrow indicated at the bottom of the figure, the space (or distance) between the points becomes narrower.

[0174] Figure 9 An example of point configuration for each LOD according to an embodiment is illustrated.

[0175] As reference Figures 1 to 8 As described, a point cloud content providing system or a point cloud video encoder (e.g., Figure 1 Point cloud video encoder 10002, Figure 4 The point cloud video encoder or LOD generator 40009) can generate LOD. LOD is generated by reorganizing points into a set of refinement levels according to the set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud video encoder, but also by the point cloud video decoder.

[0176] Figure 9 The upper part of shows examples of points (P0 to P9) of point cloud content distributed in 3D space. Figure 9 In , the original order means the order of points P0 to P9 before LOD generation. Figure 9 In , the LOD-based order means the order of points generated according to LOD. Points are reorganized by LOD. In addition, a high LOD contains points belonging to a lower LOD. Figure 9 As shown in , LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.

[0177] As reference Figure 4 As described, the point cloud video encoder according to an embodiment may selectively or in combination perform LOD-based prediction transform coding, LOD-based lifting transform coding, and RAHT transform coding.

[0178] The point cloud video encoder according to an embodiment may generate a predictor for a point to perform LOD-based prediction transform coding to set the prediction attribute (or prediction attribute value) of each point. That is, N predictors may be generated for N points. The predictor according to an embodiment may calculate a weight (=1 / distance) based on the LOD value of each point, index information about neighboring points within a set distance for each LOD, and the distance to the neighboring point.

[0179] The predicted attribute (or attribute value) according to the embodiment is set to the average value obtained by multiplying the attribute (or attribute value) of the neighboring point set in the predictor of each point (e.g., color, reflectivity, etc.) by the weight (or weight value) calculated based on the distance to each neighboring point. The point cloud video encoder (e.g., coefficient quantizer 40011) according to the embodiment can quantize and inverse quantize the residual (which can be referred to as residual attribute, residual attribute value, attribute prediction residual value, or prediction error attribute value, etc.) of each point obtained by subtracting the predicted attribute (or attribute value) of each point from the attribute of each point (i.e., the original attribute value). The quantization processing performed on the residual attribute value in the transmitting device is configured as shown in Table 2. The inverse quantization processing performed on the residual attribute value in the receiving device is configured as shown in Table 3.

[0180] [Table 2]

[0181] int PCCQuantization(int value, int quantStep) { if (value >= 0) { return floor(value / quantStep + 1.0 / 3.0); } else { return -floor(-value / quantStep + 1.0 / 3.0); } }

[0182] [Table 3]

[0183] int PCCInverseQuantization(int value, int quantStep) { if (quantStep == 0) { return value; } else { return value * quantStep; } }

[0184] When the predictor of each point has neighboring points, the point cloud video encoder (e.g., arithmetic encoder 40012) according to an embodiment may perform entropy coding on the quantized and inverse quantized residual values ​​as described above. When the predictor of each point has no neighboring points, the point cloud video encoder (e.g., arithmetic encoder 40012) according to an embodiment may perform entropy coding on the attributes of the corresponding points without performing the above operations.

[0185] A point cloud video encoder (e.g., lifting transformer 40010) according to an embodiment may generate a predictor for each point, set the calculated LOD and register the neighboring points into the predictor, and set weights according to the distance to the neighboring points to perform lifting transform coding. The lifting transform coding according to an embodiment is similar to the above-mentioned prediction transform coding, but the difference is that the weights are cumulatively applied to the attribute values. The processing of cumulatively applying weights to attribute values ​​according to an embodiment is configured as follows.

[0186] 1) Create an array Quantization Weight (QW) for storing the weight value of each point. The initial value of all elements of QW is 1.0. Multiply the QW value of the predictor index of the neighbor node registered in the predictor by the weight of the predictor of the current point, and add the values ​​obtained by the multiplication.

[0187] 2) Boosting prediction processing: A value obtained by multiplying the attribute value of a point by a weight is subtracted from the existing attribute value to calculate a predicted attribute value.

[0188] 3) Create a temporary array called updateweight, and update and initialize it to zero.

[0189] 4) The weight calculated by multiplying the weight calculated for all predictors by the weight corresponding to the predictor index stored in QW is cumulatively added to the update weight array as the index of the neighbor node. The value obtained by multiplying the attribute value of the index of the neighbor node by the calculated weight is cumulatively added to the update array.

[0190] 5) Boosting update process: Divide the attribute value of the update array for all predictors by the weight value of the update weight array indexed by the predictor, and add the existing attribute value to the value obtained by the division.

[0191] 6) The predicted attribute is calculated by multiplying the attribute value updated by the lifting update process by the weight updated by the lifting prediction process (stored in QW) for all predictors. The point cloud video encoder (e.g., coefficient quantizer 40011) according to the embodiment quantizes the predicted attribute value. In addition, the point cloud video encoder (e.g., arithmetic encoder 40012) performs entropy coding on the quantized attribute value.

[0192] A point cloud video encoder according to an embodiment (e.g., RAHT transformer 40008) may perform RAHT transform coding, in which attributes associated with lower-level nodes in an octree are used to predict attributes of higher-level nodes. RAHT transform coding is an example of attribute intra-frame coding performed by backward scanning of an octree. A point cloud video encoder according to an embodiment scans the entire area from voxels, and repeats a merging process of merging voxels into larger blocks in each step until the root node is reached. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed on the upper node immediately above the empty node.

[0193] The following equation 3 represents the RAHT transformation matrix. In equation 3, Represents the average attribute value of the voxels at level l. It can be based on and To calculate and The weight is and

[0194] [Equation 3]

[0195]

[0196] here, is the low-pass value and is used in the merging process at the next higher level. Denotes the high-pass coefficient. The high-pass coefficient in each step is quantized and undergoes entropy coding (e.g., encoded by arithmetic encoder 400012). The weight is calculated as As shown in Equation 4, and Compute the root node.

[0197] [Equation 4]

[0198]

[0199] The values ​​of gDC are also quantized and undergo entropy encoding like the high-pass coefficients.

[0200] Figure 10 A point cloud video decoder according to an embodiment is illustrated.

[0201] Figure 10 The point cloud video decoder shown in the figure is Figure 1 An example of a point cloud video decoder 10006 described in Figure 1 The point cloud video decoder 10006 illustrated in FIG. 10006 may be used to decode the point cloud video content (decoded point cloud) using the decoded geometry structure and the decoded attributes. The point cloud video decoder 10006 may be used to decode the point cloud video content (decoded point cloud) using the decoded geometry structure and the decoded attributes. The point cloud video decoder 10006 may be used to decode the point cloud video content (decoded point cloud) using the decoded geometry structure and the decoded attributes. The point cloud video decoder 10006 may be used to decode the point cloud video content (decoded point cloud) using the decoded geometry structure and the decoded attributes.

[0202] Figure 11 A point cloud video decoder according to an embodiment is illustrated.

[0203] Figure 11 The point cloud video decoder shown in the figure is Figure 10 An example of a point cloud video decoder is shown in FIG. 1 and may be executed as Figures 1 to 9 The decoding operation is the inverse of the encoding operation of the point cloud video encoder shown in FIG.

[0204] As referenceFigure 1 and Figure 10 As described, the point cloud video decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.

[0205] The point cloud video decoder according to an embodiment includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesized octree) 11001, a surface approximation synthesizer (synthesized surface approximation) 11002 and a geometry reconstructor (reconstructed geometry) 11003, an inverse coordinate transformer (inverse transformed coordinates) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantizer (inverse quantization) 11006, a RAHT transformer 11007, an LOD generator (generate LOD) 11008, an inverse lifting (inverse lifting) 11009 and / or an inverse color transformer (inverse transformed color) 11010.

[0206] The arithmetic decoder 11000, the octree synthesizer 11001, the surface approximation synthesizer 11002, the geometry reconstructor 11003, and the coordinate inverse transformer 11004 may perform geometry decoding. The geometry decoding according to the embodiment may include direct decoding and triplet geometry decoding. Direct decoding and triplet geometry decoding are selectively applied. The geometry decoding is not limited to the above examples, and is used as a reference. Figures 1 to 9 This is performed by inverse processing of the geometric encoding described.

[0207] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse process of the arithmetic encoder 40004.

[0208] The octree synthesizer 11001 according to the embodiment can generate an octree by acquiring an occupancy code from a decoded geometry bitstream (or information about a geometry structure protected as a result of decoding). Figures 1 to 9 Configure the seizure code in detail.

[0209] When triplet geometry encoding is applied, a surface approximation synthesizer 11002 according to an embodiment may synthesize a surface based on the decoded geometry and / or the generated octree.

[0210] According to an embodiment, the geometry reconstructor 11003 can regenerate the geometry based on the surface and / or the decoded geometry. Figures 1 to 9 As described, direct coding and triplet geometry coding are selectively applied. Therefore, the geometry reconstructor 11003 directly imports and adds position information about the points to which direct coding is applied. When triplet geometry coding is applied, the geometry reconstructor 11003 can reconstruct the geometry by performing the reconstruction operations (e.g., triangle reconstruction, upsampling, and voxelization) of the geometry reconstructor 40005. Details and ReferencesFigure 6 The details of the description are the same, so the description thereof is omitted. The reconstructed geometry may include a point cloud picture or frame that does not contain attributes.

[0211] The coordinate inverse transformer 11004 according to an embodiment may acquire the position of a point by transforming the coordinates based on the reconstructed geometric structure.

[0212] The arithmetic decoder 11005, the inverse quantizer 11006, the RAHT transformer 11007, the LOD generator 11008, the inverse lifter 11009 and / or the inverse color transformer 11010 may perform the reference Figure 10 Described attribute decoding. Attribute decoding according to an embodiment includes regional adaptive hierarchical transform (RAHT) decoding, hierarchical nearest neighbor prediction (prediction transform) decoding based on difference, and hierarchical nearest neighbor prediction decoding based on difference with an update / lifting step (lifting transform). The above three decoding schemes can be selectively used, or a combination of one or more decoding schemes can be used. Attribute decoding according to an embodiment is not limited to the above examples.

[0213] The arithmetic decoder 11005 according to an embodiment decodes the attribute bitstream through arithmetic coding.

[0214] The inverse quantizer 11006 according to an embodiment inversely quantizes the information about the decoded attribute bitstream or attribute protected as a decoding result, and outputs the inversely quantized attribute (or attribute value). Inverse quantization may be selectively applied based on attribute encoding of the point cloud video encoder.

[0215] According to an embodiment, the RAHT transformer 11007, the LOD generator 11008 and / or the inverse lifter 11009 may process the reconstructed geometry and the inverse quantized properties. As described above, the RAHT transformer 11007, the LOD generator 11008 and / or the inverse lifter 11009 may selectively perform a decoding operation corresponding to the encoding of the point cloud video encoder.

[0216] The color inverse transformer 11010 according to an embodiment performs inverse transform coding to inversely transform the color value (or texture) included in the decoded attribute. The operation of the color inverse transformer 11010 may be selectively performed based on the operation of the color transformer 40006 of the point cloud video encoder.

[0217] Although not shown in this figure, Figure 11 The elements of the point cloud video decoder may be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits configured to communicate with one or more memories included in the point cloud content providing device. The one or more processors may perform the above Figure 11At least one or more of the operations and / or functions of the elements of the point cloud video decoder. In addition, one or more processors can operate or execute a set of software programs and / or instructions to perform Figure 11 The operation and / or functionality of the elements of the point cloud video decoder.

[0218] Figure 12 A transmitting device according to an embodiment is illustrated.

[0219] Figure 12 The sending device shown is Figure 1 The sending device 10000 (or Figure 4 Example of a point cloud video encoder). Figure 12 The sending device shown in the figure can perform the same Figures 1 to 9 One or more of the operations and methods that are the same or similar to the operations and methods of the described point cloud video encoder. The sending device according to the embodiment may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra-frame / inter-frame coding processor 12005, a first arithmetic encoder 12006, a metadata processor 12007, a color transformation processor 12008, an attribute transformation processor 12009, a LOD / lifting / RAHT transformation processor 12010, a second arithmetic encoder 12011 and / or a sending processor 12012.

[0220] The data input unit 12000 according to the embodiment receives or acquires point cloud data. The data input unit 12000 may perform the same operation and / or acquisition method as the point cloud video acquisition unit 10001 (or refer to Figure 2 The same or similar operations and / or acquisition methods as described in the acquisition process 20000).

[0221] The data input unit 12000, the quantization processor 12001, the voxelization processor 12002, the octree occupancy code generator 12003, the surface model processor 12004, the intra / inter encoding processor 12005 and the first arithmetic encoder 12006 perform geometric encoding. Figures 1 to 9 The geometric encoding described is the same or similar, so a detailed description thereof is omitted.

[0222] According to an embodiment, the quantization processor 12001 quantizes the geometric structure (e.g., the position value of a point). The operation and / or quantization of the quantization processor 12001 are similar to the reference Figure 4 The operation and / or quantization of the quantizer 40001 described above is the same or similar. Figures 1 to 9 The details of the description are the same.

[0223] The voxelization processor 12002 according to the embodiment performs voxelization on the quantized position value of the point. The voxelization processor 120002 may perform the same operation as the reference. Figure 4 The operation and / or voxelization process of the quantizer 40001 described above is the same or similar to the operation and / or process described above. Figures 1 to 9 The details of the description are the same.

[0224] The octree occupancy code generator 12003 according to the embodiment performs octree compilation on the voxelized position of the point based on the octree structure. The octree occupancy code generator 12003 can generate an occupancy code. The octree occupancy code generator 12003 can perform the same as the reference Figure 4 and Figure 6 The operations and / or methods of the point cloud video encoder (or octree analyzer 40002) described herein are the same or similar to the operations and / or methods. Figures 1 to 9 The details of the description are the same.

[0225] According to an embodiment, the surface model processor 12004 may perform triplet geometry encoding based on the surface model to reconstruct the position of a point in a specific area (or node) based on voxels. Figure 4 The operations and / or methods described herein are the same or similar to the operations and / or methods of the point cloud video encoder (e.g., surface approximation analyzer 40003). Figures 1 to 9 The details of the description are the same.

[0226] The intra-frame / inter-frame encoding processor 12005 according to the embodiment may perform intra-frame / inter-frame encoding on the point cloud data. The intra-frame / inter-frame encoding processor 12005 may perform the same as the reference Figure 7 Same or similar coding as described for intra / inter coding. Details and references Figure 7 The details of the description are the same. According to an embodiment, the intra / inter encoding processor 12005 may be included in the first arithmetic encoder 12006.

[0227] According to an embodiment, the first arithmetic encoder 12006 performs entropy coding on the octree and / or approximate octree of the point cloud data. For example, the coding scheme includes arithmetic coding. The first arithmetic encoder 12006 performs the same or similar operations and / or methods as the operations and / or methods of the arithmetic encoder 40004.

[0228] The metadata processor 12007 according to the embodiment processes metadata (e.g., set values) about the point cloud data and provides it to necessary processing processes such as geometry coding and / or attribute coding. In addition, the metadata processor 12007 according to the embodiment can generate and / or process signaling information related to geometry coding and / or attribute coding. The signaling information according to the embodiment can be encoded separately from the geometry coding and / or attribute coding. The signaling information according to the embodiment can be interleaved.

[0229] The color transformation processor 12008, the attribute transformation processor 12009, the LOD / lifting / RAHT transformation processor 12010 and the second arithmetic encoder 12011 perform attribute encoding. Figures 1 to 9 The attribute codes described are the same or similar, so a detailed description thereof is omitted.

[0230] The color transform processor 12008 according to the embodiment performs color transform coding to transform the color value included in the attribute. The color transform processor 12008 may perform color transform coding based on the reconstructed geometric structure. The reconstructed geometric structure is consistent with the reference Figures 1 to 9 In addition, it performs the same Figure 4 The operations and / or methods of the color converter 40006 described above are the same as or similar to the operations and / or methods described above, and detailed descriptions thereof are omitted.

[0231] The attribute transformation processor 12009 according to the embodiment performs attribute transformation to transform the attribute based on the reconstructed geometry and / or the location where the geometry encoding is not performed. Figure 4 The operations and / or methods of the attribute transformer 40007 described in the embodiment are the same as or similar to the operations and / or methods of the attribute transformer 40007 described in the embodiment. The detailed description thereof is omitted. The LOD / lifting / RAHT transformation processor 12010 according to the embodiment can encode the transformed attributes through any one of RAHT coding, prediction transformation coding and lifting transformation coding or a combination thereof. The LOD / lifting / RAHT transformation processor 12010 performs the same as the reference Figure 4 The operations of the RAHT transformer 40008, the LOD generator 40009 and the lifting transformer 40010 described above are the same as or similar to at least one of the operations. In addition, the prediction transformation compilation, the lifting transformation compilation and the RAHT transformation compilation are the same as the reference transformation compilation. Figures 1 to 9 Those described are the same, so a detailed description thereof is omitted.

[0232] The second arithmetic encoder 12011 according to the embodiment may encode the encoded attribute based on arithmetic coding. The second arithmetic encoder 12011 performs the same or similar operation and / or method as that of the arithmetic encoder 400012.

[0233] The sending processor 12012 according to an embodiment may send each bitstream containing the coded geometry and / or the coded attributes and metadata information, or send a bitstream configured with the coded geometry and / or the coded attributes and metadata information. When the coded geometry and / or the coded attributes and metadata information according to an embodiment are configured into a bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information, the signaling information including a sequence parameter set (SPS) for sequence level signaling, a geometry parameter set (GPS) for signaling of geometry information coding, an attribute parameter set (APS) for signaling of attribute information coding, and a tile parameter set (TPS or tile library) for tile level signaling and slice data. The slice data may include information about one or more slices. A slice according to an embodiment may include a geometry bitstream Geom0. 0 and one or more attribute bitstreams Attr0 0 and Attr1 0 . The TPS according to an embodiment may include information about each tile of one or more tiles (for example, height / size information and coordinate information about a bounding box). The geometry bitstream may include a header and a payload. The header of the geometry bitstream according to an embodiment may include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, and information about the data contained in the payload. As described above, the metadata processor 12007 according to an embodiment may generate and / or process signaling information and send it to the transmitting processor 12012. According to an embodiment, an element for performing geometry coding and an element for performing attribute coding may share data / information with each other, as indicated by the dotted lines. The transmitting processor 12012 according to an embodiment may perform operations and / or transmitting methods that are the same or similar to those of the transmitter 10003. Details and References Figure 1 and Figure 2 The details described are the same, so their description is omitted.

[0234] Figure 13 A receiving device according to an embodiment is illustrated.

[0235] Figure 13 The receiving device shown in the figure is Figure 1 The receiving device 10004 (or Figure 10 and Figure 11 An example of a point cloud video decoder). Figure 13 The receiving device shown in the figure can perform the same Figures 1 to 11One or more of the operations and methods of the point cloud video decoder described are the same or similar to the operations and methods.

[0236] The receiving device according to the embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an octtree reconstruction processor 13003 based on an occupancy code, a surface model processor (triangle reconstruction, upsampling, voxelization) 13004, a first inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, a second inverse quantization processor 13008, a LOD / lifting / RAHT inverse transform processor 13009, a color inverse transform processor 13010 and / or a renderer 13011. Each element for decoding according to the embodiment may perform an inverse process of the operation of the corresponding element for encoding according to the embodiment.

[0237] The receiver 13000 according to the embodiment receives point cloud data. The receiver 13000 may perform the same Figure 1 The operation and / or receiving method of the receiver 10005 is the same as or similar to the operation and / or receiving method of the receiver 10005. A detailed description thereof is omitted.

[0238] The reception processor 13001 according to the embodiment may obtain a geometry bitstream and / or an attribute bitstream from the received data. The reception processor 13001 may be included in the receiver 13000.

[0239] The arithmetic decoder 13002, the octtree reconstruction processor 13003 based on the occupancy code, the surface model processor 13004 and the first inverse quantization processor 13005 can perform geometric decoding. Figures 1 to 10 The geometric decoding described is the same or similar, so a detailed description thereof is omitted.

[0240] The arithmetic decoder 13002 according to an embodiment may decode the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs the same or similar operations and / or coding as those of the arithmetic decoder 11000.

[0241] The octree reconstruction processor 13003 based on the occupancy code according to the embodiment can reconstruct the octree by obtaining the occupancy code from the decoded geometry bitstream (or the information about the geometry structure protected as a result of decoding). The octree reconstruction processor 13003 based on the occupancy code performs the same or similar operations and / or methods as the operations of the octree synthesizer 11001 and / or the octree generation method. When triplet geometry coding is applied, the surface model processor 1302 according to the embodiment can perform triplet geometry decoding and related geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on the surface model method. The surface model processor 1302 performs the same or similar operations as the operations of the surface approximation synthesizer 11002 and / or the geometry reconstructor 11003.

[0242] The inverse quantization processor 1305 according to an embodiment may inverse quantize the decoded geometry.

[0243] According to an embodiment, the metadata parser 1306 can parse the metadata contained in the received point cloud data, for example, setting values. The metadata parser 1306 can pass the metadata for geometry decoding and / or attribute decoding. Figure 12 The metadata described is the same, so a detailed description thereof is omitted.

[0244] The arithmetic decoder 13007, the second inverse quantization processor 13008, the LOD / lifting / RAHT inverse transform processor 13009 and the color inverse transform processor 13010 perform attribute decoding. Figures 1 to 10 The attribute decoding described is the same or similar, so a detailed description thereof is omitted.

[0245] The arithmetic decoder 13007 according to the embodiment can decode the attribute bitstream by arithmetic coding. The arithmetic decoder 13007 can decode the attribute bitstream based on the reconstructed geometric structure. The arithmetic decoder 13007 performs the same or similar operation and / or coding as the operation and / or coding of the arithmetic decoder 11005.

[0246] The second inverse quantization processor 13008 according to an embodiment may inverse quantize the decoded attribute bitstream. The second inverse quantization processor 13008 performs the same or similar operation and / or inverse quantization method as the inverse quantizer 11006.

[0247] The LOD / lifting / RAHT inverse transform processor 13009 according to an embodiment may process the reconstructed geometry and inverse quantized attributes. The LOD / lifting / RAHT inverse transform processor 1301 performs one or more of the same or similar operations and / or decoding as the RAHT transformer 11007, the LOD generator 11008 and / or the inverse lifter 11009. The color inverse transform processor 13010 according to an embodiment performs inverse transform coding to inversely transform the color value (or texture) included in the decoded attribute. The color inverse transform processor 13010 performs the same or similar operations and / or inverse transform coding as the color inverse transformer 11010. The renderer 13011 according to an embodiment may render point cloud data.

[0248] Figure 14 Illustrated is an architecture for G-PCC based point cloud content streaming according to an embodiment.

[0249] Figure 14 The upper part shows Figures 1 to 13 The sending device described in (for example, the sending device 10000, Figure 12 The processing and sending of point cloud content is performed by the sending device, etc.

[0250] As reference Figures 1 to 13 As described above, the sending device can obtain the audio Ba (audio acquisition) of the point cloud content, encode the obtained audio (audio encoding), and output the audio bitstream Ea. In addition, the sending device can obtain the point cloud (or point cloud video) Bv (point acquisition) of the point cloud content, and perform point cloud video encoding on the obtained point cloud to output the point cloud video bitstream EV. Point cloud video encoding of the sending device is similar to reference Figures 1 to 13 Point cloud video encoding described (e.g., Figure 4 The encoding of the point cloud video encoder is the same or similar, so the detailed description thereof will be omitted.

[0251] The sending device may encapsulate the generated audio bitstream and video bitstream into files and / or segments (file / segment encapsulation). The encapsulated files and / or segments Fs, File may include files in file formats such as ISOBMFF or Dynamic Adaptive Streaming over HTTP (DASH) segments. Point cloud-related metadata according to an embodiment may be included in the encapsulated file format and / or segments. The metadata may be included in boxes at different levels in the ISO International Organization for Standardization Base Media File Format (ISOBMFF) file format, or may be included in separate tracks within the file. According to an embodiment, the sending device may encapsulate the metadata into a separate file. The sending device according to an embodiment may deliver the encapsulated file format and / or segments over a network. The processing method for encapsulation and transmission by the sending device is the same as that of the referenceFigures 1 to 13 (For example, transmitter 10003, Figure 2 The processing method described in sending step 20002, etc. is the same, so its detailed description will be omitted.

[0252] Figure 14 The lower part shows the reference Figures 1 to 13 The receiving device described (eg, receiving device 10004, Figure 13 receiving device, etc.) to process and output point cloud content.

[0253] According to an embodiment, the receiving device may include a device (e.g., a speaker, a headset, a display) configured to output final audio data and final video data and a point cloud player (point cloud player) configured to process point cloud content. The final data output device and the point cloud player may be configured as separate physical devices. The point cloud player according to an embodiment may perform geometry-based point cloud compression (G-PCC) compilation, video-based point cloud compression (V-PCC) compilation, and / or next generation compilation.

[0254] The receiving device according to the embodiment can protect the files and / or segments F', FS' contained in the received data (e.g., broadcast signals, signals sent through a network, etc.), and decapsulate them (file / segment decapsulation). Figures 1 to 13 (e.g., receiver 10005, receiver 13000, reception processor 13001, etc.) are the same as those described, so their description will be omitted.

[0255] The receiving device according to the embodiment protects the audio bitstream E'a and the video bitstream E'v contained in the file and / or segment. As shown in the figure, the receiving device outputs the decoded audio data B'a by performing audio decoding on the audio bitstream, and renders the decoded audio data (audio rendering) to output the final audio data A'a through a speaker or earphone.

[0256] In addition, the receiving device performs point cloud video decoding on the video bit stream E'v and outputs the decoded video data B'v. Figures 1 to 13 Described point cloud video decoding (e.g., Figure 11 The decoding of the point cloud video decoder is the same or similar, so a detailed description thereof will be omitted. The receiving device can render the decoded video data and output the final video data through a display.

[0257] The receiving device according to the embodiment may perform at least one of decapsulation, audio decoding, audio rendering, point cloud video decoding, and point cloud video rendering based on the transmitted metadata. Figures 12 to 13The details of the description are the same, so their description will be omitted.

[0258] As indicated by the dashed lines shown in the figure, a receiving device (e.g., a point cloud player or a sensing / tracking unit in a point cloud player) according to an embodiment may generate feedback information (orientation, viewport). According to an embodiment, the feedback information may be used in the decapsulation process, point cloud video decoding process, and / or rendering process of the receiving device, or may be delivered to the sending device. Details of the feedback information and references Figures 1 to 13 The details of the description are the same, so their description will be omitted.

[0259] Figure 15 An exemplary transmitting device according to an embodiment is shown.

[0260] Figure 15 The sending device is a device configured to send point cloud content and corresponds to the reference Figures 1 to 14 Examples of sending devices described (e.g., Figure 1 The sending device 10000, Figure 4 Point cloud video encoder, Figure 12 The sending device Figure 14 Therefore, Figure 15 The sending device performs the reference Figures 1 to 14 The operation of the sending device is the same or similar to the operation described.

[0261] The transmitting device according to the embodiment may perform one or more of point cloud acquisition, point cloud video encoding, file / segment packaging, and delivery.

[0262] Due to the point cloud acquisition and delivery operations shown in the figure and the reference Figures 1 to 14 The operations described are the same, so a detailed description thereof will be omitted.

[0263] As reference Figures 1 to 14 As described, the sending device according to the embodiment can perform geometry encoding and attribute encoding. Geometry encoding can be called geometry compression, and attribute encoding can be called attribute compression. As described above, a point can have a geometry structure and one or more attributes. Therefore, the sending device performs attribute encoding on each attribute. The figure illustrates that the sending device performs one or more attribute compressions (attribute #1 compression, ..., attribute #N compression). In addition, the sending device according to the embodiment can perform auxiliary compression. Auxiliary compression is performed on metadata. Details of metadata and reference Figures 1 to 14 The details of the description are the same, so a detailed description thereof will be omitted. The sending device may also perform mesh data compression. The mesh data compression according to the embodiment may include reference Figures 1 to 14 Triplet geometry encoding of descriptions.

[0264] According to an embodiment, a transmitting device may encapsulate a bit stream (e.g., a point cloud stream) output according to point cloud video encoding into a file and / or a segment. According to an embodiment, a transmitting device may perform media track encapsulation for carrying data other than metadata (e.g., media data), and perform metadata track encapsulation for carrying metadata. According to an embodiment, metadata may be encapsulated into a media track.

[0265] As reference Figures 1 to 14 As described, the sending device may receive feedback information (orientation / viewport metadata) from the receiving device and perform at least one of point cloud video encoding, file / segment packaging, and delivery operations based on the received feedback information. Figures 1 to 14 The details of the description are the same, so their description will be omitted.

[0266] Figure 16 An exemplary receiving device according to an embodiment is shown.

[0267] Figure 16 The receiving device is a device for receiving point cloud content and corresponds to the reference Figures 1 to 14 Examples of receiving devices described (e.g., Figure 1 The receiving device 10004, Figure 11 Point cloud video decoder and Figure 13 receiving equipment, Figure 14 Therefore, Figure 16 The receiving device performs the reference Figures 1 to 14 The operation of the receiving device is the same or similar to the operation described. Figure 16 The receiving device can receive Figure 15 The sending device sends a signal and executes Figure 15 The reverse process of the operation of the sending device.

[0268] The receiving device according to the embodiment may perform at least one of delivery, file / segment decapsulation, point cloud video decoding, and point cloud rendering.

[0269] Since the point cloud receiving and point cloud rendering operations shown in the figure are similar to the reference Figures 1 to 14 Those described are the same, so a detailed description thereof will be omitted.

[0270] As reference Figures 1 to 14 As described, according to an embodiment, a receiving device decapsulates files and / or segments obtained from a network or a storage device. According to an embodiment, the receiving device may perform media track decapsulation for carrying data other than metadata (e.g., media data), and perform metadata track decapsulation for carrying metadata. According to an embodiment, in the case where metadata is encapsulated in a media track, metadata track decapsulation is omitted.

[0271] As referenceFigures 1 to 14 As described, the receiving device can perform geometry decoding and attribute decoding on a bit stream (e.g., a point cloud stream) protected by decapsulation. Geometry decoding can be referred to as geometry decompression, and attribute decoding can be referred to as attribute decompression. As described above, a point can have a geometry structure and one or more attributes correspondingly encoded by a transmitting device. Therefore, the receiving device performs attribute decoding on each attribute. The figure illustrates that the receiving device performs one or more attribute decompressions (attribute #1 decompression, ..., attribute #N decompression). The receiving device according to the embodiment can also perform auxiliary decompression. Auxiliary decompression is performed on metadata. Details of metadata and reference Figures 1 to 14 The details of the description are the same, so the description thereof will be omitted. The receiving device may also perform mesh data decompression. The mesh data decompression according to the embodiment may include reference Figures 1 to 14 The receiving device according to the embodiment may render the point cloud data outputted from the point cloud video decoding.

[0272] As reference Figures 1 to 14 As described, the receiving device may use a separate sensing / tracking element to protect the orientation / viewport metadata and send feedback information including the same to the sending device (e.g., Figure 15 In addition, the receiving device may perform at least one of a receiving operation, file / segment decapsulation, and point cloud video decoding based on the feedback information. Figures 1 to 14 The details of the description are the same, so their description will be omitted.

[0273] Figure 17 An exemplary structure operably connectable with a method / apparatus for transmitting and receiving point cloud data according to an embodiment is shown.

[0274] Figure 17 The structure of represents a configuration in which at least one of the server 1760, the robot 1710, the self-driving vehicle 1720, the XR / PCC device 1730, the smart phone 1740, the home appliance 1750, and / or the head mounted display (HMD) 1770 is connected to the cloud network 1700. The robot 1710, the self-driving vehicle 1720, the XR / PCC device 1730, the smart phone 1740, or the home appliance 1750 is referred to as a device. In addition, the XR / PCC device 1730 may correspond to a point cloud compressed data (PCC) device according to an embodiment, or may be operably connected to a PCC device.

[0275] The cloud network 1700 may represent a network that constitutes part of a cloud computing infrastructure or exists in a cloud computing infrastructure. Here, the cloud network 1700 may be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0276] The server 1760 may be connected to at least one of the robot 1710 , the self-driving vehicle 1720 , the XR / PCC device 1730 , the smartphone 1740 , the home appliance 1750 , and / or the HMD 1770 via the cloud network 1700 , and may assist at least a portion of the connected devices 1710 to 1770 .

[0277] The HMD 1770 represents one of implementation types of an XR device and / or a PCC device according to an embodiment. The HMD type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0278] Hereinafter, various embodiments of devices 1710 to 1750 to which the above-mentioned technology is applied will be described. According to the above-mentioned embodiments, Figure 17 The devices 1710 to 1750 illustrated in FIG. 1 may be operably connected / coupled to a point cloud data sending device and a receiving device.

[0279] <PCC + XR>

[0280] The XR / PCC device 1730 can adopt PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) set in a vehicle, a TV, a mobile phone, a smart phone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a fixed robot or a mobile robot.

[0281] The XR / PCC device 1730 can analyze 3D point cloud data or image data obtained through various sensors or from external devices, and generate position data and attribute data about 3D points. Thus, the XR / PCC device 1730 can obtain information about the surrounding space or real objects, and render and output XR objects. For example, the XR / PCC device 1730 can match an XR object including auxiliary information about the identified object with the identified object, and output the matched XR object.

[0282] <PCC + Self-driving + XR>

[0283] The self-driving vehicle 1720 can be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0284] The self-driving vehicle 1720 to which the XR / PCC technology is applied may represent a self-driving vehicle provided with a device for providing an XR image or a self-driving vehicle as a control / interaction target in the XR image. Specifically, the self-driving vehicle 1720 as a control / interaction target in the XR image may be distinguished from the XR / PCC device 1730 and may be operatively connected to the XR / PCC device 1730.

[0285] The self-driving vehicle 1720 having a device for providing an XR / PCC image may acquire sensor information from a sensor including a camera and output an XR / PCC image generated based on the acquired sensor information. For example, the self-driving vehicle 1720 may have a HUD and output an XR / PCC image thereto, thereby providing an occupant with an XR / PCC object corresponding to a real object or an object present on a screen.

[0286] When the XR / PCC object is output to the HUD, at least a portion of the XR / PCC object may be output to overlap with a real object that the occupant's eyes are looking at. On the other hand, when the XR / PCC object is output to a display provided inside the self-driving vehicle, at least a portion of the XR / PCC object may be output to overlap with an object on the screen. For example, the self-driving vehicle 1720 may output an XR / PCC object corresponding to an object such as a road, another vehicle, a traffic light, a traffic sign, a two-wheeled vehicle, a pedestrian, and a building.

[0287] The virtual reality (VR) technology, augmented reality (AR) technology, mixed reality (MR) technology and / or point cloud compression (PCC) technology according to the embodiments are applicable to various devices.

[0288] In other words, VR technology is a display technology that only provides CG images of real-world objects, backgrounds, etc. On the other hand, AR technology refers to a technology that shows a virtually created CG image on an image of a real object. MR technology is similar to the above-mentioned AR technology in that the virtual objects to be shown are mixed and combined with the real world. However, MR technology is different from AR technology in that AR technology clearly distinguishes between real objects and virtual objects created as CG images and uses virtual objects as supplementary objects to real objects, while MR technology regards virtual objects as objects with equivalent characteristics to real objects. More specifically, an example of the application of MR technology is a hologram service.

[0289] Recently, VR, AR, and MR technologies are sometimes referred to as extended reality (XR) technologies without being clearly distinguished from each other. Therefore, the embodiments of the present disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are applicable to such technologies.

[0290] The PCC method / apparatus according to the embodiment may be applied to a vehicle providing a self-driving service.

[0291] Vehicles providing self-driving services are connected to the PCC device for wired / wireless communication.

[0292] When a point cloud compressed data (PCC) transmitting and receiving device according to an embodiment is connected to a vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services that can be provided together with self-driving services, and send it to the vehicle. In the case where the PCC transmitting and receiving device is installed on the vehicle, the PCC transmitting and receiving device can receive / process content data related to AR / VR / PCC services based on a user input signal input through a user interface device, and provide it to the user. The vehicle or user interface device according to the embodiment can receive a user input signal. The user input signal according to the embodiment may include a signal indicating a self-driving service.

[0293] At the same time, the point cloud video encoder on the sending side can further perform a spatial partitioning process of spatially partitioning the point cloud data into one or more 3D blocks before encoding the point cloud data. That is, in order to perform the encoding and transmission operations of the sending device and the decoding and rendering operations of the receiving device in real time and process with low latency, the sending device can spatially partition the point cloud data into multiple regions. In addition, the sending device can encode the spatially partitioned regions (or blocks) independently or non-independently, thereby achieving random access and parallel encoding in the three-dimensional space occupied by the point cloud data. In addition, the sending device and the receiving device can independently or non-independently perform encoding and decoding on each spatially partitioned region (or block), thereby preventing error accumulation during the encoding and decoding process.

[0294] Figure 18 is a diagram illustrating another example of a point cloud transmitting device including a space partitioner according to an embodiment.

[0295] The point cloud transmission device according to an embodiment may include a space partitioner 14001, a signaling processor 14002, a geometry encoder 14003, an attribute encoder 14004, and a transmission processor 14005. According to an embodiment, the space partitioner 14001, the geometry encoder 14003, and the attribute encoder 14004 may be referred to as a point cloud video encoder.

[0296] That is, the spatial partitioner 14001 can partition the input point cloud data space into one or more 3D blocks based on the bounding box and / or sub-bounding box. Here, a 3D block can refer to a tile group, a tile, a slice, a coding unit (CU), a prediction unit (PU) or a transform unit (TU). In one embodiment, the signaling information for spatial partitioning is entropy encoded by the signaling processor 14002 and then sent via the sending processor 14005 in the form of a bitstream.

[0297] Figure 19 (a) to Figure 19 (c) illustrates an embodiment of partitioning a bounding box into one or more tiles. Figure 19 As shown in (a), the point cloud object corresponding to the point cloud data can be expressed in the form of a box based on the coordinate system, which is called a bounding box. In other words, the bounding box represents a cube that can contain all the points of the point cloud.

[0298] Figure 19 (b) and Figure 19 (c) An example diagram, where Figure 19 The bounding box of (a) is partitioned into tile 1# and tile 2#, and tile 2# is partitioned again into slice 1# and slice 2#.

[0299] A tile may represent a partial area of ​​a 3D space occupied by point cloud data according to an embodiment. According to an embodiment, a tile may include one or more slices. A tile according to an embodiment may be partitioned into one or more slices, and thus the point cloud video encoder may encode the point cloud data in parallel.

[0300] A slice may mean a data unit encoded by a point cloud video encoder according to an embodiment and / or a data unit decoded by a point cloud video decoder according to an embodiment. A slice may be a collection of data in a 3D space occupied by point cloud data, or a collection of some data among point cloud data. According to an embodiment, a slice may represent a collection of areas or points included in a tile according to an embodiment. According to an embodiment, a tile may be partitioned into one or more slices based on the number of points included in a tile. For example, a tile may be a collection of points partitioned by the number of points. According to an embodiment, a tile may be partitioned into one or more slices based on the number of points, and some data may be split or merged during the partitioning process. That is, a slice may be a unit that can be independently compiled within a corresponding tile.

[0301] The point cloud video encoder according to the embodiment may encode the point cloud data by slice or by a tile including one or more slices. In addition, the point cloud video encoder according to the embodiment may perform different quantization and / or transformation on each tile or each slice.

[0302] The positions of one or more 3D blocks spatially partitioned by the spatial partitioner 14001 are output to the geometry encoder 14003, and attribute information (or attributes) are output to the attribute encoder 14004. These positions may be position information about points included in the partition unit (frame or block), and are referred to as geometry information.

[0303] The geometry encoder 14003 constructs and encodes an octree based on the position output from the spatial partitioner 14001 to output a geometry bitstream. In addition, the geometry encoder 14003 may reconstruct an octree and / or an approximate octree and output it to the attribute encoder 14004. The reconstructed octree may be referred to as reconstructed geometry (or restored geometry).

[0304] The attribute encoder 14004 encodes the attributes output from the space partitioner 14001 based on the reconstructed geometry output from the geometry encoder 14003 and outputs an attribute bitstream.

[0305] Geometry Encoder 14003 can execute Figure 4 Some or all of the operations of the coordinate transformer 40000, the quantizer 40001, the octree analyzer 40002, the surface approximation analyzer 40003, the arithmetic encoder 40004 and the geometric reconstruction 40005, or may be performed Figure 12 The quantization processor 12001, the voxelization processor 12002, the octree occupancy code generator 12003, the surface model processor 12004, the intra / inter coding processor 12005 and the first arithmetic encoder 12006 may be used for some or all operations.

[0306] Attribute encoder 14004 can execute Figure 4 The color converter 40006, the attribute converter 40007, the RAHT converter 40008, the LOD generator 40009, the lifting converter 40010 and the coefficient quantizer 40011 and the arithmetic encoder 40012 are performed, or some or all of the operations are performed. Figure 12 Some or all operations of the color transform processor 12008, the attribute transform processor 12009, the LOD / lifting / RAHT transform processor 12010 and the second arithmetic encoder 12011.

[0307] The signaling processor 14002 may generate and / or process signaling information and output it to the transmission processor 14005 in the form of a bitstream. The signaling information generated and / or processed by the signaling processor 14002 may be provided to the geometry encoder 14003, the attribute encoder 14004, and the transmission processor 14005 for geometry encoding, attribute encoding, and transmission processing. Alternatively, the signaling processor 14002 may receive signaling information generated by the geometry encoder 14003, the attribute encoder 14004, and the transmission processor 14005. In the present specification, the signaling information may be signaled and transmitted by parameter sets (sequence parameter set (SPS), geometry parameter set (GPS), attribute parameter set (APS), tile parameter set (TPS) (also referred to as tile library), etc.). The signaling information may be signaled and transmitted in units of coding units of each image such as slices or tiles. In the present specification, the signaling information may include metadata (e.g., setting values, etc.) about the point cloud data, and may be provided to the geometry encoder 14003, the attribute encoder 14004, and / or the transmission processor 14005 for geometry encoding, attribute encoding, and transmission processing. Depending on the application, the signaling information may also be defined at a system end such as a file format, Dynamic Adaptive Streaming over HTTP (DASH), and MPEG Media Transport (MMT), or a wired interface end such as High Definition Multimedia Interface (HDMI), DisplayPort, Video Electronics Standards Association (VESA), and CTA.

[0308] The method / device according to the embodiment may send relevant information with a signal to add / perform the operation of the embodiment. The signaling information according to the embodiment may be used in a sending device and / or a receiving device.

[0309] The sending processor 14005 can execute Figure 12 The operations and / or sending methods of the sending processor 12012 are the same or similar to the operations and / or sending methods, and can be performed with Figure 1 The operation and / or transmission method of the transmitter 1003 is the same as or similar to the operation and / or transmission method of the transmitter 1003. The detailed description will be omitted and reference will be made to Figure 1 or Figure 12 Description.

[0310] The transmission processor 14005 can transmit the geometry bit stream output from the geometry encoder 14003, the attribute bit stream output from the attribute encoder 14004, and the signaling bit stream output from the signaling processor 14002, or can multiplex them into one bit stream and transmit the bit stream.

[0311] Figure 20 is a diagram illustrating another exemplary point cloud receiving device according to an embodiment.

[0312] The point cloud receiving device according to the embodiment may include a receiving processor 15001, a signaling processor 15002, a geometry decoder 15003, an attribute decoder 15004, and a post-processor 15005. According to the embodiment, the geometry decoder 15003 and the attribute decoder 15004 may be collectively referred to as a point cloud video decoder. According to the embodiment, the point cloud video decoder may be referred to as a PCC decoder, a PCC decoding unit, a point cloud video decoder, a point cloud video decoding unit, etc.

[0313] The receiving processor 15001 according to the embodiment may receive a single bit stream, or may receive a geometry bit stream, an attribute bit stream, and a signaling bit stream respectively. After receiving the single bit stream, the receiving processor 15001 demultiplexes the geometry bit stream, the attribute bit stream, and the signaling bit stream from the single bit stream. Then, the demultiplexed signaling bit stream is output to the signaling processor 15002, the geometry bit stream is output to the geometry decoder 15003, and the attribute bit stream is output to the attribute decoder 15004. After receiving each of the geometry bit stream, the attribute bit stream, and the signaling bit stream, the receiving processor 15001 may deliver the signaling bit stream to the signaling processor 15002, deliver the geometry bit stream to the geometry decoder 15003, and deliver the attribute bit stream to the attribute decoder 15004.

[0314] The signaling processor 15002 can parse and process signaling information from the input signaling bit stream, for example, information contained in SPS, GPS, APS, TPS, metadata, etc., and provide it to the geometry decoder 15003, the attribute decoder 15004, and the post-processor 15005. That is, when the point cloud data is partitioned into tiles and / or slices at the sending side, such as Fig.19 As shown in , the TPS includes the number of slices included in each tile, and thus the point cloud video decoder according to an embodiment can check the number of slices and quickly parse the information for parallel decoding.

[0315] Therefore, the point cloud video decoder according to the present disclosure can quickly parse the bitstream containing point cloud data when it receives the SPS with a reduced amount of data. The receiving device can decode the tile after receiving the tile, and can decode each slice based on the GPS and APS included in each tile. Thus, the decoding efficiency can be maximized.

[0316] That is, the geometry decoder 15003 can perform Fig.18The geometry is reconstructed by performing the inverse process of the operation of the geometry encoder 14003. The geometry recovered (or reconstructed) by the geometry decoder 15003 is provided to the attribute decoder 15004. The attribute decoder 15004 can perform the attribute decoder 15004 on the input attribute bitstream based on the signaling information (e.g., attribute related parameters) and the reconstructed geometry. Fig.18 The attributes are restored by performing the inverse process of the attribute encoder 14004. Fig.19 When the sending side shown in is partitioned into tiles and / or slices, the geometry decoder 15003 and the attribute decoder 15004 perform geometry decoding and attribute decoding on a tile-by-tile and / or slice-by-slice basis.

[0317] According to the embodiment, the geometry decoder 15003 may perform Fig.11 Some or all of the operations of the arithmetic decoder 11000, the octree synthesizer 11001, the surface approximation synthesizer 11002, the geometry reconstructor 11003 and the inverse coordinate transformer 11004 may be performed, or Fig.13 Some or all operations of the arithmetic decoder 13002, the octtree reconstruction processor 13003 based on the occupancy code, the surface model processor 13004 and the first inverse quantization processor 13005.

[0318] According to an embodiment, the attribute decoder 15004 may perform Fig.11 Some or all of the operations of the arithmetic decoder 11005, the inverse quantizer 11006, the RAHT transformer 11007, the LOD generator 11008, the inverse lifter 11009 and the color inverse transformer 11010 may be performed, or Fig.13 Some or all of the operations of the arithmetic decoder 13007, the second inverse quantization processor 13008, the LOD / lifting / RAHT inverse transform processor 13009 and the color inverse transform processor 13010.

[0319] The post-processor 15005 can reconstruct the point cloud data by matching the restored geometric structure with the restored attributes. In addition, when the reconstructed point cloud data is in units of tiles and / or slices, the post-processor 15005 can perform the inverse process of spatial partitioning on the sending side based on the signaling information. Fig.19 When the bounding box shown in (a) is partitioned into tiles and slices, as Fig.19 (b) and Fig.19 As shown in (c), tiles and / or slices may be combined based on signaling information to restore Fig.19 (b) The bounding box shown.

[0320] Fig.21 An exemplary bitstream structure for transmitting / receiving point cloud data according to an embodiment is shown.

[0321] When the geometry bitstream, attribute bitstream and signaling bitstream according to the embodiment are configured as one bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to the embodiment may include a sequence parameter set (SPS) for sequence level signaling, a geometry parameter set (GPS) for signaling of geometry information coding, one or more attribute parameter sets (APS) (APS0, APS1) for signaling of attribute information coding, a tile parameter set (TPS) (or tile library) for tile level signaling, and one or more slices (slice 0 to slice n). That is, the bitstream of the point cloud data according to the embodiment may include one or more tiles, and each tile may be a group of slices including one or more slices (slice 0 to slice n). The TPS according to the embodiment may contain information about each of the one or more tiles (for example, coordinate value information and height / size information about the bounding box). Each slice may include a geometry bitstream (Geom0) and one or more attribute bitstreams (Attr0 and Attr1). For example, the first slice (slice 0) may include a geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 、Attr1 0 ).

[0322] The geometry bitstream in each slice may be composed of a geometry slice header (geom_slice_header) and geometry slice data (geom_slice_data). According to an embodiment, the geom_slice_header may include identification information (geom_parameter_set_id) for a parameter set included in the GPS, a tile identifier (geom_tile_id) and a slice identifier (geom_slice_id), and information (geomBoxOrigin, geom_box_log2_scale, geom_max_node_size_log2, geom_num_points) about the data contained in the geometry slice data (geom_slice_data). geomBoxOrigin is geometry box origin information indicating the origin of the box of the geometry slice data, geom_box_log2_scale is information indicating the logarithmic scale of the geometry slice data, geom_max_node_size_log2 is information indicating the root geometry octree node size, and geom_num_points is information related to the number of points of the geometry slice data. According to an embodiment, geom_slice_data may include geometric information (or geometric data) about point cloud data in a corresponding slice.

[0323] Each attribute bitstream in each slice may be composed of an attribute slice header (attr_slice_header) and attribute slice data (attr_slice_data). According to an embodiment, attr_slice_header may include information about the corresponding attribute slice data. The attribute slice data may contain attribute information (or attribute data) about the point cloud data in the corresponding slice. When there are multiple attribute bitstreams in a slice, each bitstream may contain different attribute information. For example, one attribute bitstream may contain attribute information corresponding to color, while another attribute stream may contain attribute information corresponding to reflectivity.

[0324] Fig. 22 (a) and Fig. 22 (b) shows an exemplary bitstream structure for point cloud data and a connection relationship between components in a bitstream of point cloud data according to an embodiment. Fig. 22 (a) and Fig. 22 The bitstream structure of the point cloud data shown in (b) can be expressed as Fig.21 The bitstream structure of the point cloud data is shown.

[0325] According to an embodiment, the SPS may include an identifier (seq_parameter_set_id) for identifying the SPS, and the GPS may include an identifier (geom_parameter_set_id) for identifying the GPS and an identifier (seq_parameter_set_id) indicating the active SPS to which the GPS belongs. The APS may include an identifier (attr_parameter_set_id) for identifying the APS and an identifier (seq_parameter_set_id) indicating the active SPS to which the APS belongs. According to an embodiment, the geometry data may include a geometry slice header and geometry slice data. The geometry slice header may include an identifier (geom_parameter_set_id) of the active GPS to be referenced by the corresponding geometry slice. The geometry slice header may further include an identifier (geom_slice_id) for identifying the corresponding geometry slice and / or an identifier (geom_tile_id) for identifying the corresponding tile. The geometry slice data may include a geometry bitstream belonging to the corresponding slice. According to an embodiment, the attribute data may include an attribute slice header and attribute slice data. The attribute slice header may include an identifier (attr_parameter_set_id) of an active APS to be referenced by a corresponding attribute slice and an identifier (geom_slice_id) for identifying a geometry slice related to the attribute slice. The attribute slice data may include an attribute bitstream belonging to a corresponding slice.

[0326] That is, the geometry slice refers to the GPS, and the GPS refers to the SPS. In addition, the SPS lists the available attributes, assigns identifiers to each attribute, and identifies the decoding method. The attribute slices are mapped to the output attributes according to the identifiers. The attribute slices have dependencies on the aforementioned (decoded) geometry slices and APS. The APS refers to the SPS.

[0327] According to an embodiment, parameters required for encoding of point cloud data may be newly defined in a parameter set of point cloud data and / or a corresponding slice header. For example, when encoding attribute information, parameters may be added to the APS. When tile-based encoding is performed, parameters may be added to the tile and / or slice header.

[0328] As in Fig.21 , Fig. 22 (a) and Fig. 22 As shown in (b), the bitstream of the point cloud data provides tiles or slices so that the point cloud data can be partitioned and processed by region. According to an embodiment, the corresponding regions of the bitstream may have different importance. Therefore, when the point cloud data is partitioned into tiles, different filters (encoding methods) and different filter units may be applied to each tile. When the point cloud data is partitioned into slices, different filters and different filter units may be applied to each slice.

[0329] When point cloud data is partitioned and compressed, the transmitting device and the receiving device according to the embodiment may transmit and receive a bit stream in a high-level syntax structure for selectively transmitting attribute information in the partitioned regions.

[0330] The sending device according to the embodiment can Fig.21 , Fig. 22 (a) and Fig. 22 The bitstream structure shown in (b) transmits point cloud data. Therefore, a method for applying different encoding operations to important areas and using a good quality encoding method can be provided. In addition, efficient encoding and transmission can be supported according to the characteristics of point cloud data, and attribute values ​​can be provided according to user needs.

[0331] The receiving device according to the embodiment can Fig.21 , Fig. 22 (a) and Fig. 22 The bitstream structure shown in (b) receives point cloud data. Therefore, different filtering (decoding) methods can be applied to corresponding areas (areas partitioned into tiles or slices) instead of applying complex decoding (filtering) methods to the entire point cloud data. Therefore, better image quality in areas important to users and appropriate latency for the system can be ensured.

[0332] A field, as a term used in the syntax of the present disclosure described below, may have the same meaning as a parameter or an element.

[0333] Fig.23 An embodiment of a syntax structure of a sequence parameter set (SPS) (seq_parameter_set()) according to the present disclosure is shown. The SPS may contain sequence information about a point cloud data bitstream.

[0334] The SPS according to an embodiment may include a profile_compatibility_flag field, a level_idc field, a sps_bounding_box_present_flag field, a sps_source_scale_factor field, a sps_seq_parameter_set_id field, a sps_num_attribute_sets field, and a sps_extension_present_flag field.

[0335] The profile_compatibility_flag field having a value equal to 1 may indicate that the bitstream conforms to the profile.

[0336] The level_idc field indicates the level to which the bitstream conforms.

[0337] The sps_bounding_box_present_flag field indicates whether source bounding box information is signaled in the SPS. The source bounding box information may include offset and size information about the source bounding box. For example, an sps_bounding_box_present_flag field equal to 1 indicates that source bounding box information is signaled in the SPS. An sps_bounding_box_present_flag field equal to 0 indicates that source bounding box information is not signaled.

[0338] The sps_source_scale_factor field indicates the scaling factor of the source point cloud.

[0339] The sps_seq_parameter_set_id field provides an identifier of the SPS for reference by other syntax elements.

[0340] The sps_num_attribute_sets field indicates the number of attributes coded in the bitstream.

[0341] The sps_extension_present_flag field specifies whether the sps_extension_data syntax structure is present in the SPS syntax structure. For example, a sps_extension_present_flag field equal to 1 specifies that the sps_extension_data syntax structure is present in the SPS syntax structure. A sps_extension_present_flag field equal to 0 specifies that the syntax structure is not present. When not present, the value of the sps_extension_present_flag field is inferred to be equal to 0.

[0342] When the sps_bounding_box_present_flag field is equal to 1, the SPS according to an embodiment may further include a sps_bounding_box_offset_x field, a sps_bounding_box_offset_y field, a sps_bounding_box_offset_z field, a sps_bounding_box_scale_factor field, a sps_bounding_box_size_width field, a sps_bounding_box_size_height field, and a sps_bounding_box_size_depth field.

[0343] The sps_bounding_box_offset_x field indicates the x offset of the source bounding box in Cartesian coordinates. When the x offset of the source bounding box does not exist, the value of sps_bounding_box_offset_x is 0.

[0344] The sps_bounding_box_offset_y field indicates the y offset of the source bounding box in Cartesian coordinates. When the y offset of the source bounding box does not exist, the value of sps_bounding_box_offset_y is 0.

[0345] The sps_bounding_box_offset_z field indicates the z offset of the source bounding box in Cartesian coordinates. When the z offset of the source bounding box does not exist, the value of sps_bounding_box_offset_z is 0.

[0346] The sps_bounding_box_scale_factor field indicates the scaling factor of the source bounding box in Cartesian coordinates. When the scaling factor of the source bounding box does not exist, the value of sps_bounding_box_scale_factor may be 1.

[0347] The sps_bounding_box_size_width field indicates the width of the source bounding box in Cartesian coordinates. When the width of the source bounding box does not exist, the value of the sps_bounding_box_size_width field may be 1.

[0348] The sps_bounding_box_size_height field indicates the height of the source bounding box in Cartesian coordinates. When the height of the source bounding box does not exist, the value of the sps_bounding_box_size_height field may be 1.

[0349] The sps_bounding_box_size_depth field indicates the depth of the source bounding box in Cartesian coordinates. When the depth of the source bounding box does not exist, the value of the sps_bounding_box_size_depth field may be 1.

[0350] The SPS according to an embodiment includes an iteration statement that repeats as many times as the value of the sps_num_attribute_sets field. In an embodiment, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is repeated until the value of i becomes equal to the value of the sps_num_attribute_sets field. The iteration statement may include an attribute_dimension[i] field, an attribute_instance_id[i] field, an attribute_bitdepth[i] field, an attribute_cicp_colour_primaries[i] field, an attribute_cicp_transfer_characteristics[i] field, an attribute_cicp_matrix_coeffs[i] field, an attribute_cicp_video_full_range_flag[i] field, and a known_attribute_label_flag[i] field.

[0351] The attribute_dimension[i] field specifies the number of components of the i-th attribute.

[0352] The attribute_instance_id[i] field specifies the instance ID of the i-th attribute.

[0353] The attribute_bitdepth[i] field specifies the bit depth of the i-th attribute signal.

[0354] The attribute_cicp_colour_primaries[i] field indicates the chromaticity coordinates of the colour attribute source primaries of the i-th attribute.

[0355] The attribute_cicp_transfer_characteristics[i] field indicates a reference electro-optical transfer characteristic function of a color attribute as a function of source input linear light intensity having a nominal real value range of 0 to 1, or indicates the inverse of a reference electro-optical transfer characteristic as a function of output linear light intensity.

[0356] The attribute_cicp_matrix_coeffs[i] field describes the matrix coefficients used to derive the luma and chroma signals from the Green, Blue, and Red or Y, Z, and X primary colors.

[0357] The attribute_cicp_video_full_range_flag[i] field indicates the black level and range of luma and chroma signals derived from the E'Y, E'PB, and E'PR or E'R, E'G, and E'B real-valued component signals.

[0358] The known_attribute_label[i] field specifies whether the known_attribute_label field or the attribute_label_four_bytes field is signaled for the i-th attribute. For example, a value of the known_attribute_label_flag[i] field equal to 1 specifies that the known_attribute_label field is signaled for the i-th attribute. A known_attribute_label_flag[i] field equal to 1 specifies that the attribute_label_four_bytes field is signaled for the i-th attribute.

[0359] The known_attribute_label[i] field may specify the attribute type. For example, a known_attribute_label[i] field equal to 0 may specify that the i-th attribute is color. A known_attribute_label[i] field equal to 1 may specify that the i-th attribute is reflectivity. A known_attribute_label[i] field equal to 2 may specify that the i-th attribute is a frame index.

[0360] The attribute_label_four_bytes field indicates a known attribute type using a 4-byte code.

[0361] Fig.24A table listing exemplary attribute types assigned to the attribute_label_four_bytes field is shown.

[0362] In this example, the attribute_label_four_bytes field indicates color when equal to 0 and indicates reflectivity when equal to 1.

[0363] According to an embodiment, when the sps_extension_present_flag field is equal to 1, the SPS may further include a sps_extension_data_flag field.

[0364] The sps_extension_data_flag field can have any value.

[0365] Fig.25 An embodiment of a syntax structure of a tile parameter set (TPS) (tile_parameter_set()) according to the present disclosure is shown. According to an embodiment, the TPS may be referred to as a tile library.

[0366] TPS includes a num_tiles field.

[0367] The num_tiles field indicates the number of tiles signaled for the corresponding attribute.

[0368] The TPS according to an embodiment includes an iteration statement with the same number of iterations as the value of the num_tiles field. In an embodiment, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of i becomes equal to the value of the num_tiles field. The iteration statement may include a tile_bounding_box_offset_x[i] field, a tile_bounding_box_offset_y[i] field, a tile_bounding_box_offset_z[i] field, a tile_bounding_box_size_width[i] field, a tile_bounding_box_size_height[i] field, and a tile_size_bounding_box field.

[0369] The tile_bounding_box_offset_x[i] field indicates the x offset of the i-th tile in Cartesian coordinates.

[0370] The tile_bounding_box_offset_y[i] field indicates the y offset of the i-th tile in Cartesian coordinates.

[0371] The tile_bounding_box_offset_z[i] field indicates the z offset of the i-th tile in Cartesian coordinates.

[0372] The tile_bounding_box_size_width[i] field indicates the width of the i-th tile in Cartesian coordinates.

[0373] The tile_bounding_box_size_height[i] field indicates the height of the i-th tile in Cartesian coordinates.

[0374] The tile_bounding_box_size_depth[i] field indicates the depth of the i-th tile in Cartesian coordinates.

[0375] Fig.26 An embodiment of a syntax structure of a geometry parameter set (GPS) (geometry_parameter_set()) according to the present disclosure is shown. The GPS according to the embodiment may include information on a method of encoding geometric information on point cloud data included in one or more slices.

[0376] According to an embodiment, the GPS may include a gps_geom_parameter_set_id field, a gps_seq_parameter_set_id field, a gps_box_present_flag field, a unique_geometry_points_flag field, a neighbor_context_restriction_flag field, an inferred_direct_coding_mode_enabled_flag field, a bitwise_occupancy_coding_flag field, an adjacent_child_contextualization_enabled_flag field, a log2_neighbour_avail_boundary field, a log2_intra_pred_max_node_size field, a log2_trisoup_node_size field, and a gps_extension_present_flag field.

[0377] The gps_geom_parameter_set_id field provides an identifier for the GPS for reference by other syntax elements.

[0378] The gps_seq_parameter_set_id field specifies the value of the sps_seq_parameter_set_id of the active SPS.

[0379] The gps_box_present_flag field specifies whether additional bounding box information is provided in the geometry slice header with reference to the current GPS. For example, a gps_box_present_flag field equal to 1 may specify that additional bounding box information is provided in the geometry header with reference to the current GPS. Accordingly, when the gps_box_present_flag field is equal to 1, the GPS may further include a gps_gsh_box_log2_scale_present_flag field.

[0380] The gps_gsh_box_log2_scale_present_flag field specifies whether the gps_gsh_box_log2_scale field is signaled in each geometry slice header with reference to the current GPS. For example, a gps_gsh_box_log2_scale_present_flag field equal to 1 may specify that the gps_gsh_box_log2_scale field is signaled in each geometry slice header with reference to the current GPS. As another example, a gps_gsh_box_log2_scale_present_flag field equal to 0 may specify that the gps_gsh_box_log2_scale field is not signaled in each geometry slice header and a common scaling of all slices is signaled in the gps_gsh_box_log2_scale field of the current GPS.

[0381] When the gps_gsh_box_log2_scale_present_flag field is equal to 0, GPS may further include a gps_gsh_box_log2_scale field.

[0382] The gps_gsh_box_log2_scale field indicates the common scaling factor of the bounding box origin of all slices referenced to the current GPS.

[0383] The unique_geometry_points_flag field indicates whether all output points have unique positions. For example, a unique_geometry_points_flag field equal to 1 indicates that all output points have unique positions. A unique_geometry_points_flag field equal to 0 indicates that two or more output points may have the same position.

[0384] The neighbor_context_restriction_flag field indicates the context used for octree occupancy coding. For example, a neighbor_context_restriction_flag field equal to 0 indicates that the octree occupancy coding uses the context determined from the six neighboring parent nodes. A neighbor_context_restriction_flag field equal to 1 indicates that the octree occupancy coding uses the context determined only from the sibling nodes.

[0385] The inferred_direct_coding_mode_enabled_flag field indicates whether the direct_mode_flag field is present in the geometry node syntax. For example, an inferred_direct_coding_mode_enabled_flag field equal to 1 indicates that the direct_mode_flag field may be present in the geometry node syntax. For example, an inferred_direct_coding_mode_enabled_flag field equal to 0 indicates that the direct_mode_flag field is not present in the geometry node syntax.

[0386] The bitwise_occupancy_coding_flag field indicates whether the geometry node occupancy is encoded using the bitwise contextualization of the syntax element occupancy map. For example, a bitwise_occupancy_coding_flag field equal to 1 indicates that the geometry node occupancy is encoded using the bitwise contextualization of the syntax element ocupancy_map. For example, a bitwise_occupancy_coding_flag field equal to 0 indicates that the geometry node occupancy is encoded using the dictionary coded syntax element occupancy_byte.

[0387] The adjacent_child_contextualization_enabled_flag field indicates whether the adjacent child nodes of the adjacent octree node are used for bit-by-bit occupancy contextualization. For example, an adjacent_child_contextualization_enabled_flag field equal to 1 indicates that the adjacent child nodes of the adjacent octree node are used for bit-by-bit occupancy contextualization. For example, an adjacent_child_contextualization_enabled_flag field equal to 0 indicates that the child nodes of the adjacent octree node are not used for occupancy contextualization.

[0388] The log2_neighbour_avail_boundary field specifies the value of the variable NeighbAvailBoundary used in the decoding process as follows:

[0389] NeighbAvailBoundary=2log2_neighbour_avail_boundary

[0390] For example, when the neighbor_context_restriction_flag field is equal to 1, the NeighbAvailabilityMask may be set equal to 1. For example, when the neighbor_context_restriction_flag field is equal to 0, the NeighbAvailabilityMask may be set equal to 1. <log2_neighbour_avail_boundary。

[0391] The log2_intra_pred_max_node_size field specifies the octree node size suitable for occupied intra prediction.

[0392] The log2_trisoup_node_size field specifies the variable TrisoupNodeSize as the size of the triangle nodes as follows.

[0393] TrisoupNodeSize=1< <log2_trisoup_node_size

[0394] The gps_extension_present_flag field specifies whether the gps_extension_data syntax structure is present in the GPS syntax structure. For example, gps_extension_present_flag equal to 1 specifies that the gps_extension_data syntax structure is present in the GPS syntax. For example, gps_extension_present_flag equal to 0 specifies that this syntax structure is not present in the GPS syntax.

[0395] When the value of the gps_extension_present_flag field is equal to 1, the GPS according to the embodiment may further include a gps_extension_data_flag field.

[0396] The gps_extension_data_flag field can have any value. Its presence and value does not affect the conformance of the decoder to the profile.

[0397] Fig. 27 An embodiment of a syntax structure of an attribute parameter set (APS) (attribute_parameter_set()) according to the present disclosure is shown. The APS according to the embodiment may include information on a method of encoding attribute information on point cloud data included in one or more slices.

[0398] The APS according to an embodiment may include an aps_attr_parameter_set_id field, an aps_seq_parameter_set_id field, an attr_coding_type field, an aps_attr_initial_qp field, an aps_attr_chroma_qp_offset field, an aps_slice_qp_delta_present_flag field, and an aps_present field.

[0399] The aps_attr_parameter_set_id field provides an identifier for the APS for reference by other syntax elements.

[0400] The aps_seq_parameter_set_id field specifies the value of sps_seq_parameter_set_id of the active SPS.

[0401] The attr_coding_type field indicates the coding type of the attribute.

[0402] Fig.28 is a table showing exemplary attribute coding types assigned to the attr_coding_type field.

[0403] In this example, an attr_coding_type field equal to 0 indicates prediction weight boosting as the coding type. An attr_coding_type field equal to 1 indicates RAHT as the coding type. An attr_coding_type field equal to 2 indicates fixed weight boosting.

[0404] The aps_attr_initial_qp field specifies the initial value of the variable SliceQp for each slice that references the APS. When a non-zero value of slice_qp_delta_luma or slice_qp_delta_luma is decoded, the initial value of SliceQp is modified at the attribute slice segment layer.

[0405] The aps_attr_chroma_qp_offset field specifies the offset to the initial quantization parameter signaled by the syntax aps_attr_initial_qp.

[0406] The aps_slice_qp_delta_present_flag field specifies whether the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements are present in the attribute slice header (ASH). For example, an aps_slice_qp_delta_present_flag field equal to 1 specifies that the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements are present in the ASH. For example, the aps_slice_qp_delta_present_flag field specifies that the ash_attr_qp_delta_luma and ash_attr_qp_delta_chroma syntax elements are not present in the ASH.

[0407] When the value of the attr_coding_type field is 0 or 2, that is, the coding type is prediction weight boosting or fixed weight boosting, the APS according to an embodiment may further include a lifting_num_pred_nearest_neighbors field, a lifting_search_range_minus1 field, a lifting_num_detail_levels_minus1 field, and a lifting_neighbour_bias[k] field.

[0408] The lifting_num_pred_nearest_neighbours field specifies the maximum number of nearest neighbors to be used for prediction.

[0409] The lifting_max_num_direct_predictors field specifies the maximum number of predictors used for direct prediction. The value of the variable MaxNumPredictors used in the decoding process can be expressed as follows:

[0410] MaxNumPredictors=lifting_max_num_direct_predicots field+1

[0411] The lifting_lifting_search field specifies the search range used to determine the nearest neighbors to be used for prediction and building distance-based levels of detail.

[0412] The lifting_lod_regular_sampling_enabled_flag field specifies whether the level of detail (LOD) is constructed by a regular sampling strategy. For example, a lifting_lod_regular_sampling_enabled_flag equal to 1 specifies that the level of detail (LOD) is constructed by using a regular sampling strategy. A lifting_lod_regular_sampling_enabled_flag equal to 0 specifies that a distance-based sampling strategy is used instead.

[0413] The lifting_num_detail_levels_minus1 field specifies the number of detail levels used for attribute compilation.

[0414] The APS according to an embodiment includes repeating an iteration statement as many times as the value of the lifting_num_detail_levels_minus1 field. In an embodiment, the index (idx) is initialized to 0 and incremented by 1 each time the iteration statement is executed, and the iteration statement is repeated until the index (idx) is greater than the value of the lifting_num_detail_levels_minus1 field. When the value of the lifting_lod_decimation_enabled_flag field is true (e.g., 1), this iteration statement may include a lifting_sampling_period[idx] field, and when the value of the lifting_lod_decimation_enabled_flag field is false (e.g., 0), a lifting_sampling_distance_squared[idx] field may be included.

[0415] The lifting_sampling_period[idx] field specifies the sampling period for level of detail idx.

[0416] The lifting_sampling_distance_squared[idx] field specifies the square of the sampling distance for level of detail idx.

[0417] When the value of the attr_coding_type field is 0, that is, the coding type is prediction weight lifting, the APS according to an embodiment may further include a lifting_adaptive_prediction_threshold field and a lifting_intra_lod_prediction_num_layers field.

[0418] The lifting_adaptive_prediction_threshold field specifies the threshold to enable adaptive prediction.

[0419] The lifting_intra_lod_prediction_num_layers field specifies the number of LOD layers in which the prediction value of the target point can be generated with reference to the decoded points in the same LOD layer. For example, a lifting_intra_lod_prediction_num_layers field equal to num_detail_levels_minus1 plus 1 indicates that for all LOD layers, the target point can refer to the decoded points in the same LOD layer. For example, a lifting_intra_lod_prediction_num_layers field equal to 0 indicates that for any LOD layer, the target point cannot refer to the decoded points in the same LoD layer.

[0420] The aps_extension_present_flag field specifies whether the aps_extension_data syntax structure is present in the APS syntax structure. For example, an aps_extension_present_flag field equal to 1 specifies that the aps_extension_data syntax structure is present in the APS syntax structure. For example, an aps_extension_present_flag field equal to 0 specifies that this syntax structure is not present in the APS syntax structure.

[0421] When a value of the aps_extension_present_flag field is 1, the APS according to an embodiment may further include an aps_extension_data_flag field.

[0422] The aps_extension_data_flag field can have any value. Its presence and value does not affect the decoder's compliance with the profile.

[0423] Fig.29 An embodiment of the syntax structure of a geometry slice bitstream() according to the present disclosure is shown.

[0424] The geometry slice bitstream (geometry_slice_bitstream()) according to an embodiment may include a geometry slice header (geometry_slice_header()) and geometry slice data (geometry_slice_data()). The geometry slice bitstream may be referred to as a geometry slice. In addition, the attribute slice bitstream may be referred to as an attribute slice.

[0425] Fig.30 An embodiment of a syntax structure of a geometry slice header (geometry_slice_header()) according to the present disclosure is shown.

[0426] According to an embodiment, a bit stream sent by a transmitting device (or a bit stream received by a receiving device) may include one or more slices. Each slice may include a geometry slice and an attribute slice. The geometry slice includes a geometry slice header (GSH). The attribute slice includes an attribute slice header (ASH).

[0427] A geometry slice header (geometry_slice_header()) according to an embodiment may include a gsh_geom_parameter_set_id field, a gsh_tile_id field, a gsh_slice_id field, a gsh_max_node_size_log2 field, a gsh_num_points field, and a byte_alignment() field.

[0428] When the value of the gps_box_present_flag field included in GPS is "true" (e.g., 1), and the value of the gps_gsh_box_log2_scale_present_flag field is "true" (e.g., 1), the geometry slice header (geometry_slice_header()) according to an embodiment may further include a gsh_box_log2_scale field, a gsh_box_origin_x field, a gsh_box_origin_y field, and a gsh_box_origin_z field.

[0429] The gsh_geom_parameter_set_id field specifies the value of the gps_geom_parameter_set_id of the active GPS.

[0430] The gsh_tile_id field specifies the value of the tile id referenced by GSH.

[0431] gsh_slice_id specifies the slice id for reference by other syntax elements.

[0432] The gsh_box_log2_scale field specifies the scaling factor of the slice's bounding box origin.

[0433] The gsh_box_origin_x field specifies the x value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0434] The gsh_box_origin_y field specifies the y value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0435] The gsh_box_origin_z field specifies the z value of the bounding box origin scaled by the value of the gsh_box_log2_scale field.

[0436] gsh_max_node_size_log2 specifies the size of the root geometry octree node.

[0437] gbh_points_number specifies the number of points compiled within the corresponding slice.

[0438] Fig.31 An embodiment of a syntax structure of geometry slice data (geometry_slice_data()) according to the present disclosure is shown. The geometry slice data (geometry_slice_data()) according to an embodiment may carry a geometry bitstream belonging to a corresponding slice.

[0439] According to an embodiment, geometry_slice_data() may include a first iteration statement that repeats as many times as the value of MaxGeometryOctreeDepth. In an embodiment, the depth is initialized to 0 and incremented by 1 each time the iteration statement is executed, and the first iteration statement is repeated until the depth becomes equal to MaxGeometryOctreeDepth. The first iteration statement may include a second loop statement that repeats as many times as the value of NumNodesAtDepth. In an embodiment, nodeidx is initialized to 0 and incremented by 1 each time the iteration statement is executed. Repeat the second iteration statement until nodeidx becomes equal to NumNodesAtDepth. The second iteration statement may include xN=NodeX[depth][nodeIdx], yN=NodeY[depth][nodeIdx], zN=NodeZ[depth][nodeIdx], and geometry_node(depth,nodeIdx,xN,yN,zN). MaxGeometryOctreeDepth indicates the maximum value of the depth of the geometry octree, and NumNodesAtDepth indicates the number of nodes to be decoded at the corresponding depth. The variables NodeX[depth][nodeIdx], NodeY[depth][nodeIdx], and NodeZ[depth][nodeIdx] indicate the x, y, z coordinates of the Idxth node in decoding order at a given depth. The geometry bitstream for a depth node is sent via geometry_node(depth, nodeIdx, xN, yN, zN).

[0440] When the value of the log2_trisoup_node_size field is greater than 0, the geometry slice data (geometry_slice_data()) according to an embodiment may further include geometry_trisoup_data(). That is, when the size of the triangle node is greater than 0, the geometry bitstream subjected to triplet geometry encoding is transmitted through geometry_trisoup_data().

[0441] Fig.32 An embodiment of a syntax structure of attribute_slice_bitstream() according to the present disclosure is shown.

[0442] An attribute slice bitstream (attribute_slice_bitstream()) according to an embodiment may include an attribute slice header (attribute_slice_header()) and attribute slice data (attribute_slice_data()).

[0443] Fig.33 An embodiment of a syntax structure of an attribute slice header (attribute_slice_header()) according to the present disclosure is shown.

[0444] The attribute slice header (attribute_slice_header()) according to an embodiment may include an ash_attr_parameter_set_id field, an ash_attr_sps_attr_idx field, and an ash_attr_geom_slice_id field.

[0445] When a value of the aps_slice_qp_delta_present_flag field of the APS is 'true' (for example, 1), the attribute slice header (attribute_slice_header()) according to an embodiment may further include an ash_qp_delta_luma field and an ash_qp_delta_chroma field.

[0446] The ash_attr_parameter_set_id field specifies the value of the aps_attr_parameter_set_id field of the currently active APS.

[0447] The ash_attr_sps_attr_idx field specifies the attribute to be set in the currently active SPS.

[0448] The ash_attr_geom_slice_id field specifies the value of the gsh_slice_id field of the current geometry slice header.

[0449] The ash_qp_delta_luma field specifies the luma delta quantization parameter (QP) derived from the initial slice QP in the active attribute parameter set.

[0450] The ash_qp_delta_chroma field specifies the chroma delta qp derived from the initial slice qp in the active attribute parameter set.

[0451] Fig.34 An embodiment of a syntax structure of attribute slice data (attribute_slice_data()) according to the present disclosure is shown. The attribute slice data (attribute_slice_data()) according to the embodiment may carry an attribute bitstream belonging to a corresponding slice.

[0452] exist Fig.34In the zerorun field, the zerorun field specifies the number of zeros before predIndex or residual.

[0453] In addition, a predIndex[i] field specifies a predictor index for decoding a value of an i-th point of an attribute. The value of the predIndex[i] field ranges from 0 to the value of the max_num_predictors field.

[0454] As described above, the bit stream of the point cloud data output from the transmitting processor 14005 may include SPS, GPS, one or more APS, a tile library, and one or more slices. The one or more slices may include a geometry slice, one or more attribute slices, and one or more metadata slices. The geometry slice according to the embodiment is composed of a geometry slice header and geometry slice data, and each attribute slice includes an attribute slice header and attribute slice data. Each metadata slice includes a metadata slice header and metadata slice data. For example, in Fig.18 In the point cloud sending device, the geometric slice structure, the attribute slice structure and the metadata slice structure can be generated by the geometric encoder 14003, the attribute encoder 14004 and the signaling processor 14002 respectively, can be generated by the sending processor 14005, or can be generated using separate modules / components.

[0455] In an embodiment of the present disclosure, the bit stream of the point cloud data may be configured in a G-PCC bit stream structure. The G-PCC bit stream may be sent to the receiving side as is, or may be composed of Fig.14 or Fig.15 The file / fragment encapsulator encapsulates the file / fragment into a file / fragment and sends it to the receiving end.

[0456] Depending on the embodiment, the G-PCC bitstream structure may be generated by the transmit processor 14005 or a separate module / component.

[0457] Fig.35An example of a G-PCC bitstream structure according to an embodiment is shown. The G-PCC bitstream consists of one or more G-PCC units. That is, the G-PCC bitstream is a collection of G-PCC units. Each G-PCC unit consists of a G-PCC unit header and a G-PCC unit payload. In the present disclosure, the data contained in the G-PCC unit payload is distinguished by the G-PCC unit header. To this end, the G-PCC unit header contains type information indicating the G-PCC unit type. The G-PCC unit payload of the G-PCC unit contains SPS, GPS, one or more APS, TPS, geometric slices, one or more attribute slices, and one or more metadata slices. According to an embodiment, according to the type information, each G-PCC unit payload may contain one of SPS, GPS, one or more APS, TPS, geometric slices, one or more attribute slices, and one or more metadata slices.

[0458] The syntax structure of SPS and the detailed information and references included in SPS Fig.23 The syntax structure of TPS and the detailed information contained in TPS are the same as those in reference Fig.25 The syntax structure of GPS and the detailed information included in GPS are the same as those in reference 1. Fig.26 The syntax structure of APS and the detailed information included in APS are the same as those in reference 1. Fig. 27 The descriptions are the same, and thus a detailed description thereof will be omitted.

[0459] Details and references of geometric slices Figure 29 to Figure 31 The details of the description are the same as those of the reference 1 and thus the description thereof will be omitted. Figure 32 to Figure 34 The descriptions are the same, and therefore the descriptions thereof will be omitted.

[0460] The receiver may use the metadata to decode the geometry or attribute slices or to render the reconstructed point cloud.According to an embodiment, the metadata may be included in the G-PCC bitstream.

[0461] For example, when the point cloud is Figure 2 or Fig.14When the viewing orientation (or viewpoint) shown in has different color values, the metadata can be a viewing orientation (or viewpoint) associated with information about each color among the attribute values ​​of the point cloud. For example, when the color of the points constituting the point cloud displayed when viewed from (0,0,0) is rendered as different from the color displayed when viewed from (0,90,0), there may be two colors associated with each point. In addition, in order to render appropriate color information according to the user's viewing orientation (or viewpoint) in the rendering operation, it is necessary to send the viewing orientation (or viewpoint) associated with the corresponding color information. To this end, each metadata slice may contain one or more viewing orientations (or viewpoints), and may contain information about the slice containing attribute information associated therewith. Thus, the player can find the associated attribute slice based on the information contained in the appropriate metadata slice according to the user's viewing orientation (or viewpoint), decode it, and perform rendering based on the decoding result. Therefore, attribute values ​​according to the user's viewing orientation (viewpoint) can be presented and provided.

[0462] Fig.36 An embodiment of a syntax structure of metadata_slice_bitstream() according to the present disclosure is shown.

[0463] The metadata slice bitstream (metadata_slice_bitstream()) according to an embodiment may include a metadata slice header (metadata_slice_header()) and metadata slice data (metadata_slice_data()).

[0464] Fig.37 An embodiment of a syntax structure of a metadata slice header (metadata_slice_header()) according to the present disclosure is shown.

[0465] The metadata slice header (metadata_slice_header()) according to an embodiment may include an msh_slice_id field, an msh_geom_slice_id field, and a num_metadata field.

[0466] The msh_slice_id field specifies an identifier for identifying a metadata slice bitstream.

[0467] The msh_geom_slice_id field specifies an identifier used to identify the geometry slice associated with the metadata carried in the metadata slice.

[0468] The num_metadata field indicates the number of metadata included in the metadata slice bitstream.

[0469] According to an embodiment, metadata_slice_header() may further include an iteration statement, the number of iterations of which is as many as the value of the num_metadata field. In an embodiment, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of i becomes equal to the value of the num_metadata field. The iteration statement may include an msh_type field, an msh_attr_id field, and an msh_attr_slice_id field.

[0470] The msh_type field indicates the type of the i-th metadata.

[0471] The msh_attr_id field specifies an identifier for identifying an attribute associated with the i-th metadata carried in the metadata slice.

[0472] The msh_attr_slice_id field specifies an identifier for identifying an i-th metadata-related attribute slice carried in a metadata slice.

[0473] Fig.38 An embodiment of a syntax structure of metadata slice data (metadata_slice_data()) according to the present disclosure is shown.

[0474] The metadata slice data (metadata_slice_data()) according to an embodiment may include an iteration statement, the number of iterations of which is as many as the value of the num_metadata field in the metadata slice header. In an embodiment, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of i becomes equal to the value of the num_metadata field. The iteration statement includes a metadata bitstream (metadata_bitstream()).

[0475] Fig.39 An exemplary syntax structure of each G-PCC unit according to an embodiment is shown. Each G-PCC unit consists of a G-PCC unit header and a G-PCC unit payload.

[0476] Fig.40 An exemplary syntax structure of a G-PCC unit header according to an embodiment is shown. In an embodiment, Fig.40 The G-PCC unit header (gpcc_unit_header()) includes a gpcc_unit_type field. The gpcc_unit_type field indicates a G-PCC unit type or a data type contained in a G-PCC unit payload.

[0477] Fig.41An example of a G-PCC unit type allocated to the gpcc_unit_type field according to an embodiment is shown.

[0478] refer to Fig.41 According to an embodiment, a gpcc_unit_type field equal to 0 indicates that the data contained in the G-PCC unit payload of the G-PCC unit is a sequence parameter set (GPCC_SPS), and a gpcc_unit_type field equal to 1 indicates that the data is a geometry parameter set (GPCC_GPS). A gpcc_unit_type field equal to 2 indicates that the data is an attribute parameter set (GPCC_APS). A gpcc_unit_type field equal to 3 indicates that the data is a tile parameter set (GPCC_TPS). A gpcc_unit_type field equal to 4 indicates that the data is a geometry slice (GPCC_GS). A gpcc_unit_type field equal to 5 indicates that the data is an attribute slice (GPCC_AS). A gpcc_unit_type field equal to 6 indicates that the data is a metadata slice (GPCC_MS). A geometry slice according to an embodiment contains geometry data decoded independently of another slice. An attribute slice according to an embodiment contains attribute data decoded independently of another slice. A metadata slice according to an embodiment contains metadata decoded independently of another slice.

[0479] Those skilled in the art can easily change the meaning, order, deletion, addition, etc. of the values ​​assigned to the gpcc_unit_type field, and therefore the present invention should not be limited to the above-described embodiments.

[0480] The G-PCC unit payload conforms to the format of the HEVC NAL unit.

[0481] Fig.42 An exemplary syntax structure of a G-PCC unit payload (gpcc_unit_payload()) according to an embodiment is shown.

[0482] Fig.42 The G-PCC unit payload includes one of a slice sequence parameter set (SPS), a geometry parameter set (GPS), an attribute parameter set (APS), a tile parameter set (TPS), a geometry slice bitstream, an attribute slice bitstream, and a metadata slice bitstream according to the value of the gpcc_unit_type field in the G-PCC unit header.

[0483] when Fig.41When the value of the gpcc_unit_type field in the G-PCC unit header indicates a sequence parameter set (GPCC_SPS), the G-PCC unit payload (gpcc_unit_payload()) may contain a sequence parameter set (seq_parameter_set()). For detailed information contained in seq_parameter_set(), refer to Fig.23 Description.

[0484] When the value of the gpcc_unit_type field indicates a geometry parameter set (GPCC_GPS), the G-PCC unit payload (gpcc_unit_payload()) may contain a geometry parameter set (geometry_parameter_set()). For detailed information contained in geometry_parameter_set(), refer to Fig.26 Description.

[0485] When the value of the gpcc_unit_type field indicates an attribute parameter set (GPCC_APS), the G-PCC unit payload (gpcc_unit_payload()) may contain an attribute parameter set (attribute_parameter_set()). For detailed information contained in attribute_parameter_set(), refer to Fig. 27 Description.

[0486] When the value of the gpcc_unit_type field indicates a tile parameter set (GPCC_TPS), the G-PCC unit payload (gpcc_unit_payload()) may contain a tile parameter set (tile_parameter_set()). For detailed information contained in tile_parameter_set(), refer to Fig.25 Description.

[0487] When the value of the gpcc_unit_type field indicates geometry slice (GPCC_GS), the G-PCC unit payload (gpcc_unit_payload()) may contain a geometry slice bitstream (geometry_slice_bitstream()). For detailed information contained in geometry_slice_bitstream(), refer to Figure 29 to Figure 31 Description.

[0488] When the value of the gpcc_unit_type field indicates an attribute slice (GPCC_AS), the G-PCC unit payload (gpcc_unit_payload()) may contain an attribute slice bitstream (attribute_slice_bitstream()). For detailed information contained in attribute_slice_bitstream(), refer to Figure 36 to Figure 38 Description.

[0489] When the value of the gpcc_unit_type field indicates a metadata slice (GPCC_MS), the G-PCC unit payload (gpcc_unit_payload()) may contain a metadata slice bitstream (metadata_slice_bitstream()). For detailed information contained in metadata_slice_bitstream(), refer to Figure 36 to Figure 38 Description.

[0490] One or more G-PCC units may be encapsulated into a G-PCC access unit.

[0491] Fig.43 An exemplary structure of a G-PCC access unit according to an embodiment is shown. The G-PCC access unit consists of a G-PCC access unit header and one or more G-PCC units. According to an embodiment, the G-PCC unit included in the G-PCC access unit may be a G-PCC unit that needs to be decoded and rendered simultaneously. That is, the G-PCC access unit consists of a G-PCC access unit header and one or more G-PCC units, which include SPS, GPS, APS or TPS.

[0492] Fig.44 An exemplary syntax structure of a G-PCC access unit header according to an embodiment is shown.

[0493] The G-PCC access unit header (gpcc_access_unit_header()) according to an embodiment may include a num_gpcc_units field and an attribute_included field.

[0494] The num_gpcc_units field indicates the number of G-PCC units included in the G-PCC access unit.

[0495] The attribute_included field indicates whether the G-PCC access unit includes an attribute slice.

[0496] When the attribute_included field indicates that an attribute slice is included in the G-PCC access unit, the G-PCC access unit header (gpcc_access_unit_header()) may further include a num_attribute_sets field.

[0497] The num_attribute_sets field indicates the number of attribute slices included in the G-PCC access unit.

[0498] The G-PCC access unit header (gpcc_access_unit_header()) may further include an iteration statement, the number of iterations of which is as many as the value of the num_attribute_sets field. In an embodiment, i is initialized to 0 and is incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of i becomes equal to the value of the num_attribute_sets field. The iteration statement may include an attribute_type[i] field, an attribute_dimension[i] field, and an attribute_instance_id[i] field.

[0499] The attribute_type[i] field indicates the attribute type of the i-th attribute slice in the G-PCC access unit.

[0500] The attribute_dimension[i] field indicates the attribute dimension of the attribute type of the i-th attribute slice in the G-PCC access unit.

[0501] The attribute_instance_id[i] field indicates the instance identifier of the i-th attribute slice in the G-PCC access unit.

[0502] Fig.45 An exemplary syntax structure of a G-PCC access unit payload according to an embodiment is shown.

[0503] According to an embodiment, a G-PCC access unit payload (gpcc_access_unit_payload()) may include an iteration statement, the number of iterations of which is as many as the value of the num_gpcc_units field in the G-PCC access unit header. In an embodiment, i is initialized to 0 and incremented by 1 each time the iteration statement is executed. The iteration statement is iterated until the value of i becomes equal to the value of the num_gpcc_units field. The iteration statement includes a G-PCC unit (gpcc_unit()). For details of gpcc_unit(), reference will be made to Fig.35 and Figure 39 to Figure 42.

[0504] Fig.46 is a flowchart illustrating a method of transmitting point cloud data according to an embodiment.

[0505] The method of transmitting point cloud data according to an embodiment may include encoding the point cloud data ( 18001 ) and transmitting a bit stream including the encoded point cloud data and signaling information ( 18002 ).

[0506] In operation 18001 of encoding point cloud data, Figure 1 Point cloud video encoder 10002, Figure 2 Code 20001, Figure 4 Point cloud video encoder, Fig.12 Point cloud video encoder, Fig.14 Point cloud coding, Fig.15 Point cloud encoding and Fig.18 Some or all operations of the point cloud video encoder are performed.

[0507] The operation 18001 of encoding the point cloud data may include encoding geometric information of the point cloud data and encoding attribute information of the point cloud data. In this case, the encoding according to the embodiment may be performed in units of slices or tiles including one or more slices.

[0508] Operation 18002 of sending a bit stream including the encoded point cloud data and signaling information may be sending Fig.35 The G-PCC bitstream structure or Fig.43 Operation 17002 of sending a bit stream including coded point cloud data and signaling information may be performed by Figure 1 Transmitter 10003, Figure 2 Send 20002, Fig.14 Point cloud encoder or file / segment encapsulation unit, Fig.12 The sending processor 12012, or Fig.18 The sending processor 14005 executes.

[0509] The G-PCC bitstream according to an embodiment may have Fig.35 and Figure 39 to Figure 42 The structure shown in the figure, and the G-PCC access unit may have Figure 43 to Figure 45 The illustrated structure. The G-PCC bitstream and / or the G-PCC access unit may include SPS, GPS, APS, TPS, geometry slices, attribute slices, and metadata slices.

[0510] Fig.47 is a flowchart illustrating a method of receiving point cloud data according to an embodiment.

[0511] According to an embodiment, a method for receiving point cloud data may include receiving a bit stream including point cloud data and signaling information (19001), decoding the point cloud data (19002), and rendering the decoded point cloud data operation (19003).

[0512] The bit stream received in the receiving operation 19001 according to the embodiment may have a G-PCC bit stream structure or a G-PCC access unit structure. The G-PCC bit stream may be configured as follows: Fig.35 and 39 to Fig.42 The structure shown in the figure, and the G-PCC access unit can be configured as follows Figure 43 to Figure 45 The structure shown in FIG. 1 is a block diagram of a G-PCC bitstream and / or a G-PCC access unit. The G-PCC bitstream and / or the G-PCC access unit may include SPS, GPS, APS, TPS, geometry slices, attribute slices, and metadata slices.

[0513] Operation 19001 of receiving a bit stream including point cloud data and signaling information may be performed by Figure 1 Receiver 10005, Figure 2 Send 20002 or decode 2003, Fig.13 The receiver 13000 or the receiving processor 13001, Fig.14 A file / fragment decapsulation unit or a point cloud decoding unit, or Fig. 20 The receiving processor 15001 executes.

[0514] The operation 19002 of decoding the point cloud data may include decoding geometric information of the point cloud data based on the signaling information and decoding attribute information of the point cloud data. The decoding operation according to the embodiment may be performed in units of slices or tiles including one or more slices.

[0515] In operation 19002 of decoding point cloud data, Figure 1 Point cloud video decoder 10006, Figure 2 Decoding 20003, Fig.11 Point cloud video decoder, Fig.13 Point cloud video decoder, Fig.14 Point cloud decoding, Fig.16 Point cloud decoding, and Fig. 20 Some or all operations of the point cloud video decoder are performed.

[0516] In operation 19003 of rendering point cloud data, the decoded point cloud data may be rendered according to various rendering methods. For example, points of the point cloud content may be rendered to vertices with certain thicknesses, to cubes of a certain minimum size centered at the vertex position, or to circles centered at the vertex position. All or part of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).

[0517] Operation 19003 of rendering point cloud data can be performed by Figure 1 Renderer 10007, Figure 2 Rendering 20004, Fig.13 Renderer 13011, Fig.14 Rendering units or Fig.16 The point cloud rendering unit is executed.

[0518] Each of the above parts, modules or units can be software, processors or hardware parts that perform the continuous process stored in the memory (or storage unit). Each of the steps described in the above embodiments can be performed by a processor, software or hardware part. Each module / block / unit described in the above embodiments can be used as a processor, software or hardware operation. In addition, the method proposed in the embodiment can be performed as a code. The code can be written on a processor-readable storage medium, and is therefore read by a processor provided by the device.

[0519] In the specification, when a part "includes" or "comprises" an element, it means that the part also includes or comprises another element unless otherwise mentioned. In addition, the term "...module (or unit)" disclosed in the specification means a unit for processing at least one function or operation, and can be implemented by hardware, software, or a combination of hardware and software.

[0520] Although the embodiments are described with reference to each of the accompanying drawings for the sake of simplicity, new embodiments may be designed by combining the embodiments illustrated in the accompanying drawings. If a person skilled in the art designs a computer-readable recording medium having a program for executing the embodiments mentioned in the above description recorded therein, it may fall within the scope of the appended claims and their equivalents.

[0521] The device and method are not limited to the configuration and method of the above-described embodiments. The above-described embodiments can be configured by selectively combining with each other in whole or in part to enable various modifications.

[0522] Although the preferred embodiment has been described with reference to the accompanying drawings, it will be appreciated by those skilled in the art that various modifications and variations may be made in the embodiments without departing from the spirit or scope of the present disclosure described in the appended claims. Such modifications will not be understood independently of the technical ideas or viewpoints of the embodiments.

[0523] The various elements of the device of the embodiment can be implemented by hardware, software, firmware or a combination thereof. The various elements in the embodiment can be implemented by a single chip (e.g., a single hardware circuit). According to an embodiment, the components according to the embodiment can be implemented as separate chips respectively. According to an embodiment, at least one or more of the components of the device according to the embodiment may include one or more processors capable of executing one or more programs. The one or more programs may execute any one or more of the operations / methods according to the embodiment, or include instructions for executing them. The executable instructions for executing the method / operation of the device according to the embodiment may be stored in a non-transient CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transient CRM or other computer program product configured to be executed by one or more processors. In addition, the memory according to the embodiment can be used as a concept that not only covers volatile memory (e.g., RAM) but also covers non-volatile memory, flash memory and PROM. In addition, it can also be implemented in the form of a carrier such as sending through the Internet. In addition, the processor-readable recording medium can be distributed in a computer system connected by a network so that the processor-readable code can be stored and executed in a distributed manner.

[0524] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". In addition, "A, B" may mean "A and / or B". In addition, "A / B / C" may mean "at least one of A, B, and / or C". "A, B, C" may also mean "at least one of A, B, and / or C".

[0525] In addition, in this document, the term "or" should be interpreted as "and / or". For example, the expression "A or B" may mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as "additionally or alternatively".

[0526] The various elements of the embodiments can be described using terms such as first and second. However, the various components according to the embodiments should not be limited by the above terms. These terms are only used to distinguish one element from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of these terms should be interpreted as not departing from the scope of the various embodiments.

[0527] The first user input signal and the second user input signal are both user input signals but are not meant to be the same user input signal unless the context clearly indicates otherwise.The terms used to describe the embodiments are used only for the purpose of describing particular embodiments and are not intended to be limiting of the embodiments.

[0528] As used in the description of the embodiments and in the claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. The expression "and / or" is used to include all possible combinations of terms. Terms such as "including" or "having" are intended to indicate the presence of figures, numbers, steps, elements, and / or components, and should be understood as not excluding the possibility of the presence of additional figures, numbers, steps, elements, and / or components.

[0529] As used herein, conditional expressions such as “if” and “when” are not limited to optional cases, but are intended to be interpreted as performing relevant operations or interpreting relevant definitions according to the specific conditions when the specific conditions are met.

[0530] Mode for the Invention

[0531] As described above, the relevant contents have been described in the best mode for carrying out the embodiment.

[0532] Industrial Applicability

[0533] As described above, the embodiments may be applied in whole or in part to point cloud data sending / receiving devices and systems. It will be clear to those skilled in the art that various changes or modifications may be made to the embodiments within the scope of the embodiments. Therefore, the embodiments are intended to cover modified forms and variations of the present disclosure, provided that they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for sending point cloud data, the method comprising: Encoding the point cloud data including geometric information and attribute information by an encoder; as well as The transmitter sends a bit stream including the encoded point cloud data and signaling information. The signaling information includes a sequence parameter set including information related to the point cloud data sequence, a geometry parameter set including information related to geometry, and an attribute parameter set including information related to attributes. The bit stream is composed of data units, each of which has type information and a payload. wherein the type information indicates the type of data included in the payload, wherein the payload includes a geometry data unit, an attribute data unit, a sequence parameter set, one of the geometry parameter set and the attribute parameter set, The geometry data unit includes a geometry data unit header and a portion of the geometry information. The geometry data unit header includes at least one of identification information for specifying the geometry parameter set associated with the geometry data unit or slice information associated with the geometry data unit, The attribute data unit includes an attribute data unit header and a portion of the attribute information. The attribute data unit header includes at least one of identification information for specifying the attribute parameter set related to the attribute data unit or information for specifying a geometric slice related to the attribute data unit.

2. The method according to claim 1, wherein: The encoding includes: spatially partitioning the geometric information and the attribute information into slices or tiles, each of the tiles comprising one or more slices; and The geometric information and the attribute information are encoded based on the slice or the tile.

3. A device for sending point cloud data, the device comprising: an encoder configured to encode the point cloud data including geometric information and attribute information; as well as a transmitter configured to transmit a bit stream including encoded point cloud data and signaling information, The signaling information includes a sequence parameter set including information related to the point cloud data sequence, a geometry parameter set including information related to geometry, and an attribute parameter set including information related to attributes. The bit stream is composed of data units, each of which has type information and a payload. wherein the type information indicates the type of data included in the payload, wherein the payload includes a geometry data unit, an attribute data unit, a sequence parameter set, one of the geometry parameter set and the attribute parameter set, The geometry data unit includes a geometry data unit header and a portion of the geometry information. The geometry data unit header includes at least one of identification information for specifying the geometry parameter set associated with the geometry data unit or slice information associated with the geometry data unit, The attribute data unit includes an attribute data unit header and a portion of the attribute information. The attribute data unit header includes at least one of identification information for specifying the attribute parameter set related to the attribute data unit or information for specifying a geometric slice related to the attribute data unit.

4. The device according to claim 3, wherein: The encoder comprises: a spatial partitioner configured to spatially partition the geometric information and the attribute information into slices or tiles, each of the tiles comprising one or more slices; and A video encoder is configured to encode the geometric information and the attribute information based on the slice or the tile.

5. A method for receiving point cloud data, the method comprising: A receiver receives a bit stream including the point cloud data and signaling information, wherein the point cloud data includes geometric information and attribute information; as well as Decoding the point cloud data by a decoder based on the signaling information; The signaling information includes a sequence parameter set including information related to the point cloud data sequence, a geometry parameter set including information related to geometry, and an attribute parameter set including information related to attributes. The bit stream is composed of data units, each of which has type information and a payload. wherein the type information indicates the type of data included in the payload, wherein the payload includes a geometry data unit, an attribute data unit, a sequence parameter set, one of the geometry parameter set and the attribute parameter set, The geometry data unit includes a geometry data unit header and a portion of the geometry information. The geometry data unit header includes at least one of identification information for specifying the geometry parameter set associated with the geometry data unit or slice information associated with the geometry data unit, The attribute data unit includes an attribute data unit header and a portion of the attribute information. The attribute data unit header includes at least one of identification information for specifying the attribute parameter set related to the attribute data unit or information for specifying a geometric slice related to the attribute data unit.

6. The method according to claim 5, wherein: The decoding includes: Based on the signaling information, decoding the geometric information and the attribute information on a slice basis or a tile including one or more slices; and Based on the signaling information, the slice-based or tile-based decoded geometry information and attribute information are spatially combined.

7. A device for receiving point cloud data, the device comprising: a receiver configured to receive a bit stream including the point cloud data and signaling information, the point cloud data including geometric information and attribute information; as well as a decoder configured to decode the point cloud data based on the signaling information; The signaling information includes a sequence parameter set including information related to the point cloud data sequence, a geometry parameter set including information related to geometry, and an attribute parameter set including information related to attributes. The bit stream is composed of data units, each of which has type information and a payload. wherein the type information indicates the type of data included in the payload, wherein the payload includes a geometry data unit, an attribute data unit, a sequence parameter set, one of the geometry parameter set and the attribute parameter set, The geometry data unit includes a geometry data unit header and a portion of the geometry information. The geometry data unit header includes at least one of identification information for specifying the geometry parameter set associated with the geometry data unit or slice information associated with the geometry data unit, The attribute data unit includes an attribute data unit header and a portion of the attribute information. The attribute data unit header includes at least one of identification information for specifying the attribute parameter set related to the attribute data unit or information for specifying a geometric slice related to the attribute data unit.

8. The device according to claim 7, wherein: The decoder comprises: a video decoder configured to decode the geometric information and the attribute information on a slice basis or a tile including one or more slices based on the signaling information; and A post-processor is configured to spatially combine the slice-based or tile-based decoded geometry information and attribute information based on the signaling information.

Citation Information

Patent Citations

  • Point cloud compression

    US20190087979A1