Point cloud data processing apparatus and method

The method and apparatus efficiently process point cloud data by encoding geometry and attribute information in bitstreams, addressing latency and complexity issues to support high-quality VR and autonomous driving services.

JP7798974B2Active Publication Date: 2026-01-14LG ELECTRONICS INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024115905
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-28
Filing Date
2024-07-19
Publication Date
2026-01-14
Estimated Expiration
2040-06-01

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing massive amounts of point cloud data required for VR, AR, MR, and autonomous driving services due to high latency and encoding/decoding complexity.

Method used

A method and apparatus for encoding and decoding point cloud data, including geometry and attribute information, using bitstreams to facilitate efficient processing and transmission.

Benefits of technology

The solution enables high-efficiency processing of point cloud data, providing high-quality services such as VR and autonomous driving by optimizing encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798974000022
    Figure 0007798974000022
  • Figure 0007798974000023
    Figure 0007798974000023
  • Figure 0007798974000024
    Figure 0007798974000024
Patent Text Reader

Abstract

To provide an apparatus and method for processing point cloud data.SOLUTION: A method for processing point cloud data according to embodiments comprises encoding and transmitting point cloud data. The method for processing point cloud data according to embodiments comprises receiving and decoding point cloud data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiment provides a method for providing point cloud content to provide users with various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality) and autonomous driving services. [Background technology]

[0002] Point cloud content is content expressed as a point cloud, which is a collection of points belonging to a coordinate system that represents three-dimensional space. Point cloud content can represent three-dimensional media and is used to provide various services such as VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and autonomous driving services. However, expressing point cloud content requires tens of thousands to hundreds of thousands of point data. Therefore, a method for efficiently processing massive amounts of point data is required. Summary of the Invention [Problem to be solved by the invention]

[0003] SUMMARY OF THE INVENTION The embodiments provide an apparatus and method for efficiently processing point cloud data.The embodiments provide a method and apparatus for processing point cloud data to address latency and encoding / decoding complexity.

[0004] However, the scope of the invention is not limited to the above-mentioned technical problems, but can be extended to other technical problems that a person skilled in the art can derive based on all the contents described herein. [Means for solving the problem]

[0005] Therefore, in order to efficiently process point cloud data, a point cloud processing method according to an embodiment includes encoding point cloud data including geometry information and attribute information, and transmitting a bitstream including the encoded point cloud data. The geometry information according to an embodiment is information indicating positions of points in the point cloud data, and the attribute information is information indicating attributes of points in the point cloud data.

[0006] A point cloud processing device according to an embodiment includes an encoder for encoding point cloud data including geometry information and attribute information, and a transmitter for transmitting a bitstream including the encoded point cloud data. The geometry information according to the embodiment is information indicating positions of points in the point cloud data, and the attribute information is information indicating attributes of the points in the point cloud data.

[0007] A point cloud processing method according to an embodiment includes receiving a bitstream including point cloud data, and decoding the point cloud data. The point cloud data according to an embodiment includes geometry information and attribute information. The geometry information according to an embodiment is information indicating positions of points of the point cloud data, and the attribute information according to an embodiment is information indicating one or more attributes of the points of the point cloud data.

[0008] A point cloud processing device according to an embodiment includes a receiver for receiving a bitstream including point cloud data, and a decoder for decoding the point cloud data. The point cloud data according to an embodiment includes geometry information and attribute information. The geometry information according to an embodiment is information indicating positions of points of the point cloud data, and the attribute information according to an embodiment is information indicating one or more attributes of the points of the point cloud data. [Effects of the Invention]

[0009] The apparatus and method according to the embodiment can process point cloud data with high efficiency.

[0010] The apparatus and method according to the embodiment can provide a high-quality point cloud service.

[0011] The apparatus and method according to the embodiment can provide point cloud content for providing general-purpose services such as VR services and autonomous driving services. [Brief explanation of the drawings]

[0012] The accompanying drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0013] The drawings are attached for a better understanding of the embodiments and together with the description relating to the embodiments illustrate the embodiments.

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 3l

Figure 32

Figure 33

[0015] Preferred embodiments will be described in detail with reference to the accompanying drawings. The following detailed description with reference to the accompanying drawings is intended to illustrate preferred embodiments rather than merely showing embodiments that can be implemented by the embodiments. The following detailed description includes details to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that embodiments can be practiced without such details.

[0016] Although most of the terms used in the examples are common and widely used in the relevant fields, some of them have been arbitrarily selected by the applicant, and their meanings will be explained in detail below as necessary. Therefore, the examples should be understood based on the intended meaning of the terms, rather than the simple names or meanings of the terms.

[0017] FIG. 1 is a diagram illustrating an example of a point cloud content providing system according to an embodiment.

[0018] The point cloud content providing system shown in Figure 1 includes a transmission device 10000 and a reception device 10004. The transmission device 10000 and the reception device 10004 are capable of wired and wireless communication to transmit and receive point cloud data.

[0019] According to an embodiment, the transmitting device 10000 acquires, processes, and transmits a point cloud video (or point cloud content). In the embodiment, the transmitting device 10000 includes a fixed station, a base transceiver system (BTS), a network, an AI (Artificial Intelligence) device and / or system, a robot, an AR / VR / XR device and / or server, etc. In the embodiment, the transmitting device 10000 also includes a device that communicates with a base station and / or other wireless devices using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)), a robot, a vehicle, an AR / VR / XR device, a mobile device, a home appliance, an IoT (Internet of Things) device, an AI device / server, etc.

[0020] The transmitting device 10000 according to the embodiment includes a Point Cloud Video Acquisition unit 10001, a Point Cloud Video Encoder 10002, and / or a Transmitter (or communication module) 10003.

[0021] The point cloud video acquisition unit 10001 according to the embodiment acquires a point cloud video through a process such as capturing, synthesizing, or generating. The point cloud video is point cloud content represented by a point cloud, which is a collection of points located in a three-dimensional space, and is also called point cloud video data. The point cloud video according to the embodiment includes one or more frames. One frame represents a still image / picture. Therefore, the point cloud video includes a point cloud image / frame / picture, and is called any of a point cloud image, a frame, and a picture.

[0022] The point cloud video encoder 10002 according to the embodiment encodes the secured point cloud video data. The point cloud video encoder 10002 encodes the point cloud video data based on point cloud compression coding. The point cloud compression coding according to the embodiment includes Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding, or next-generation coding. Note that the point cloud compression coding according to the embodiment is not limited to the above-described embodiments. The point cloud video encoder 10002 can output a bitstream including encoded point cloud video data. The bitstream includes not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0023] According to an embodiment, the transmitter 10003 transmits a bitstream including encoded point cloud video data. According to an embodiment, the bitstream is encapsulated into a file or a segment (e.g., a streaming segment) and transmitted via various networks such as a broadcast network and / or a broadband network. Although not shown, the transmitting device 10000 includes an encapsulation unit (or encapsulation module) that performs the encapsulation operation. In an embodiment, the encapsulation unit is included in the transmitter 10003. According to an embodiment, the file or segment is transmitted to the receiving device 10004 via a network or stored on a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). According to an embodiment, the transmitter 10003 can communicate with the receiving device 10004 (or receiver 10005) via wired or wireless communication via a network such as 4G, 5G, or 6G. Furthermore, the transmitter 10003 can perform necessary data processing operations via a network system (e.g., a communication network system such as 4G, 5G, or 6G). The transmitting device 10000 can also transmit encapsulated data on an on-demand basis.

[0024] The receiving device 10004 according to the embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. In the embodiment, the receiving device 10004 includes a device, robot, vehicle, AR / VR / XR device, mobile device, home appliance, IoT (Internet of Things) device, AI device / server, etc. that communicates with a base station and / or other wireless device using a wireless connection technology (e.g., 5G NR (New RAT), LTE (Long Term Evolution)).

[0025] The receiver 10005 according to the embodiment receives a bitstream containing point cloud video data or a file / segment in which the bitstream is encapsulated from a network or storage medium. The receiver 10005 performs data processing operations required by a network system (e.g., a communication network system such as 4G, 5G, or 6G). The receiver 10005 according to the embodiment decapsulates the received file / segment and outputs a bitstream. In addition, in the embodiment, the receiver 10005 includes a decapsulation unit (or decapsulation module) for performing the decapsulation operation. The decapsulation unit is embodied as an element (or component) separate from the receiver 10005.

[0026] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data in the manner in which it was encoded (e.g., the reverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud reconstruction coding, which is the reverse process of point cloud compression. Point cloud reconstruction coding includes G-PCC coding.

[0027] The renderer 10007 renders the decoded point cloud video data. The renderer 10007 renders not only the point cloud video data but also the audio data to output point cloud content. In an embodiment, the renderer 10007 includes a display for displaying the point cloud content. In an embodiment, the display is not included in the renderer 10007, but is embodied as a separate device or component.

[0028] In the drawing, dotted arrows indicate the transmission path of feedback information obtained by the receiving device 10004. The feedback information is information for reflecting interaction with a user consuming point cloud content and includes user information (e.g., head orientation information, viewport information, etc.). In particular, if the point cloud content is for a service that requires interaction with a user (e.g., an autonomous driving service), the feedback information may be transmitted to a content transmitting side (e.g., the transmitting device 10000) and / or a service provider. In an embodiment, the feedback information may be used not only by the transmitting device 10000 but also by the receiving device 10004, or may not be provided.

[0029] According to an embodiment, head orientation information is information regarding the position, direction, angle, movement, etc. of the user's head. According to an embodiment, the receiving device 10004 calculates viewport information based on the head orientation information. The viewport information is information regarding the area of ​​the point cloud video viewed by the user. The viewpoint is the point at which the user views the point cloud video and refers to the center of the viewport area. That is, the viewport is an area centered on the viewpoint, and the size, shape, etc. of the area are determined by the FOV (Field of View). Therefore, the receiving device 10004 can extract viewport information based on the vertical or horizontal FOV supported by the device in addition to the head orientation information. The receiving device 10004 also performs gaze analysis to determine the user's point cloud consumption method, the point cloud video area the user is gazing at, the gaze time, etc. In an embodiment, the receiving device 10004 can transmit feedback information including the results of the gaze analysis to the transmitting device 10000. According to an embodiment, the feedback information is obtained during the rendering and / or display process. According to this embodiment, feedback information is obtained by one or more sensors included in the receiving device 10004. In another embodiment, feedback information is obtained by the renderer 10007 or another external element (or device, component, etc.). The dotted lines in FIG. 1 indicate the transmission process of feedback information obtained by the renderer 10007. The point cloud content providing system processes (encodes / decodes) point cloud data based on the feedback information. Therefore, the point cloud video data decoder 10006 can perform a decoding operation based on the feedback information. Furthermore, the receiving device 10004 can transmit the feedback information to the transmitting device 10000. The transmitting device 10000 (or point cloud video data encoder 10002) can perform an encoding operation based on the feedback information.Therefore, the point cloud content providing system does not process (encode / decode) all point cloud data, but can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on feedback information to provide point cloud content to the user.

[0030] In an embodiment, the sending device 10000 may be referred to as an encoder, a sending device, a transmitter, etc., and the receiving device 10004 may be referred to as a decoder, a receiving device, a receiver, etc.

[0031] 1 according to an embodiment (processed through a series of steps of acquisition / encoding / transmission / decoding / rendering), the point cloud data may also be referred to as point cloud content data or point cloud video data. In an embodiment, the point cloud content data may be used as a concept including metadata or signaling information related to the point cloud data.

[0032] The elements of the point cloud content providing system shown in FIG. 1 may be implemented in hardware, software, a processor, and / or a combination thereof.

[0033] FIG. 2 is a block diagram illustrating the operation of providing point cloud content according to an embodiment.

[0034] Figure 2 is a block diagram showing the operation of the point cloud content providing system described in Figure 1. As described above, the point cloud content providing system processes point cloud data based on point cloud compression coding (e.g., G-PCC).

[0035] A point cloud content providing system (e.g., a point cloud transmitting device 10000 or a point cloud video acquiring unit 10001) according to an embodiment acquires a point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system representing a three-dimensional space. The point cloud video according to an embodiment includes a Ply (Polygon File format or the Stanford Triangle format) file. If the point cloud video has one or more frames, the acquired point cloud video includes one or more Ply files. A Ply file includes point cloud data such as the geometry and / or attributes of points. The geometry includes the position of the point. The position of each point is expressed by parameters (e.g., values ​​on the X-axis, Y-axis, and Z-axis) indicating a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). The attributes include the characteristics of the point (e.g., texture information of each point, color (YCbCr or RGB), reflectance (r), transparency, etc.). A point has one or more characteristics (or attributes). For example, one point may have one characteristic of hue, or two characteristics of hue and reflectance. In the embodiment, geometry may be referred to as position, geometry information, geometry data, etc., and characteristics may be referred to as characteristics, characteristic information, characteristic data, etc. In addition, the point cloud content providing system (e.g., the point cloud transmitting device 10000 or the point cloud video acquiring unit 10001) may acquire point cloud data from information related to the point cloud video acquiring process (e.g., depth information, color information, etc.).

[0036] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment encodes point cloud data (20001). The point cloud content providing system encodes point cloud data based on point cloud compression coding. As described above, point cloud data includes the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry coding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute coding to encode the attributes and output a attribute bitstream. In an embodiment, the point cloud content providing system can perform attribute coding based on the geometry coding. The geometry bitstream and the attribute bitstream according to an embodiment are multiplexed and output as a single bitstream. The bitstream according to an embodiment further includes signaling information related to the geometry coding and the attribute coding.

[0037] A point cloud content providing system (e.g., transmitting device 10000 or transmitter 10003) according to an embodiment transmits encoded point cloud data (20002). As described in FIG. 1, the encoded point cloud data is represented by a geometry bitstream and a feature bitstream. The encoded point cloud data is transmitted in the form of a bitstream together with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometry encoding and feature encoding). The point cloud content providing system encapsulates the bitstream for transmitting the encoded point cloud data and transmits it in the form of a file or segment.

[0038] A point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to an embodiment receives a bitstream including encoded point cloud data, and can demultiplex the bitstream.

[0039] The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the encoded point cloud data (e.g., geometry bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the point cloud video data based on signaling information related to the encoding of the point cloud video data included in the bitstream. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) decodes the geometry bitstream to restore the position (geometry) of the point. The point cloud content providing system decodes the attribute bitstream based on the restored geometry to restore the attribute of the point. The point cloud content providing system (e.g., the receiving device 10004 or the point cloud video decoder 10005) restores the point cloud video based on the position according to the restored geometry and the decoded attribute.

[0040] A point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) according to an embodiment renders the decoded point cloud data (20004). The point cloud content providing system (e.g., a receiving device 10004 or a renderer 10007) renders the geometry and characteristics decoded during the decoding process using various rendering methods. Points of the point cloud content are rendered as fixed points with a certain thickness, cubes with a predetermined minimum size centered at the position of the fixed points, or circles centered at the position of the fixed points. All or part of the region of the rendered point cloud content is provided to a user via a display (e.g., a VR / AR display, a general display, etc.).

[0041] A point cloud content providing system (e.g., receiving device 10004) according to the embodiment may obtain feedback information (20005). The point cloud content providing system encodes and / or decodes point cloud data based on the feedback information. The feedback information and operation of the point cloud content providing system according to the embodiment are the same as the feedback information and operation described in FIG. 1, so a detailed description will be omitted.

[0042] FIG. 3 illustrates an example of a point cloud video capturing process according to an embodiment.

[0043] FIG. 3 illustrates an example of a point cloud video capture process in the point cloud content providing system described in FIGS.

[0044] Point cloud content includes point cloud video (images and / or video) showing objects and / or environments located in various three-dimensional spaces (e.g., a three-dimensional space showing a real environment, a three-dimensional space showing a virtual environment, etc.). Accordingly, a point cloud content providing system according to an embodiment captures point cloud video using one or more cameras (e.g., an infrared camera capable of obtaining depth information, an RGB camera capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., an infrared pattern projector for obtaining depth information), LiDAR, etc. to generate point cloud content. The point cloud content providing system according to an embodiment extracts a geometric form composed of points in three-dimensional space from the depth information and extracts characteristics of each point from the color information to obtain point cloud data. Images and / or video according to an embodiment are captured based on either an inward-facing approach or an outward-facing approach.

[0045] The left side of Figure 3 shows the inward-facing method. The inward-facing method is a method in which one or more cameras (or camera sensors) positioned around a central object capture the central object. The inward-facing method is used to generate point cloud content that provides the user with a 360-degree image of the core object (e.g., VR / AR content that provides the user with a 360-degree image of an object (e.g., a core object such as a character, player, item, or actor)).

[0046] The right side of Figure 3 shows an outward-facing approach. The outward-facing approach is a method in which one or more cameras (or camera sensors) positioned around a central object capture the environment of the central object, which is not the central object. The outward-facing approach is used to generate point cloud content to provide the surrounding environment from a user's perspective (e.g., content showing the external environment provided to a user of an autonomous vehicle).

[0047] As shown in the figure, point cloud content is generated based on the capture operation of one or more cameras. In this case, since each camera has a different coordinate system, the point cloud content providing system calibrates one or more cameras to set a global coordinate system before the capture operation. The point cloud content providing system also generates point cloud content by combining images and / or video captured using the above capture method with an arbitrary image and / or video. When generating point cloud content representing a virtual space, the point cloud content providing system does not perform the capture operation described in FIG. 3. The point cloud content providing system according to the embodiment can also perform post-processing on the captured images and / or video. That is, the point cloud content providing system can remove unwanted areas (e.g., background) or fill in any spatial holes by recognizing the space where the captured images and / or video are connected.

[0048] The point cloud content providing system can also generate a single point cloud content by performing coordinate system transformation on points in the point cloud video acquired from each camera. The point cloud content providing system performs coordinate system transformation on points based on the position coordinates of each camera. This allows the point cloud content providing system to generate content showing a single wide area or point cloud content with a high point density.

[0049] FIG. 4 is a diagram illustrating an example of a point cloud encoder according to an embodiment.

[0050] 4 shows an example of the point cloud video encoder 10002 of FIG. 1. The point cloud encoder performs encoding by reconstructing point cloud data (e.g., point positions and / or characteristics) to adjust the quality of the point cloud content (e.g., lossless, lossy, near-lossless) depending on the network conditions or application. If the overall size of the point cloud content is large (e.g., point cloud content of 60 Gbps for 30 fps), the point cloud content providing system cannot stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide it according to the network environment, etc.

[0051] As shown in Figures 1 and 2, a point cloud encoder can perform geometry encoding and feature encoding, where geometry encoding is performed before feature encoding.

[0052] The point cloud encoder according to the embodiment includes a coordinate system transformation unit (Transformation Coordinates) 40000, a quantization unit (Quantize and Remove Points (Voxelize)) 40001, an octree analysis unit (Analyze Octree) 40002, a surface approximation analysis unit (Analyze Surface Approximation) 40003, an arithmetic encoder (Arithmetic Encode) 40004, a geometry reconstruction unit (Reconstruct Geometry) 40005, a color transformation unit (Transform Colors) 40006, a characteristic transformation unit (Transfer Attributes) 40007, a RAHT transformation unit 40008, an LOD generation unit (Generated LOD) 40009, a lift transformation unit (Lifting) 40010, a coefficient quantization unit (Quantize Coefficients) 40011 and / or an arithmetic encoder (Arithmetic Encode) 40012.

[0053] The coordinate system transformation unit 40000, the quantization unit 40001, the octree analysis unit 40002, the surface approximation analysis unit 40003, the arithmetic encoder 40004, and the geometry reconstruction unit 40005 can perform geometry coding. Geometry coding according to the embodiment includes octree geometry coding, direct coding, trisoup geometry encoding, and entropy coding. Direct coding and trisoup geometry encoding are applied selectively or in combination. Note that geometry coding is not limited to the above examples.

[0054] As shown in the figure, a coordinate system conversion unit 40000 according to an embodiment receives a position and converts it into a coordinate system. For example, the position is converted into position information in a three-dimensional space (e.g., a three-dimensional space expressed in an XYZ coordinate system). The position information in the three-dimensional space according to an embodiment is also referred to as geometry information.

[0055] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 quantizes points based on the minimum position value of all points (e.g., the minimum value on each axis for the X, Y, and Z axes). The quantizer 40001 performs a quantization operation by multiplying the difference between the minimum position value and the position value of each point by a predetermined quantization scale value and then rounding down or up to find the nearest integer value. Therefore, one or more points may have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. Just as the smallest unit containing 2D image / video information is a pixel, points in the point cloud content (or 3D point cloud video) according to the embodiment are included in one or more voxels. A voxel is a combination of the words volume and pixel and refers to a 3D cubic space generated when a 3D space is divided into units (unit=1.0) based on the axes (e.g., X-axis, Y-axis, Z-axis) that represent the 3D space. The quantization unit 40001 can match a group of points in the 3D space with voxels. In some embodiments, a voxel may contain only one point. In some embodiments, a voxel may contain one or more points. To represent a voxel as a point, the center of the voxel may be set based on the positions of one or more points contained in the voxel. In this case, the characteristics of all points contained in a voxel are combined and assigned to the voxel.

[0056] The octree analysis unit 40002 according to the embodiment performs octree geometry coding (or octree coding) to represent voxels in an octree structure, which represents points matched to voxels based on an octet structure.

[0057] The surface approximation analysis unit 40003 according to the embodiment analyzes and approximates the octree. The octree analysis and approximation according to the embodiment is a process of analyzing a region containing a large number of points to voxelize it in order to efficiently provide the octree and voxelization.

[0058] The arithmetic encoder 40004 according to the embodiment entropy encodes the octree and / or the approximated octree. For example, the encoding method includes an arithmetic encoding method. As a result of the encoding, a geometry bitstream is generated.

[0059] The color transform unit 40006, the feature transform unit 40007, the RAHT transform unit 40008, the LOD generation unit 40009, the lift transform unit 40010, the coefficient quantization unit 40011, and / or the arithmetic encoder 40012 perform feature coding. As described above, one point has one or more features. Feature coding according to the embodiment is applied equally to all features of one point. However, if one feature (e.g., hue) includes one or more elements, independent feature coding is applied to each element. Feature coding according to the embodiment includes color transform coding, feature transform coding, RAHT (Region Adaptive Hierarchical Transform) coding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) coding. Depending on the point cloud content, the above-mentioned RAHT coding, predictive transform coding, and lift transform coding may be selectively used, or a combination of one or more coding methods may be used. Furthermore, the feature coding according to the embodiment is not limited to the above examples.

[0060] The color converter 40006 according to the embodiment performs color conversion coding to convert the color values ​​(or textures) included in the attributes. For example, the color converter 40006 converts the format of the hue information (e.g., converts from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment is optionally applied depending on the color values ​​included in the attributes.

[0061] The geometry reconstruction unit 40005 according to the embodiment reconstructs (restores) the octree and / or the approximated octree. The geometry reconstruction unit 40005 reconstructs the octree / voxel based on the result of analyzing the distribution of points. The reconstructed octree / voxel is also called a reconstructed geometry (or restored geometry).

[0062] According to an embodiment, the feature converter 40007 performs feature conversion, converting features based on a position where geometry encoding has not been performed and / or reconstructed geometry. As described above, because features depend on geometry, the feature converter 40007 can convert features based on reconstructed geometry information. For example, the feature converter 40007 can convert features of points included in a voxel based on their position values. As described above, if the center point of a voxel is set based on the positions of one or more points included in the voxel, the feature converter 40007 converts features of one or more points. If trisoup geometry encoding is performed, the feature converter 40007 can convert features based on the trisoup geometry encoding.

[0063] The feature conversion unit 40007 performs feature conversion by calculating the average value of the features or feature values ​​(e.g., hue or reflectance of each point) of adjacent points within a specific position / radius from the position (or position value) of the center point of each voxel. When calculating the average value, the feature conversion unit 40007 applies a weight based on the distance from the center point to each point. Therefore, each voxel has a position and a calculated feature (or feature value).

[0064] The feature conversion unit 40007 searches for neighboring points within a specific position / radius from the center point of each voxel based on a KD tree or Moulton code. A KD tree supports a data structure that manages points based on their position, enabling a fast Nearest Neighbor Search (NNS) using a binary search tree. A Moulton code is generated by mixing bits, representing the coordinate values ​​(e.g., (x, y, z)) that indicate the 3D position of all points. For example, if the coordinate value indicating the point's position is (5, 9, 1), the bit values ​​of the coordinate value are (0101, 1001, 0001). Mixing the bit values ​​in the order of z, y, and x according to the bit index results in 010001000111. This value is expressed in decimal as 1095. In other words, the Moulton code value of a point with coordinate values ​​(5, 9, 1) is 1095. The feature conversion unit 40007 aligns points based on the Moulton code value and performs nearest neighbor search (NNS) using a depth-first traversal process. After the feature conversion operation, if nearest neighbor search (NNS) is required in other conversion processes for feature coding, a KD tree or Moulton code is used.

[0065] As shown, the transformed attributes are input to a RAHT transformer 40008 and / or an LOD generator 40009 .

[0066] The RAHT converter 40008 according to the embodiment performs RAHT coding to predict feature information based on the reconstructed geometry information. For example, the RAHT converter 40008 can predict feature information of a node at a higher level of the octree based on feature information associated with a node at a lower level of the octree.

[0067] According to an embodiment, the LOD generator 40009 generates a Level of Detail (LOD) for predictive coding. The LOD according to an embodiment indicates the degree of detail of the point cloud content, and a smaller LOD value indicates less detail of the point cloud content, while a larger LOD value indicates more detail of the point cloud content. Points can be classified according to the LOD.

[0068] The lift transform unit 40010 according to the embodiment performs lift transform coding, which transforms the characteristics of the point cloud based on weights. As described above, the lift transform coding is selectively applied.

[0069] The coefficient quantization unit 40011 according to the embodiment quantizes the feature-coded feature based on the coefficients.

[0070] The arithmetic encoder 40012 according to the embodiment encodes the quantized characteristics based on arithmetic coding.

[0071] The elements of the point cloud encoder of FIG. 4 may be implemented in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors may perform any one of the operations and / or functions of the elements of the point cloud encoder of FIG. 4 described above. The one or more processors may also operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud encoder of FIG. 4. According to embodiments, the one or more memories may include high-speed random access memory or non-volatile memory (e.g., one or more magnetic disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).

[0072] FIG. 5 is a diagram showing an example of a voxel according to the embodiment.

[0073] FIG. 5 shows voxels located in a three-dimensional space expressed by a coordinate system consisting of three axes: X, Y, and Z. As shown in FIG. 4, a point cloud encoder (e.g., quantization unit 40001) performs voxelization. A voxel is a three-dimensional cubic space that is generated when the three-dimensional space is divided into units (unit=1.0) based on the axes (e.g., X, Y, and Z) that represent the three-dimensional space. FIG. 5 shows two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is an example of a voxel generated by an octree structure that recursively subdivides a bounding box (cubical axis-aligned bounding box) defined by the cubic axis-aligned bounding box (B). One voxel contains at least one point. The spatial coordinates of a voxel can be estimated from its positional relationship with other voxel groups. As mentioned above, a voxel has characteristics (such as color or reflectance) just like a pixel in a 2D image / video. A detailed explanation of voxels is omitted here as it has been explained in FIG. 4.

[0074] FIG. 6 is a diagram illustrating an example of an octree and occupancy code according to an embodiment.

[0075] As shown in Figures 1 to 4, a point cloud content providing system (point cloud video encoder 10002) or a point cloud encoder (e.g., octree analysis unit 40002) performs octree geometry coding (or octree coding) based on an octree structure to efficiently manage the area and / or position of voxels.

[0076] The upper part of Figure 6 shows an octree structure. The three-dimensional space of the point cloud content according to the embodiment is represented by the axes of a coordinate system (e.g., X-axis, Y-axis, Z-axis). The octree structure is defined by two extreme points (0,0,0) and (2 d , 2 d , 2 d ) is generated by recursively subdividing the cubic axis-aligned bounding box defined by the point cloud content (or point cloud video). 2d is set to the value that constitutes the smallest bounding box that encloses all points in the point cloud content (or point cloud video). d indicates the depth of the octree. The d value is determined by the following formula: In the following formula, (x int n , y int n , z int n ) indicates the quantized position (or position value) of the point.

[0077]

number

[0078] As shown in the upper center of Figure 6, the entire 3D space is divided into eight spaces through division. Each divided space is represented by a cube with six faces. As shown in the upper right of Figure 6, each of the eight spaces is again divided by the axes of the coordinate system (e.g., X-axis, Y-axis, Z-axis). Therefore, each space is again divided into eight smaller spaces. The divided smaller spaces are also represented by cubes with six faces. This division method is applied until the leaf nodes of the octree become voxels.

[0079] The bottom of Figure 6 shows an occupancy code for an octree. An occupancy code for an octree is generated to indicate whether each of the eight subspaces generated by dividing a space contains at least one point. Therefore, one occupancy code is represented by eight child nodes. Each child node indicates the occupancy of the divided space and has a 1-bit value. Therefore, the occupancy code is represented by an 8-bit code. That is, if the space corresponding to a child node contains at least one point, the corresponding node has a value of 1. If the space corresponding to a node does not contain any points (is empty), the corresponding node has a value of 0. The occupancy code shown in Figure 6 is 00100001, which indicates that the spaces corresponding to the third and eighth child nodes of the eight child nodes each contain at least one point. As shown in the figure, the third and eighth child nodes each have eight child nodes, and each child node is represented by an 8-bit occupancy code. In the figure, the occupied code of the third child node is 10000111, and the occupied code of the eighth child node is 01001111. A point cloud encoder (e.g., arithmetic encoder 40004) according to an embodiment can entropy encode the occupied code. To improve compression efficiency, the point cloud encoder can also intra / inter-code the occupied code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs an octree based on the occupied code.

[0080] A point cloud encoder according to an embodiment (e.g., the point cloud encoder or octree analyzer 40002 in FIG. 4) performs voxelization and octree coding to store point positions. However, since points in a three-dimensional space are not always uniformly distributed, there may be certain areas where there are not many points. Therefore, performing voxelization on the entire three-dimensional space is inefficient. For example, if there are almost no points in a certain area, there is no need to perform voxelization on that area.

[0081] Therefore, the point cloud encoder according to the embodiment does not perform voxelization for the specific region (or nodes other than the leaf nodes of the octree) described above, but performs direct coding, which directly codes the positions of points included in the specific region. The coordinates of direct coding points according to the embodiment are called Direct Coding Mode (DCM). The point cloud encoder according to the embodiment can also perform trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) based on voxels based on a surface model. Trisoup geometry encoding is a geometry encoding that represents an object as a series of triangle meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Direct coding and trisoup geometry encoding according to the embodiment can be performed selectively. Furthermore, direct coding and trisoup geometry encoding according to the embodiment can be combined with octree geometry coding (or octree coding).

[0082] To perform direct coding, a direct mode option for applying direct coding must be activated, and the node to which direct coding is applied must not be a leaf node, but must have points below a threshold within a specific node. Furthermore, the total number of points to be subjected to direct coding must not exceed a predetermined threshold. If the above conditions are met, the point cloud encoder (or the operation encoder 40004) according to the embodiment can entropy code the positions (or position values) of the points.

[0083] A point cloud encoder (e.g., the surface approximation analysis unit 40003) according to an embodiment can determine a specific level of the octree (if the level is smaller than the depth d of the octree) and perform trisoup geometry encoding (trisoup mode) from that level, which uses a surface model to reconstruct the positions of points within the node area based on voxels. The point cloud encoder according to an embodiment can specify the level to which trisoup geometry encoding is applied. For example, if the specified level is the same as the depth of the octree, the point cloud encoder does not operate in trisoup mode. In other words, the point cloud encoder according to an embodiment can operate in trisoup mode only when the specified level is smaller than the depth value of the octree. The 3D cubic area of ​​a node at a specified level according to an embodiment is called a block. One block contains one or more voxels. A block or voxel can also correspond to a brick. Within each block, geometry is represented as a surface. According to an embodiment, a surface can intersect each edge of the block at most once.

[0084] Since one block has 12 edges, there are at least 12 intersections within one block. Each intersection is called a vertex. A vertex along an edge is detected if there is at least one occupied voxel adjacent to the edge among all blocks that share the edge. In this embodiment, an occupied voxel refers to a voxel that contains a point. The position of the vertex detected along the edge is the average position along the edge of all voxels adjacent to the edge among all blocks that share the edge.

[0085] When a vertex is detected, the point cloud encoder according to the embodiment can entropy code the edge start point (x, y, z), the edge direction vector (Δx, Δy, Δz), and the vertex position value (relative position value within the edge). When trisoup geometry encoding is applied, the point cloud encoder according to the embodiment (e.g., the geometry reconstruction unit 40005) can perform triangle reconstruction, up-sampling, and voxelization processes to generate a restored geometry (reconstructed geometry).

[0086] JPEG0007798974000002.jpg34170

[0087]

number

[0088] The minimum value of the added values ​​is found, and a projection process is performed along the axis where the minimum value is. For example, if the x element is the smallest, each vertex is projected onto the x-axis based on the center of the block, and then onto the (y,z) plane. If the value obtained by projecting onto the (y,z) plane is (ai,bi), the θ value is found using atan2(bi,ai), and the vertices are aligned based on the θ value. The following table shows the vertex combinations required to create a triangle depending on the number of vertices. Vertices are aligned in order from 1 to n. The following table shows that for four vertices, two triangles are formed by combining the vertices. The first triangle is formed by the 1st, 2nd, and 3rd vertices of the aligned vertices, and the second triangle is formed by the 3rd, 4th, and 1st vertices of the aligned vertices.

[0089] Table.Triangles formed from vertices ordered 1

[0090] [Table 1]

[0091] The upsampling process is performed to add intermediate points along the edges of triangles for voxelization. The additional points are generated based on the upsampling factor and the block width. These additional points are called refined vertices. The point cloud encoder according to the embodiment can voxelize the refined vertices. The point cloud encoder can also perform feature encoding based on the voxelized positions (or position values).

[0092] FIG. 7 is a diagram illustrating an example of an adjacent node pattern according to the embodiment.

[0093] To increase the compression efficiency of point cloud videos, the point cloud encoder according to the embodiment performs entropy coding based on context adaptive arithmetic coding.

[0094] As described with reference to FIGS. 1 to 6, a point cloud content providing system or a point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or arithmetic encoder 40004 in FIG. 4) can immediately entropy code the occupied code. The point cloud content providing system or the point cloud encoder can also perform entropy coding (intra coding) based on the occupied code of the current node and the occupied rate of neighboring nodes, or can perform entropy coding (inter coding) based on the occupied code of a previous frame. A frame according to the present embodiment refers to a collection of point cloud videos generated at the same time. The compression efficiency of intra coding / inter coding according to the present embodiment varies depending on the number of neighboring nodes to reference. Although the complexity increases as the number of bits increases, the compression efficiency can be improved by focusing on one side. For example, a 3-bit context requires eight coding methods, which is equal to the cube of two. The separately coded portion affects the complexity of the implementation. Therefore, it is necessary to balance compression efficiency and complexity at an appropriate level.

[0095] FIG. 7 shows a process for determining an occupancy pattern based on the occupancy of neighboring nodes. The point cloud encoder according to the embodiment obtains a neighbor pattern value by determining the occupancy of neighboring nodes for each node in an occupancy tree. The neighboring node pattern is used to infer the occupancy pattern of the corresponding node. The left side of FIG. 7 shows a cube corresponding to the node (the cube located in the middle) and six cubes (neighboring nodes) that share at least one side with the corresponding cube. The nodes shown are nodes at the same depth. The numbers shown indicate the weights (1, 2, 4, 8, 16, 32, etc.) associated with each of the six nodes. Each weight is assigned in order according to the position of the neighboring node.

[0096] The right side of FIG. 7 shows the adjacent node pattern value. The adjacent node pattern value is the sum of values ​​multiplied by the weight values ​​of occupied adjacent nodes (adjacent nodes with points). Therefore, the adjacent node pattern value ranges from 0 to 63. An adjacent node pattern value of 0 means that there are no nodes with points (occupied nodes) among the adjacent nodes of the corresponding node. An adjacent node pattern value of 63 means that all adjacent nodes are occupied nodes. As shown in the figure, adjacent nodes assigned weight values ​​of 1, 2, 4, and 8 are occupied nodes, so the adjacent node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform coding according to the adjacent node pattern value (e.g., if the adjacent node pattern value is 63, 64 coding is performed). In an embodiment, the point cloud encoder can reduce coding complexity by modifying the adjacent node pattern value (e.g., based on a table that modifies 64 to 10 or 6).

[0097] FIG. 8 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.

[0098] As described in Figures 1 to 7, before feature encoding, the coded geometry is reconstructed (restored). When direct coding is applied, the geometry reconstruction operation involves changing the placement of the direct coded points (e.g., placing the direct coded points in front of the point cloud data). When trisoup geometry encoding is applied, the geometry reconstruction process involves the processes of triangulation, upsampling, and voxelization. Since features are dependent on the geometry, feature encoding is performed based on the reconstructed geometry.

[0099] A point cloud encoder (e.g., LOD generator 40009) reorganizes points by LOD. The diagram shows point cloud content corresponding to LOD. The left side of the diagram shows the original point cloud content. The second from the left in the diagram shows the distribution of points with the lowest LOD, and the rightmost side shows the distribution of points with the highest LOD. That is, points with the lowest LOD have a sparse distribution, and points with the highest LOD have a fine distribution. That is, as the LOD increases along the arrow direction shown at the bottom of the diagram, the spacing (or distance) between points becomes shorter.

[0100] FIG. 9 is a diagram illustrating an example of a point configuration for each LOD according to the embodiment.

[0101] As described in Figures 1 to 8, a point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or LOD generator 40009 in Figure 4) generates LOD. LOD is generated by rearranging points into a set of refinement levels according to a set LOD distance value (or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.

[0102] The upper part of Figure 9 shows an example of points (P0 to P9) of point cloud content distributed in 3D space. The original order in Figure 9 shows the order of points P0 to P9 before LOD generation. The LOD-based order in Figure 9 shows the order of points after LOD generation. Points are re-sorted for each LOD. Also, higher LODs include points belonging to lower LODs. As shown in Figure 9, LOD0 includes P0, P5, P4, and P2. LOD1 includes points of LOD0, P1, P6, and P3. LOD2 includes points of LOD0, points of LOD1, and P9, P8, and P7.

[0103] As described in FIG. 4, the point cloud encoder according to the embodiment can selectively or in combination perform predictive transform coding, lift transform coding, and RAHT transform coding.

[0104] The point cloud encoder according to the embodiment generates a predictor for each point and performs predictive transformation coding to set a prediction characteristic (or a prediction characteristic value) for each point. That is, N predictors are generated for N points. The predictor according to the embodiment can calculate a weight (=1 / distance) based on the LOD value of each point, index information for neighboring points within a distance set for each LOD, and the distance value to the neighboring point.

[0105] According to an embodiment, the predicted feature (or feature value) is set as the average value of the feature (or feature value, e.g., hue, reflectance, etc.) of neighboring points set in the predictor of each point multiplied by a weight (or weight value) calculated based on the distance to each neighboring point. A point cloud encoder (e.g., coefficient quantization unit 40011) according to an embodiment may quantize and inverse quantize residuals (also called residual features, residual feature values, feature prediction residual values, etc.) obtained by subtracting the predicted feature (feature value) from the feature (feature value) of each point. The quantization process is shown in the following table.

[0106] Attribute prediction residuals quantization pseudo code

[0107] [Table 2]

[0108] Attribute prediction residuals inverse quantization pseudo Code

[0109] [Table 3]

[0110] The point cloud encoder (e.g., the arithmetic encoder 40012) according to the embodiment performs entropy coding on the quantized and dequantized residual values ​​as described above if there are adjacent points in the predictor of each point. If there are no adjacent points in the predictor of each point, the point cloud encoder (e.g., the arithmetic encoder 40012) according to the embodiment does not perform the above process and performs entropy coding on the characteristics of the corresponding point.

[0111] The point cloud encoder (e.g., lift transform unit 40010) according to the embodiment generates a predictor for each point, sets the calculated LOD in the predictor, registers neighboring points, and sets weights according to the distance to the neighboring points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the above-mentioned point cloud transform coding, but differs in that weights are cumulatively applied to feature values. The process of cumulatively applying weights to feature values ​​according to the embodiment is as follows.

[0112] 1) Create an array QW (QuantizationWeight) that stores the weight value of each point. The initial value of all elements of QW is 1.0. The QW value of the predictor index of the adjacent node registered in the predictor is multiplied by the weight value of the predictor of the current point and added.

[0113] 2) Lift prediction process: To calculate the predicted attribute value, the attribute value of the point is multiplied by the weight and subtracted from the existing attribute value.

[0114] 3) Create temporary arrays called updateweight and update, and initialize the temporary arrays to 0.

[0115] 4) The calculated weights for all predictors are multiplied by the weights stored in the QW corresponding to the predictor index, and the resulting weights are accumulated and summed in the update weight array as the index of the adjacent node. The update array is accumulated and summed by multiplying the calculated weights by the characteristic values ​​of the index of the adjacent node.

[0116] 5) Lift update process: For every predictor, divide the feature value in the update array by the weight value in the update weight array of the predictor index, and add the result to the existing feature value again.

[0117] 6) For all predictors, predicted feature values ​​are calculated by multiplying the feature values ​​updated in the lift update process by the weights updated (stored in the QW) in the lift prediction process. According to an embodiment, a point cloud encoder (e.g., coefficient quantizer 40011) quantizes the predicted feature values. Furthermore, a point cloud encoder (e.g., arithmetic encoder 40012) entropy codes the quantized feature values.

[0118] A point cloud encoder (e.g., the RAHT transform unit 40008) according to an embodiment performs RAHT transform coding, which predicts the characteristics of higher-level nodes using characteristics associated with lower-level nodes in an octree. RAHT transform coding is an example of characteristic intra-coding using octree backward scanning. The point cloud encoder according to an embodiment scans the entire region from a voxel, and at each step, repeats a merging process up to the root node while combining the voxels into larger blocks. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes, but is performed on the nodes immediately above the empty nodes.

[0119] The following equation shows the RAHT transformation matrix: JPEG0007798974000007.jpg1023 denotes the average quality value of the voxels at level l. JPEG0007798974000008.jpg922 JPEG0007798974000009.jpg1031 and Calculated from JPEG0007798974000010.jpg1137. JPEG0007798974000011.jpg1135 and The weighting of JPEG0007798974000012.jpg1130 is JPEG0007798974000013.jpg944 and JPEG0007798974000014.jpg1048.

[0120]

number

[0121] JPEG0007798974000016.jpg826 is a low-pass value that will be used in the merging process at the next higher level. JPEG0007798974000017.jpg1025 is a high-pass coefficient, and the high-pass coefficient at each step is quantized and entropy coded (for example, the encoding of the computational encoder 400012). The weight value is JPEG0007798974000018.jpg1288 is calculated. The root node is the last JPEG0007798974000019.jpg1122 and JPEG0007798974000020.jpg1021 generates the following:

[0122]

number

[0123] The gDC values ​​are also quantized and entropy coded like the high-pass coefficients.

[0124] FIG. 10 is a diagram illustrating an example of a point cloud decoder according to an embodiment.

[0125] The point cloud decoder shown in FIG. 10 is an example of the point cloud video decoder 10006 shown in FIG. 1 and performs operations that are the same as or similar to those of the point cloud video decoder 10006 described in FIG. 1. As shown, the point cloud decoder receives a geometry bitstream and an attribute bitstream included in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and attribute bitstreams and outputs decoded attributes. The decoded geometry and decoded attributes are used to restore point cloud content (decoded point cloud).

[0126] FIG. 11 is a diagram illustrating an example of a point cloud decoder according to an embodiment.

[0127] The point cloud decoder shown in FIG. 11 is an example of the point cloud decoder described in FIG. 10, and performs a decoding operation that is the reverse process of the encoding operation of the point cloud encoder described in FIGS.

[0128] As explained in Figures 1 and 10, the point cloud decoder performs geometry decoding and feature decoding, with geometry decoding occurring before feature decoding.

[0129] The point cloud decoder according to the embodiment includes an arithmetic decoder 11000, an octree synthesis unit 11001, a surface approximation synthesis unit 11002, a geometry reconstruction unit 11003, an inverse transform coordinates unit 11004, an arithmetic decoder 11005, an inverse quantization unit 11006, an RAHT transform unit 11007, a LOD generation unit 11008, an inverse lifting unit 11009 and / or an inverse color transformation unit 11010.

[0130] The arithmetic decoder 11000, octree synthesis unit 11001, surface approximation synthesis unit 11002, geometry reconstruction unit 11003, and coordinate system inverse transformation unit 11004 perform geometry decoding. Geometry decoding according to the embodiment includes direct coding and trisoup geometry decoding. Direct coding and trisoup geometry decoding are selectively applied. Furthermore, geometry decoding is not limited to the above examples and is performed in the reverse process of the geometry encoding described with reference to FIGS. 1 to 9.

[0131] The arithmetic decoder 11000 according to the embodiment decodes the received geometry bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the reverse process of the arithmetic encoder 40004.

[0132] The octree synthesis unit 11001 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or from the decoded result, information on the secured geometry) to generate an octree. Specific details regarding exclusive codes are as described in FIGS. 1 to 9.

[0133] The surface approximation synthesis unit 11002 according to the embodiment synthesizes a surface based on the decoded geometry and / or the generated octree if trisoup geometry encoding is applied.

[0134] According to the embodiment, the geometry reconstruction unit 11003 regenerates geometry based on the surface and / or decoded geometry. As described with reference to FIGS. 1 to 9, direct coding and trisoup geometry coding are selectively applied. Therefore, the geometry reconstruction unit 11003 directly retrieves and adds position information of points to which direct coding is applied. Also, when trisoup geometry coding is applied, the geometry reconstruction unit 11003 reconstructs geometry by performing reconstruction operations of the geometry reconstruction unit 40005, such as triangulation, upsampling, and voxelization operations. The detailed contents are the same as those described with reference to FIG. 6, and therefore will not be repeated. The reconstructed geometry includes a point cloud picture or frame that does not include features.

[0135] The coordinate system inverse transform unit 11004 according to the embodiment transforms the coordinate system based on the reconstructed geometry to obtain the position of the point.

[0136] The arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generation unit 11008, the inverse lift unit 11009, and / or the color inverse transform unit 11010 perform the feature decoding described in FIG. 10. Feature decoding according to the embodiment includes RAHT (Region Adaptive Hierarchical Transform) decoding, Interpolarization-based hierarchical nearest-neighbor prediction-Prediction Transform (Interpolarization-based hierarchical nearest-neighbor prediction with an update / lifting step (Lifting Transform)) decoding. The above three decoding methods may be used selectively, or a combination of one or more of them may be used. Also, feature decoding according to the embodiment is not limited to the above examples.

[0137] The arithmetic decoder 11005 according to the embodiment decodes the attribute bitstream into arithmetic coding.

[0138] The inverse quantization unit 11006 according to the embodiment inverse quantizes the decoded feature bitstream or the feature information obtained as a result of decoding, and outputs the inverse quantized feature (or feature value). The inverse quantization is selectively applied based on the feature encoding of the point cloud encoder.

[0139] In some embodiments, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 process the reconstructed geometry and dequantized features. As described above, the RAHT transform unit 11007, LOD generator 11008, and / or inverse lifting unit 11009 selectively perform decoding operations corresponding to the encoding of the point cloud encoder.

[0140] According to an embodiment, the color inverse transform unit 11010 performs inverse transform coding to inversely transform the color values ​​(or textures) included in the decoded features. The operation of the color inverse transform unit 11010 is selectively performed based on the operation of the color transform unit 40006 of the point cloud encoder.

[0141] The elements of the point cloud decoder of Figure 11 may be embodied in hardware, software, firmware, or a combination thereof, including one or more processors or integrated circuits (not shown) configured to communicate with one or more memories included in the point cloud providing device. The one or more processors perform any of the operations and / or functions of the elements of the point cloud decoder of Figure 11 described above. Additionally, the one or more processors operate or execute a software program and / or set of instructions to perform the operations and / or functions of the elements of the point cloud decoder of Figure 11.

[0142] FIG. 12 shows an example of a transmitting device according to the embodiment.

[0143] The transmitting device shown in Fig. 12 is an example of the transmitting device 10000 of Fig. 1 (or the point cloud encoder of Fig. 4). The transmitting device shown in Fig. 12 performs any of the same or similar operations and methods as the operations and encoding methods of the point cloud encoder described with reference to Figs. 1 to 9. The transmitting device according to the embodiment includes a data input unit 12000, a quantization processing unit 12001, a voxelization processing unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, an arithmetic coder 12006, a metadata processing unit 12007, a hue conversion processing unit 12008, a characteristic conversion processing unit (or attribute conversion processing unit) 12009, a prediction / lift / RAHT conversion processing unit 12010, an arithmetic coder 12011, and / or a transmission processing unit 12012.

[0144] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 performs operations and / or acquisition methods that are the same as or similar to the operations and / or acquisition methods of the point cloud video acquisition unit 10001 (or the acquisition process 20000 shown in FIG. 2).

[0145] Geometry encoding is performed by a data input unit 12000, a quantization unit 12001, a voxelization unit 12002, an octree occupation code generation unit 12003, a surface model processing unit 12004, an intra / inter coding processing unit 12005, and an arithmetic coder 12006. The geometry encoding according to this embodiment is the same as or similar to the geometry encoding described with reference to Figures 1 to 9, and therefore a detailed description thereof will be omitted.

[0146] The quantization unit 12001 according to the embodiment quantizes geometry (e.g., position values ​​of points). The operation and / or quantization of the quantization unit 12001 is the same as or similar to the operation and / or quantization of the quantization unit 40001 shown in Fig. 4. The specific description is as described in Figs. 1 to 9.

[0147] The voxelization processing unit 12002 according to the embodiment voxels the position values ​​of the quantized points. The voxelization processing unit 12002 performs operations and / or processes that are the same as or similar to the operations and / or voxelization processes of the quantization unit 40001 shown in Fig. 4. Specific details are as described in Figs. 1 to 9.

[0148] The octree occupation code generator 12003 according to the embodiment performs octree coding on the positions of voxelized points based on an octree structure. The octree occupation code generator 12003 generates occupation codes. The octree occupation code generator 12003 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (or the octree analyzer 40002) described in FIGS. 4 and 6. Specific descriptions are as described in FIGS. 1 to 9.

[0149] The surface model processing unit 12004 according to the embodiment performs trisoup geometry encoding, which reconstructs the positions of points within a specific region (or node) on a voxel basis based on a surface model. The surface model processing unit 12004 performs operations and / or methods that are the same as or similar to the operations and / or methods of the point cloud encoder (e.g., the surface approximation analysis unit 40003) shown in Figure 4. Specific descriptions are as described in Figures 1 to 9.

[0150] According to the embodiment, the intra / inter coding processor 12005 performs intra / inter coding of the point cloud data. The intra / inter coding processor 12005 performs coding that is the same as or similar to the intra / inter coding described in FIG. 7, as detailed in FIG. 7. In the embodiment, the intra / inter coding processor 12005 is included in the arithmetic coder 12006.

[0151] According to an embodiment, the arithmetic coder 12006 entropy encodes the octree and / or approximated octree of the point cloud data. For example, the encoding method includes an arithmetic encoding method. The arithmetic coder 12006 performs operations and / or methods that are the same as or similar to those of the arithmetic encoder 40004.

[0152] The metadata processing unit 12007 according to the embodiment processes metadata related to point cloud data, such as setting values, and provides the metadata to necessary processing steps such as geometry encoding and / or feature encoding. The metadata processing unit 12007 according to the embodiment also generates and / or processes signaling information related to geometry encoding and / or feature encoding. The signaling information according to the embodiment is encoded separately from the geometry encoding and / or feature encoding. The signaling information according to the embodiment may also be interleaved.

[0153] The hue conversion processor 12008, the feature conversion processor 12009, the prediction / lift / RAHT conversion processor 12010, and the arithmetic coder 12011 perform feature coding. The feature coding according to the embodiment is the same as or similar to the feature coding described with reference to FIGS. 1 to 9, so a detailed description thereof will be omitted.

[0154] The color conversion unit 12008 according to this embodiment performs color conversion coding to convert the hue value included in the feature. The color conversion unit 12008 performs color conversion coding based on the reconstructed geometry. The reconstructed geometry has been described with reference to FIGS. 1 to 9. The color conversion unit 12008 also performs operations and / or methods that are the same as or similar to the operations and / or methods of the color conversion unit 40006 described with reference to FIG. 4. Detailed description thereof will be omitted.

[0155] The feature conversion processor 12009 according to the embodiment performs feature conversion, which converts features based on positions where geometry coding has not been performed and / or reconstructed geometry. The feature conversion processor 12009 performs operations and / or methods identical to or similar to those of the feature conversion unit 40007 described in FIG. 4, detailed descriptions of which will be omitted. The prediction / lift / RAHT conversion processor 12010 according to the embodiment codes the converted features using any one or a combination of RAHT coding, predictive transformation coding, and lift transformation coding. The prediction / lift / RAHT conversion processor 12010 performs any one of operations identical to or similar to those of the RAHT conversion unit 40008, LOD generation unit 40009, and lift transformation unit 40010 described in FIG. 4. Since the predictive transformation coding, lift transformation coding, and RAHT transformation coding have been described with reference to FIGS. 1 to 9, detailed descriptions of these will be omitted.

[0156] The arithmetic coder 12011 according to the embodiment encodes the coded characteristic based on arithmetic coding. The arithmetic coder 12011 performs operations and / or methods that are the same as or similar to the operations and / or methods of the arithmetic encoder 400012.

[0157] The transmission processing unit 12012 according to the embodiment transmits each bitstream including the coded geometry and / or coded attribute and metadata information, or transmits the coded geometry and / or coded attribute and metadata information in one bitstream. When the coded geometry and / or coded attribute and metadata information according to the embodiment is configured in one bitstream, the bitstream includes one or more sub-bitstreams. The bitstream according to the embodiment includes signaling information including a Sequence Parameter Set (SPS) for sequence-level signaling, a Geometry Parameter Set (GPS) for signaling geometry information coding, an Attribute Parameter Set (APS) for signaling attribute information coding, and a Tile Parameter Set (TPS) for tile-level signaling, and slice data. The slice data includes information about one or more slices. According to the embodiment, one slice is one geometry bitstream (Geometry 0). 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0) according to the embodiment. A TPS according to the embodiment includes information about each tile (e.g., coordinate value information and height / size information of a bounding box, etc.) for one or more tiles. A geometry bitstream includes a header and a payload. The header of a geometry bitstream according to the embodiment includes identification information of a parameter set included in the GPS (geom_parameter_set_id), a tile identifier (geom_tile_id), a slice identifier (geom_slice_id), and information about data included in the payload. As described above, the metadata processing unit 12007 according to the embodiment can generate and / or process signaling information and transmit it to the transmission processing unit 12012. In the embodiment, an element that performs geometry coding and an element that performs attribute coding can share data / information with each other, as shown by the dotted lines. The transmission processing unit 12012 according to the embodiment performs an operation and / or a transmission method that is the same as or similar to the operation and / or transmission method of the transmitter 10003. As it is the same as that described with reference to FIGS. 1 and 2, a detailed description thereof will be omitted.

[0158] FIG. 13 shows an example of a receiving device according to the embodiment.

[0159] The receiving device shown in Fig. 13 is an example of the receiving device 10004 in Fig. 1 (or the point cloud decoder in Figs. 10 and 11). The receiving device shown in Fig. 13 performs any of the same or similar operations and methods as the operations and decoding methods of the point cloud decoder described in Figs. 1 to 11.

[0160] The receiving device according to the embodiment includes a receiving unit 13000, a receiving processing unit 13001, an arithmetic decoder 13002, an occupancy code-based octree reconstruction processing unit 13003, a surface model processing unit (triangle reconstruction, upsampling, voxelization) 13004, an inverse quantization processing unit 13005, a metadata analysis 13006, an arithmetic decoder 13007, an inverse quantization processing unit 13008, a prediction / lift / RAHT inverse transform processing unit 13009, a color inverse transform processing unit 13010, and / or a renderer 13011. Each component of the decoding according to the embodiment performs the reverse process of the component of the encoding according to the embodiment.

[0161] The receiving unit 13000 according to the embodiment receives point cloud data. The receiving unit 13000 performs operations and / or a receiving method that are the same as or similar to the operations and / or a receiving method of the receiver 10005 of Fig. 1. Detailed description thereof will be omitted.

[0162] The receiving unit 13001 according to the embodiment obtains a geometry bitstream and / or a characteristic bitstream from the received data. The receiving unit 13000 includes the receiving unit 13000.

[0163] Geometry decoding is performed by an arithmetic decoder 13002, an exclusive code based octree reconstruction processor 13003, a surface model processor 13004, and an inverse quantization processor 13005. The geometry decoding according to this embodiment is the same as or similar to the geometry decoding described with reference to Figures 1 to 10, so a detailed description thereof will be omitted.

[0164] The arithmetic decoder 13002 according to the embodiment decodes the geometry bitstream based on arithmetic coding. The arithmetic decoder 13002 performs operations and / or coding that are the same as or similar to the operations and / or coding of the arithmetic decoder 11000.

[0165] The exclusive code-based octree reconstruction processor 13003 according to the embodiment obtains exclusive codes from the decoded geometry bitstream (or information on the secured geometry as a result of decoding) and reconstructs an octree. The exclusive code-based octree reconstruction processor 13003 performs operations and / or methods identical to or similar to the operations and / or octree generation method of the octree synthesis unit 11001. The surface model processor 13004 according to the embodiment performs trisoup geometry decoding and associated geometry reconstruction (e.g., triangle reconstruction, upsampling, voxelization) based on a surface model scheme when trisoup geometry encoding is applied. The surface model processor 13004 performs operations identical to or similar to the operations of the surface approximation synthesis unit 11002 and / or the geometry reconstruction unit 11003.

[0166] The inverse quantization unit 13005 according to the embodiment inverse quantizes the decoded geometry.

[0167] According to an embodiment, the metadata analysis 13006 analyzes metadata, such as setting values, included in the received point cloud data. The metadata analysis 13006 transmits the metadata to geometry decoding and / or feature decoding. A detailed description of the metadata is omitted here as it has been described with reference to FIG.

[0168] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / lift / RAHT inverse transform processor 13009, and the hue inverse transform processor 13010 perform feature decoding. The feature decoding is the same as or similar to the feature decoding described in Figures 1 and 10, so a detailed description thereof will be omitted.

[0169] The operation decoder 13007 according to the embodiment decodes the attribute bitstream into operation coding. The operation decoder 13007 performs decoding of the attribute bitstream based on the reconstructed geometry. The operation decoder 13007 performs operations and / or coding that are the same as or similar to the operations and / or coding of the operation decoder 11005.

[0170] The inverse quantization unit 13008 according to the embodiment inversely quantizes the decoded characteristic bitstream. The inverse quantization unit 13008 performs the same or similar operations and / or methods as the inverse quantization unit 11006.

[0171] The prediction / lift / RAHT inverse transform processing unit 13009 according to the embodiment processes the reconstructed geometry and dequantized features. The prediction / lift / RAHT inverse transform processing unit 13009 performs any of the same or similar operations and / or decoding as those of the RAHT transform unit 11007, the LOD generation unit 11008, and / or the inverse lift unit 11009. The hue inverse transform processing unit 13010 according to the embodiment performs inverse transform coding to inversely transform the color values ​​(or textures) included in the decoded features. The hue inverse transform processing unit 13010 performs the same or similar operations and / or inverse transform coding as those of the color inverse transform unit 11010. The renderer 13011 according to the embodiment renders point cloud data.

[0172] FIG. 14 is a diagram illustrating an architecture for G-PCC-based point cloud content streaming according to an embodiment.

[0173] The upper part of FIG. 14 illustrates a process in which the transmitting device described in FIGS. 1 to 13 (eg, the transmitting device 10000, the transmitting device in FIG. 12, etc.) processes and transmits point cloud content.

[0174] As described in FIGS. 1 to 13, the transmitting device acquires audio (Ba) of the point cloud content (Audio Acquisition), encodes the acquired audio (Audio Encoding), and outputs an audio bitstream (Ea). The transmitting device also acquires point cloud (Bv) (or point cloud video) of the point cloud content (Point Acqusition), performs point cloud encoding on the acquired point cloud, and outputs a point cloud video bitstream (Eb). The point cloud encoding of the transmitting device is the same as or similar to the point cloud encoding described in FIGS. 1 to 13 (e.g., encoding by the point cloud encoder in FIG. 4), so a detailed description thereof will be omitted.

[0175] The transmitting device encapsulates the generated audio and video bitstreams into files and / or segments (File / segment encapsulation). The encapsulated files and / or segments (Fs, File) include files in file formats such as ISOBMFF or DASH segments. According to an embodiment, point cloud-related metadata is included in the encapsulated file formats and / or segments. The metadata may be included in boxes at various levels in the ISOBMFF file format or in separate tracks within the file. In an embodiment, the transmitting device may encapsulate the metadata itself in a separate file. According to an embodiment, the transmitting device transmits the encapsulated file formats and / or segments via a network. The encapsulation and transmission processing method of the transmitting device is the same as that described in FIGS. 1 to 13 (e.g., transmitter 10003, transmission step 20002 of FIG. 2, etc.), so detailed description thereof will be omitted.

[0176] The lower part of FIG. 14 shows a process in which the receiving device described in FIGS. 1 to 13 (for example, receiving device 10004, receiving device of FIG. 13, etc.) processes and outputs point cloud content.

[0177] In an embodiment, the receiving device includes a device (e.g., loudspeakers, headphones, display) that outputs final audio data and final video data, and a point cloud player that processes point cloud content. The final data output device and the point cloud player are configured as separate physical devices. The point cloud player in this embodiment performs Geometry-based Point Cloud Compression (G-PCC) coding and / or Video-based Point Cloud Compression (V-PCC) coding and / or next-generation coding.

[0178] A receiving device according to an embodiment secures and decapsulates files and / or segments (F', Fs') included in received data (e.g., broadcast signals, signals transmitted over a network, etc.). The receiving and decapsulation methods of the receiving device are the same as those described in Figures 1 to 13 (e.g., receiver 10005, receiving unit 13000, receiving processing unit 13001, etc.), so detailed description thereof will be omitted.

[0179] A receiving device according to an embodiment of the present invention obtains an audio bitstream (E'a) and a video bitstream (E'v) included in a file and / or a segment. As shown in the figure, the receiving device performs audio decoding on the audio bitstream, outputs decoded audio data (B'a), and performs audio rendering on the decoded audio data to output final audio data (A'a) via a speaker or headphones.

[0180] The receiving device also performs point cloud decoding on the video bitstream (E'v) and outputs decoded video data (B'v). The point cloud decoding according to the embodiment is the same as or similar to the point cloud decoding described in Figures 1 to 13 (e.g., the decoding of the point cloud decoder in Figure 11), so a detailed description will be omitted. The receiving device renders the decoded video data and outputs the final video data to a display.

[0181] The receiving device according to the embodiment performs one of decapsulation, audio decoding, audio rendering, point cloud decoding, and rendering operations based on the transmitted metadata. The description of the metadata is omitted here as it is the same as that described in Figures 12 and 13.

[0182] As shown by the dotted lines, a receiving device according to an embodiment (e.g., a point cloud player or a sensing / tracking unit within the point cloud player) generates feedback information (orientation, viewport). The feedback information according to the embodiment is used in the decapsulation, point cloud decoding and / or rendering processes of the receiving device and can also be transmitted to the transmitting device. The description of the feedback information is omitted as it has been described in Figures 1 to 13.

[0183] FIG. 15 is a diagram illustrating an example of a transmitting device according to an embodiment.

[0184] The transmitting device in Fig. 15 is a device for transmitting point cloud content, and corresponds to the examples of the transmitting devices described in Fig. 1 to Fig. 14 (for example, the transmitting device 10000 in Fig. 1, the point cloud encoder in Fig. 4, the transmitting device in Fig. 12, the transmitting device in Fig. 14, etc.). Therefore, the transmitting device in Fig. 15 performs the same or similar operations as the transmitting devices described in Fig. 1 to Fig. 14.

[0185] A transmitting device according to an embodiment may perform one or more of point cloud acquisition, point cloud encoding, file / segment encapsulation, and delivery.

[0186] The point cloud acquisition and transmission operations shown in the figures are the same as those described with reference to FIGS. 1 to 14, so a detailed description thereof will be omitted.

[0187] As described with reference to FIGS. 1 to 14, the transmitting device according to the embodiment performs geometry encoding and attribute encoding. Geometry encoding according to the embodiment is also called geometry compression, and attribute encoding is also called attribute compression. As described above, one point has one geometry and one or more attributes. Therefore, the transmitting device performs attribute encoding for each attribute. The drawings show an example in which the transmitting device performs one or more attribute compressions (Attribute #1 compression, ... Attribute #N compression). The transmitting device according to the embodiment can also perform auxiliary compression. The auxiliary compression is performed on metadata. A description of the metadata is omitted as it has been described with reference to FIGS. 1 to 14. The transmitting device can also perform mesh data compression. Mesh data compression according to the embodiment includes the trisoup geometry encoding described with reference to FIGS. 1 to 14.

[0188] In an embodiment, a sending device encapsulates a bitstream (e.g., a point cloud stream) output by point cloud encoding into files and / or segments. In an embodiment, the sending device performs media track encapsulation to carry data other than metadata (e.g., media data) and metadata track encapsulation to carry metadata. In an embodiment, metadata is encapsulated in a media track.

[0189] As described in Figures 1 to 14, the transmitting device receives feedback information (orientation / viewport metadata) from the receiving device, and performs one of point cloud encoding, file / segment encapsulation, and transmission operations based on the received feedback information. Detailed descriptions are omitted as they are the same as those described in Figures 1 to 14.

[0190] FIG. 16 is a diagram illustrating an example of a receiving device according to an embodiment.

[0191] The receiving device of Figure 16 is a device for receiving point cloud content, and corresponds to an example of the receiving device described in Figures 1 to 14 (e.g., receiving device 10004 of Figure 1, point cloud decoder of Figure 11, receiving device of Figure 13, receiving device of Figure 14, etc.). Therefore, the receiving device of Figure 16 performs the same or similar operations as the receiving devices described in Figures 1 to 14. In addition, the receiving device of Figure 16 can receive signals transmitted by the transmitting device of Figure 15 and perform the reverse process of the operations of the transmitting device of Figure 15.

[0192] A receiving device according to an embodiment performs one or more of delivery, file / segment decapsulation, point cloud decoding, and point cloud rendering.

[0193] The illustrated point cloud receiving and point cloud rendering operations are the same as those described with reference to FIGS. 1 to 14, and therefore detailed description thereof will be omitted.

[0194] As illustrated in Figures 1 to 14, a receiving device according to an embodiment performs decapsulation on files and / or segments obtained from a network or storage device. In an embodiment, the receiving device may perform media track decapsulation, which carries data other than metadata (e.g., media data), and may perform metadata track decapsulation, which carries metadata. In an embodiment, if metadata is encapsulated in a media track, metadata track decapsulation is omitted.

[0195] As illustrated in FIGS. 1 to 14, the receiving device performs geometry decoding and attribute decoding on the bitstream (e.g., point cloud stream) secured by decapsulation. Geometry decoding according to the embodiment is also referred to as geometry decompression, and attribute decoding is also referred to as attribute decompression. As described above, one point has one geometry and one or more attributes, which are encoded separately. Therefore, the receiving device performs attribute decoding for each attribute. The drawings illustrate an example in which the receiving device performs one or more attribute decompressions (Attribute #1 decompression, ..., Attribute #N decompression). The receiving device according to the embodiment can also perform auxiliary decompression. The auxiliary decompression is performed on metadata. The description of the metadata is omitted here, as it has been described with reference to FIGS. 1 to 14. The receiving device also performs mesh data decompression. Mesh data decompression according to the embodiment includes the trisoup geometry decoding described with reference to FIGS. 1 to 14. The receiving device according to the embodiment renders the point cloud data output by point cloud decoding.

[0196] As described in Figures 1 to 14, the receiving device can obtain orientation / viewport metadata using a separate sensing / tracking element, etc., and transmit feedback information including the same to the transmitting device (e.g., the transmitting device in Figure 15). The receiving device can also perform one of a receiving operation, file / segment decapsulation, and point cloud decoding based on the feedback information. Detailed descriptions are omitted here as they are the same as those described in Figures 1 to 14.

[0197] FIG. 17 is a diagram illustrating an example of a structure that can be linked to a method / apparatus for transmitting and receiving point cloud data according to an embodiment.

[0198] 17 illustrates a configuration in which any of a server 1760, a robot 1710, an autonomous vehicle 1720, an XR device 1730, a smartphone 1740, a home appliance 1750, and / or an HMD 1770 are connected to a cloud network 1710. The robot 1710, the autonomous vehicle 1720, the XR device 1730, the smartphone 1740, or the home appliance 1750 may also be referred to as a device. The XR device 1730 may correspond to or be linked to a point cloud data (PCC) device according to an embodiment.

[0199] Cloud network 1700 refers to a network that forms part of a cloud computing infrastructure or exists within a cloud computing infrastructure, where cloud network 1700 is configured using a 3G network, a 4G or LTE network, a 5G network, or the like.

[0200] The server 1760 is connected to any of the robot 1710, autonomous vehicle 1720, XR device 1730, smartphone 1740, home appliance 1750, and / or HMD 1770 via the cloud network 1700 and can assist with at least some of the processing of the connected devices 1710-1770.

[0201] An HMD (Head-Mounted Display) 1770 represents any type of XR device and / or PCC device according to an embodiment. An HMD type device according to an embodiment includes a communications unit, a control unit, a memory unit, an I / O unit, a sensor unit, and a power supply unit.

[0202] Various embodiments of the devices 1710 to 1750 to which the above-described technology is applied will be described below. Here, the devices 1710 to 1750 shown in Fig. 17 can be linked / coupled to the point cloud data transmitting / receiving device according to the above-described embodiments.

[0203] <PCC+XR>

[0204] The XR / PCC device 1730 may be implemented by applying PCC and / or XR (AR+VR) technology to a head-mounted display (HMD), a head-up display (HUD) installed in a vehicle, a TV, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital sign, a vehicle, a fixed robot, a mobile robot, etc.

[0205] The XR / PCC device 1730 can obtain information about the surrounding space or real objects by analyzing 3D point cloud data or image data acquired by various sensors or from an external device to generate position data and attribute data for 3D points, and can render and output the XR object to be output. For example, the XR / PCC device 1730 can output an XR object including additional information about the recognized object corresponding to the recognized object.

[0206] <PCC+Self-propelled+XR>

[0207] The autonomous vehicle 1720 is realized as a mobile robot, vehicle, unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0208] The autonomous vehicle 1720 to which XR / PCC technology is applied refers to an autonomous vehicle equipped with a means for providing XR images, an autonomous vehicle that can be controlled / interacted with within the XR images, etc. In particular, the autonomous vehicle 1720 that can be controlled / interacted with within the XR images can be separated from the XR device 1730 and can be linked to each other.

[0209] The autonomous vehicle 1720, which is equipped with a means for providing XR / PCC images, obtains sensor information from sensors including a camera and outputs XR / PCC images generated based on the obtained sensor information. For example, the autonomous vehicle 1720 may be equipped with a HUD and output XR / PCC images, thereby providing passengers with XR / PCC objects corresponding to real objects or objects on a screen.

[0210] At this time, when the XR / PCC object is output to the HUD, at least a portion of the XR / PCC object is output to overlap with an actual object toward which the passenger's gaze is directed. On the other hand, when the XR / PCC object is output to a display provided in the autonomous vehicle, at least a portion of the XR / PCC object is output to overlap with an object on the screen. For example, the autonomous vehicle 1220 may output XR / PCC objects corresponding to objects such as a roadway, another vehicle, a traffic light, a traffic sign, a motorcycle, a pedestrian, a building, etc.

[0211] The VR (Virtual Reality) technology, AR (Augmented Reality) technology, MR (Mixed Reality) technology, and / or PCC (Point Cloud Compression) technology according to the embodiments can be applied to various devices.

[0212] In other words, VR technology is a display technology that presents real objects and backgrounds only as CG images. In contrast, AR technology is a technology that displays virtual CG images on top of images of real things. MR technology is similar to AR technology in that it mixes virtual objects into the real world. However, AR technology clearly distinguishes between real objects and virtual objects made of CG images, and uses virtual objects to complement real objects, while MR technology is different from AR technology in that virtual objects and real objects are considered to have the same characteristics. More specifically, for example, hologram services are an application of the MR technology.

[0213] However, recently, rather than clearly distinguishing between VR, AR, and MR technologies, they are also referred to as XR (extended reality) technologies. Therefore, embodiments of the present invention can be applied to any of VR, AR, MR, and XR technologies. These technologies apply encoding / decoding based on PCC, V-PCC, and G-PCC technologies.

[0214] The PCC method / apparatus according to the embodiment can be applied to vehicles that provide autonomous driving services.

[0215] Vehicles that provide autonomous driving services are connected to the PCC device via wired or wireless communication.

[0216] When a point cloud data (PCC) transceiver according to an embodiment is connected to a vehicle via wired or wireless communication, it can receive / process content data related to AR / VR / PCC services that can be provided along with an autonomous driving service and transmit the content data to the vehicle. Furthermore, when the point cloud data transceiver is installed in a vehicle, it can receive / process content data related to AR / VR / PCC services according to a user input signal input through a user interface device and provide the content data to a user. According to an embodiment, the vehicle or user interface device receives a user input signal. According to an embodiment, the user input signal includes a signal instructing an autonomous driving service.

[0217] Scalable decoding according to the embodiments is decoding selectively performed by a receiving device on part or all of the geometry and / or attributes depending on the decoding capabilities of the receiving device (e.g., the receiving device 10004 of FIG. 1, the point cloud decoders of FIGS. 10 and 11, and the receiving device of FIG. 13). Parts of the geometry and attributes according to the embodiments are called partial geometry and partial attributes. Scalable decoding applied to geometry according to the embodiments is called scalable geometry decoding or geometry scalable decoding. Scalable decoding applied to attributes according to the embodiments is called scalable attribute decoding or attribute scalable decoding. As described in FIGS. 1 to 17, points in point cloud content are distributed in a three-dimensional space, and the distributed points are represented by an octree structure. The octree structure is an octree structure, and the depth increases from the upper node to the lower node. The depth according to the embodiments is called a level and / or layer. Therefore, a receiving device can perform geometry decoding (or geometry scalable decoding) for partial geometries corresponding to a specific depth or level in the direction from the upper node to the lower node of the octree structure and / or feature decoding (or feature scalable decoding) for partial features corresponding to a specific depth or level in the direction from the upper node to the lower node of the octree structure to provide low-resolution point cloud content. Furthermore, a receiving device can perform geometry and feature decoding corresponding to the entire octree structure to provide high-resolution point cloud content. According to an embodiment, a receiving device performs scalable decoding to provide a scalable point cloud representation or a scalable representation for representing not only the entire point cloud but also a portion of it. The level of a scalable representation according to an embodiment corresponds to the depth of the octree structure. According to an embodiment, the resolution or detail increases as the level value increases.

[0218] Scalable decoding supports receiving devices with various capabilities and enables receiving devices to provide point cloud services even in adaptive bitrate environments. However, because feature decoding is based on geometry decoding, accurate feature decoding requires geometry information. For example, transform coefficients for RAHT coding are determined based on geometry distribution information (or geometry structure information, such as an octree structure). Furthermore, predictive transform coding and lift transform coding require overall geometry distribution information (or geometry structure information, such as an octree structure) to determine points belonging to each LOD.

[0219] Therefore, the receiving device can receive and process all the geometry to perform stable quality decoding. However, transmitting and receiving geometry information that is not actually displayed depending on the performance of the receiving device is inefficient in terms of bit rate. In addition, if the receiving device decodes all the geometry, delays may occur in providing point cloud content services. Furthermore, if the decoder performance of the receiving device is low, it may not be able to decode all the geometry.

[0220] FIG. 18 shows an arrangement for transmitting and receiving point cloud data according to an embodiment.

[0221] FIG. 18 shows an arrangement for transmitting and receiving point cloud data for scalable decoding and representation.

[0222] The example 1800 shown at the top of FIG. 18 is an example of a configuration for transmitting and receiving a partial PCC bitstream. A transmitting device according to an embodiment (e.g., the transmitting device 10000 described in FIG. 1 or the transmitting device described in FIG. 12) does not perform scalable encoding on source geometry and source attribute, but stores them after performing full encoding. That is, the transmitting device cannot encode part of the geometry and attribute, and therefore cannot selectively transmit part of the data required by the receiving device. Therefore, the transmitting device according to an embodiment decodes the stored encoded geometry and attribute, and performs transcoding for partial encoding on some or all of the decoded geometry and attribute again to generate a bitstream (e.g., a partial PCC bitstream) including the partial geometry and attribute 1801. A receiving device according to an embodiment (e.g., receiving device 10004 of FIG. 1, point cloud decoders of FIGS. 10 and 11, receiving device of FIG. 13) receives and decodes the bitstream to output the part geometry and part characteristics 1802. The configuration shown in this example 1800 requires further data processing (e.g., decoding and transcoding) at the transmitting device, which may cause delays in the data processing process.

[0223] An example 1810 shown at the bottom of FIG. 18 is an example of a configuration for transmitting and receiving a complete PCC bitstream. A transmitting device according to an embodiment (e.g., the transmitting device 10000 described in FIG. 1 or the transmitting device described in FIG. 12) performs full encoding without scalable encoding and stores the encoded geometry and attributes. The transmitting device according to the embodiment performs transcoding for low-quality point cloud content on a portion of the stored geometry and attributes to generate a bitstream (e.g., a complete PCC bitstream) including the entire geometry and attributes 1811. The bitstream according to the embodiment includes signaling information for scalable decoding by a receiving device. A receiving device according to the embodiment receives the bitstream and performs scalable decoding to output the source geometry and source attributes, or selects a portion of the decoded data and outputs a partial geometry and partial attributes 1812. The configuration shown in this example 1810 requires transmitting unnecessary data (e.g., a portion of the unselected geometry and attributes), resulting in reduced bandwidth efficiency. Also, when using a fixed bandwidth, the transmitter may transmit the encoded point cloud data at a reduced quality.

[0224] FIG. 19 shows an octree structure and bitstream of point cloud data according to an embodiment.

[0225] According to an embodiment, points in a point cloud content are distributed in a three-dimensional space, and the distributed points are represented by an octree. As described in FIG. 6, the octree is generated by recursively subdividing a bounding box. The octree has an occupancy code including nodes corresponding to the recursively divided regions. The nodes according to the embodiment correspond to spaces or regions generated by the recursive subdivision. Therefore, the nodes according to the embodiment have a value of 0 or 1 depending on whether one or more points exist within the corresponding region. For example, if one or more points exist within the region corresponding to the node, the value of the corresponding node is 1, and if one or more points do not exist, the value of the corresponding node is 0.

[0226] The example 1900 shown in FIG. 19 illustrates an octree structure. The top node of the octree according to the embodiment is called a root node, and the bottom node is called a leaf node. Because the amount of data or data density increases from the top node to the bottom node due to the recursive division described above, the octree structure according to the embodiment is expressed in a triangular shape. The root node according to the embodiment corresponds to the first depth or the lowest level (e.g., level 0). The leaf node according to the embodiment corresponds to the final depth or the highest level (e.g., level n). The illustrated octree has a top level of 7.

[0227] As illustrated in FIG. 8, a point cloud encoder according to an embodiment (e.g., the point cloud video encoder 10002 of FIG. 1, the point cloud encoder of FIG. 4, or the point cloud encoders illustrated in FIGS. 12, 14, and 15) classifies points of a point cloud into one or more Levels of Detail (LOD). The LODs according to the embodiment are used for feature encoding. The LODs according to the embodiment correspond to levels of an octree structure. One LOD corresponds to one level of the octree structure, or one LOD corresponds to one or more levels of the octree structure. As illustrated in example 1900, LOD0 corresponds to levels 0 to 3 (or depths 0 to 3) of the octree structure, LOD1 corresponds to levels 0 to 5 (or depths 0 to 5), and LOD2 corresponds to levels 0 to the highest level 7 (or depths 0 to 7). The relationship between LODs and levels (or depths) of the octree is not limited to the above example.

[0228] Example 1910 shown at the bottom of Figure 19 illustrates a geometry bitstream and an attribute bitstream. A transmitting device according to an embodiment (e.g., transmitting device 10000 described in Figure 1 or transmitting device described in Figure 12) generates and transmits a bitstream including a geometry bitstream and an attribute bitstream (e.g., the bitstream described in Figure 10). Furthermore, the transmitting device configures and transmits each of the geometry bitstream and the attribute bitstream in slices, regardless of the octree structure or LOD. Therefore, as shown in example 1910, the geometry bitstream includes geometries corresponding to LOD 0 to LOD N (the highest level N, e.g., 7). The attribute bitstream includes attributes corresponding to LOD 0 to LOD N. That is, in order for the transmitting device according to the embodiment to transmit a partial bitstream as shown in example 1800 of Figure 18, it is necessary to decode each of the geometry bitstream and the attribute bitstream, select the partial geometry and partial attribute required for transmission, and re-encode them.

[0229] Furthermore, when a receiving device (e.g., receiving device 10004 in Figure 1, the point cloud decoder in Figures 10 and 11, or the receiving device in Figure 13) receives the bitstream shown in example 1910, as described in example 1810 in Figure 18, it decodes the entire bitstream, selects only the necessary partial data, and outputs partial geometry and partial characteristics (e.g., operation 1812 of the receiving device described in Figure 18).

[0230] Therefore, a transmitting device (or point cloud encoder) according to an embodiment transmits a geometry bitstream and / or a feature bitstream in layers to reduce unnecessary data transmission and shorten data processing time. Point cloud data according to an embodiment has a layer structure based on various parameters and variables, such as SNR (Signal to Noise Ratio), spatial resolution, color, temporal frequency, and bit depth. The layer structure of point cloud data according to an embodiment is based on the depth (level) of the octree structure and / or the LOD level described above. For example, a layer of point cloud data corresponds to each depth or one or more depths of the octree structure. A layer of point cloud data corresponds to each level or one or more levels of LOD. The layer of a geometry bitstream according to an embodiment may be the same as or different from the layer of a feature bitstream. For example, if the layer of the geometry bitstream is 3, the layer of the feature bitstream is 3. If the layer of the geometry bitstream is 3, the layer of the feature bitstream is 2 or 4.

[0231] FIG. 20 illustrates an example of a point cloud data processing device according to an embodiment.

[0232] The point cloud data processing device 2000 according to the embodiment shown in Fig. 20 is an example of the transmitting device 10000 described in Fig. 1 or the transmitting device described in Fig. 12. The point cloud data processing device 2000 performs operations that are the same as or similar to those of the transmitting device described in Fig. 1 to Fig. 17. Although not shown in Fig. 20, the point cloud data processing device 2000 further includes one or more elements for performing the encoding operations described in Fig. 1 to Fig. 17.

[0233] The point cloud data processing device 2000 includes a geometry encoder 2010, an attribute encoder 2020, a sub-bitstream generator 2030, a metadata generator 2040, a multiplexer (Mux) 2050, and a transmitter (Transmitter) 2060.

[0234] Point cloud data (or Point Cloud Compression (PCC) data) according to an embodiment includes geometry and / or features as input data for the point cloud encoder 1800. Geometry according to an embodiment is information indicating the location (e.g., position) of a point, and is expressed by parameters in a coordinate system such as a Cartesian coordinate system, a cylindrical coordinate system, or a spherical coordinate system. Features according to an embodiment indicate the features of a point (e.g., color, transparency, reflectivity, grayscale, etc.). Geometry is also referred to as geometry information (or geometry data), and features are also referred to as feature information (or feature data).

[0235] The geometry encoder 2010 performs geometry coding (or geometry encoding) described with reference to Figures 1 to 17 and outputs a geometry bitstream. The operation of the geometry encoder 2010 is the same as or similar to the operations of the coordinate system conversion unit 40000, quantization unit 40001, octree analysis unit 40002, surface approximation analysis unit 40003, arithmetic encoder 40004, and geometry reconstruction unit 40005 described with reference to Figure 4. In addition, the operation of the geometry encoder 2010 is the same as or similar to the operations of the data input unit 12100, quantization unit 12001, voxelization unit 12002, octree occupation code generation unit 12003, surface model processing unit 12004, intra / inter coding processing unit 12005, arithmetic coder 12006, and metadata processing unit 12007 described with reference to Figure 12.

[0236] The feature encoder 2020 performs the feature coding (or feature encoding) described in Figures 1 to 17 and outputs a feature bitstream. The feature encoder 2020 operates as follows:

[0237] The operations of the feature encoder 2020 are the same as or similar to those of the geometry reconstruction unit 40005, hue conversion unit 40006, attribute conversion unit 40007, RAHT conversion unit 40008, LOD generation unit 40009, lift conversion unit 40010, coefficient quantization unit 40011, and / or arithmetic encoder 40012 described in Fig. 4. In addition, the operations of the feature encoder 2020 are the same as or similar to those of the hue conversion processing unit 12008, attribute conversion processing unit 12009, prediction / lift / RAHT conversion processing unit 12110, and arithmetic coder 12011 described in Fig. 12.

[0238] The sub-bitstream generator 2030 according to the embodiment receives a geometry bitstream and an attribute bitstream and layers the geometry bitstream and the attribute bitstream in units of layers to generate one or more sub-bitstreams. The sub-bitstream according to the embodiment includes a geometry sub-bitstream and an attribute sub-bitstream. As described in FIG. 19 , the layer structure according to the embodiment is based on an octree structure and / or LOD. The sub-bitstream generator 2030 rearranges the sorting order of the sub-bitstreams corresponding to each layer and outputs one or more sub-bitstreams for each sorted layer (e.g., a geometry sub-bitstream and an attribute sub-bitstream corresponding to the same layer). Layer division or layering structure information for the geometry bitstream and the attribute bitstream according to the embodiment is transmitted to the metadata generator 2040.

[0239] The metadata generator 2040 generates and / or processes signaling information related to geometry coding by the geometry encoder 2010, feature coding by the feature encoder 2020, and layer structure by the sub-bitstream generator 2030. The operation of the metadata generator 2020 is the same as or similar to the operation of the metadata processing unit 12007 described in FIG.

[0240] The multiplexer 2050 multiplexes and outputs one or more sub-bitstreams and the parameters output by the metadata generator 2040. The transmitter 2060 according to the embodiment transmits the data output by the multiplexer 2050 to a receiving device. The transmitter 2060 is an example of the transmitter 10003 described in FIG. 1, and performs the same or similar operations as the transmitter 10003.

[0241] FIG. 21 is an example of layers of geometry bitstreams and attribute bitstreams.

[0242] As described in Figures 18 to 20, a point cloud data processing device according to an embodiment (e.g., the point cloud data processing device 2000 or the sub-bitstream generator 2040 in Figure 20) generates one or more sub-bitstreams by layering the geometry bitstream and the attribute bitstream based on layers. Example 2100 in Figure 21 shows a layer structure of the geometry bitstream and the attribute bitstream based on LOD.

[0243] As described in FIG. 9, points included in a lower level of detail are included in a higher level of detail. As described in FIG. 19, the geometry bitstream 2110 corresponds to the entire level of detail. Therefore, the point cloud data processing device generates one or more sub-bitstreams by layering the geometry bitstream based on geometry information included only in the highest level of detail. The point cloud data processing device layers the geometry bitstream 2110 to generate a first geometry sub-bitstream 2111 including a first geometry corresponding to LOD0, LOD1, and LOD2, a second geometry sub-bitstream 2112 including a second geometry R1 corresponding to LOD1 and LOD2, and a third geometry sub-bitstream 2113 including a third geometry R2 corresponding only to LOD2. Each geometry sub-bitstream includes a header (grayed box in the figure) and a payload.

[0244] As described in Figure 19, the feature bitstream 2120 corresponds to the entire LOD. Therefore, the point cloud data processing device generates one or more sub-bitstreams by layering the feature bitstream based on feature information included only in the highest LOD. The point cloud data processing device layers the feature bitstream 2120 to generate a first feature sub-bitstream 2121 including a first feature corresponding to LOD0, LOD1, and LOD2, a second feature sub-bitstream 2122 including a second feature R1 corresponding to LOD1 and LOD2, and a third feature sub-bitstream 2123 including a third feature R2 corresponding only to LOD2. Each feature sub-bitstream includes a header (grayed box in the figure) and a payload.

[0245] The point cloud data processing device according to the embodiment can change the sort order of one or more sub-bitstreams.

[0246] FIG. 22 shows an example of a method for aligning sub-bitstreams according to an embodiment.

[0247] Example 2200 of FIG. 22 shows serially aligned geometry sub-bitstreams and attribute sub-bitstreams.

[0248] A point cloud data processing device according to an embodiment (e.g., the transmitting device 10000 described in FIG. 1 or the transmitting device described in FIG. 12) first aligns one or more geometry sub-bitstreams 2201 (e.g., a first geometry sub-bitstream 2111 including a first geometry corresponding to LOD0, LOD1, and LOD2, a second geometry sub-bitstream 2112 including a second geometry R1 corresponding to LOD1 and LOD2, and a third geometry sub-bitstream 2113 including a third geometry R2 corresponding only to LOD2, as described in FIG. 21), and then aligns and transmits one or more attribute sub-bitstreams (e.g., a first attribute sub-bitstream 2121 including a first attribute corresponding to LOD0, LOD1, and LOD2, a second attribute sub-bitstream 2122 including a second attribute R1 corresponding to LOD1 and LOD2, and a third attribute sub-bitstream 2123 including a third attribute R2 corresponding only to LOD2, as described in FIG. 21). A receiving device according to an embodiment (e.g., receiving device 10004 in Figure 1, the point cloud decoder in Figures 10 and 11, or the receiving device in Figure 13) first reconstructs the geometry sub-bitstream and then reconstructs the feature sub-bitstream based on the reconstructed geometry (or geometry information).

[0249] Example 2210 in Figure 22 shows geometry sub-bitstreams and attribute sub-bitstreams aligned in parallel. According to an embodiment, a point cloud data processing device aligns and transmits sub-bitstreams corresponding to the same layer together. The point cloud data processing device first aligns 2211 a first geometry sub-bitstream (e.g., first geometry sub-bitstream 2111 in Figure 21 ) including a first geometry corresponding to LOD0, LOD1, and LOD2 and a first attribute sub-bitstream (e.g., first attribute sub-bitstream 2121 in Figure 21 ) including a first attribute corresponding to LOD0, LOD1, and LOD2. The point cloud data processing device then aligns 2212 a second geometry sub-bitstream (e.g., second geometry sub-bitstream 2112 in Figure 21 ) including a second geometry R1 corresponding to LOD1 and LOD2 and a second attribute sub-bitstream (e.g., second attribute sub-bitstream 2122 in Figure 21 ) including a second attribute R1 corresponding to LOD1 and LOD2. The point cloud data processing device aligns 2213 a third geometry sub-bitstream (e.g., third geometry sub-bitstream 2113 in FIG. 21 ) including a third geometry R2 corresponding only to LOD2 and a third attribute sub-bitstream (e.g., third attribute sub-bitstream 2123 in FIG. 21 ) including a third attribute R2 corresponding only to LOD2. Therefore, the receiving device can perform geometry decoding and attribute decoding in parallel to reduce the decoding execution time. In addition, since the attribute decoding is performed based on the geometry decoding, the receiving device can process the geometry corresponding to a smaller layer (LOD0) and the attributes of the same layer.

[0250] A point cloud data processing device (for example, the point cloud data processing device 2000 described in FIG. 20) multiplexes the geometry sub-bitstream and the attribute sub-bitstream described in FIGS. 21 and 22 and transmits them in the form of a bitstream.

[0251] The point cloud data transmitting device divides a point cloud data image into one or more packets to account for transmission channel errors and transmits the divided image over a network. According to an embodiment, a bitstream includes one or more packets (e.g., Network Abstraction Layer (NAL) units). Therefore, even if some packets are lost in a poor network environment, the point cloud data receiving device can restore the corresponding image using the remaining packets. Point cloud data can be divided into one or more slices or one or more tiles for processing. According to an embodiment, tiles and slices are regions for partitioning a picture of point cloud data and performing point cloud compression coding processing. The point cloud data transmitting device processes data corresponding to each divided region of the point cloud data according to the importance of each region, thereby providing high-quality point cloud content. That is, the point cloud data transmitting device according to an embodiment can perform point cloud compression coding processing with better compression efficiency and appropriate delay for data corresponding to regions important to the user.

[0252] According to the embodiment, an image (or picture) of point cloud content is divided into basic processing units for point cloud compression coding. The basic processing units for point cloud compression coding according to the embodiment include, but are not limited to, a coding tree unit (CTU) and a brick.

[0253] According to the embodiment, a slice is an area including one or more integer basic processing units for point cloud compression coding, and is not rectangular. According to the embodiment, a slice includes data transmitted by packets. According to the embodiment, a tile is an area divided into rectangular shapes in an image, and includes one or more basic processing units for point cloud compression coding. According to the embodiment, one slice is included in one or more tiles. According to the embodiment, one tile is included in one or more slices.

[0254] The bitstream according to the embodiment includes signaling information including an SPS (Sequence Parameter Set) for sequence-level signaling, a GPS (Geometry Parameter Set) for signaling geometry information coding, an APS (Attribute Parameter Set) for signaling attribute information coding, and a TPS (Tile Parameter Set) for tile-level signaling, and one or more slices.

[0255] The SPS according to the embodiment includes coding information for the entire sequence, such as profile and level, as well as comprehensive information for the entire file, such as picture resolution and video format.

[0256] According to the embodiment, one slice includes a slice header and slice data. The slice data includes one geometry bitstream (Geom0 0 ) and one or more attribute bitstreams (Attr0 0 , Attr1 0). The geometry bitstream includes a header (e.g., geometry slice header) and a payload (e.g., geometry slice data). The header of the geometry bitstream according to the embodiment includes identification information of the parameter set included in the GPS (geom_geom_parameter_set_id), a tile identifier (geom_tile id), a slice identifier (geom_slice_id), and information about the data included in the payload. The attribute bitstream includes a header (e.g., attribute slice header or attribute brick header) and a payload (e.g., attribute slice data or attribute brick data).

[0257] As described in Figures 18 to 22, a point cloud data processing device according to an embodiment (e.g., the transmitting device 10000 described in Figure 1 or the transmitting device described in Figure 12) generates layering structure information regarding layer division or layering of a geometry bitstream and a feature bitstream for generating sub-bitstreams. The point cloud data processing device according to an embodiment transmits a part (e.g., the geometry sub-bitstream and the feature sub-bitstream described in Figures 20 to 22) or the whole of the geometry bitstream and the feature bitstream for scalable representation. Furthermore, the layer-divided geometry bitstream according to the embodiment, i.e., the geometry sub-bitstream, is included in a slice. A slice including the geometry sub-bitstream according to the embodiment is called a geometry slice. The layer-divided feature bitstream according to the embodiment, i.e., the feature sub-bitstream, is included in a slice (feature slice). A slice including the feature sub-bitstream according to the embodiment is called a feature slice.

[0258] FIG. 23 shows an example of an SPS according to an embodiment.

[0259] FIG. 23 is an example of syntax for an SPS according to an embodiment, which includes the following information (fields, parameters, etc.):

[0260] profile_idc indicates the profile of the bitstream. This information has a specific value. Values ​​other than the specific value are reserved for future use.

[0261] profile_compatibility_flags indicates whether the bitstream is a bitstream according to the profile indicated by the profile_idc value (e.g., j). If the value of profile_compatibility_flags is 1, the bitstream to which SPS applies is a bitstream according to the profile indicated by the profile_idc value (e.g., j). If the value of profile_compatibility_flags is 0, the bitstream to which SPS applies is a bitstream according to the profile indicated by a value that is not an allowed profile_idc value.

[0262] level_idc indicates the level of the bitstream. This information has a specific value. Values ​​other than the specific value are reserved for future use.

[0263] sps_bounding_box_present_flag indicates whether the source bounding box offset and size information are signaled in the SPS. According to this embodiment, the source bounding box is a box in 3D space (e.g., the bounding box described in FIG. 5) expressed in a 3D coordinate system (e.g., X, Y, Z axes) and contains the points of the point cloud data. When the value of sps_bounding_box_present_flag is 1, sps_bounding_box_present_flag indicates that the source bounding box offset and size information are signaled in the SPS. When the value of sps_bounding_box_present_flag is 0, sps_bounding_box_present_flag indicates that the source bounding box offset and size information are not signaled in the SPS.

[0264] The following is information related to the source bounding box when sps_bounding_box_present_flag has a value of 1:

[0265] sps_bounding_box_offset_x indicates the x offset of the source bounding box in cartesian coordinates. If no sps_bounding_box_offset_x information is present, a value of 0 is inferred for sps_bounding_box_offset_x.

[0266] sps_bounding_box_offset_y indicates the y offset of the source bounding box in cartesian coordinates. If sps_bounding_box_offset_y information is not present, a value of 0 is inferred for sps_bounding_box_offset_y.

[0267] sps_bounding_box_offset_z indicates the z offset of the source bounding box in Cartesian coordinates. If no sps_bounding_box_offset_z information is present, a value of 0 is inferred for sps_bounding_box_offset_z.

[0268] sps_bounding_box_scale_factor indicates the scale factor of the source bounding box in the Cartesian coordinate system. If sps_bounding_box_scale_factor is not present, a value of 1 is inferred for sps_bounding_box_scale_factor.

[0269] sps_bounding_box_size_width indicates the width of the source bounding box in Cartesian coordinates. If sps_bounding_box_size_width is not present, a value of 1 is inferred for sps_bounding_box_size_width.

[0270] sps_bounding_box_size_height indicates the height of the source bounding box in Cartesian coordinates. If sps_bounding_box_size_height is not present, a value of 1 is inferred for sps_bounding_box_size_height.

[0271] sps_bounding_box_size_depth indicates the depth of the source bounding box in the Cartesian coordinate system. If sps_bounding_box_size_depth is not present, a value of 1 is inferred for sps_bounding_box_size_depth.

[0272] sps_source_scale_factor indicates the scale factor of the source point cloud (point cloud data). Depending on the embodiment, the scale factor is expressed as a floating point or integer.

[0273] sps_seq_parameter_set_id provides an identifier for identifying the corresponding SPS. This information is provided for other syntaxes (e.g., GPS, APS) that reference the SPS. Within the scope of this embodiment, the value of sps_seq_parameter_set_id is 0. Other values ​​other than 0 are reserved for future use.

[0274] sps_num_attribute_sets indicates the number of coded attributes in the bitstream. The value of sps_num_attribute_sets according to the embodiment is in the range of 0 to 63. Below is information about each attribute for one or more coded attributes indicated by SPS_num_attribute_sets. The "i" in the figure indicates each attribute.

[0275] attribute_dimension[i] indicates the number of components of the i-th attribute. Depending on the embodiment, the attribute may indicate reflectance, hue, etc. Therefore, the number of components that an attribute has may vary. For example, an attribute corresponding to hue may have three hue components (e.g., RGB). Therefore, an attribute corresponding to reflectance may be a mono-dimensional attribute, and an attribute corresponding to hue may be a three-dimensional attribute.

[0276] attribute_instance_id[i] indicates the instance ID (instance_id) of the i-th attribute.

[0277] attribute_bitdepth[i] indicates the bit depth of the i-th attribute.

[0278] attribute_cicp_colour_primaries[i] indicates the chromaticity coordinates of the colour attribute source primaries (e.g. white and black) of the i-th attribute.

[0279] attribute_cicp_transfer_characteristics[i] indicates the OTF (Optical-electro Transfer characteristic Function) or inverse OTF of the i-th characteristic (hue characteristic). In this embodiment, the OTF is a function of the Source Input linear optical intensity (Lc) in a nominal real-valued range from 0 to 1. In this embodiment, the inverse OTF is a function of the Output linear optical intensity (Lo) in a nominal real-valued range from 0 to 1.

[0280] attribute_cicp_matrix_coeffs[i] indicates the coefficients of the matrix used to derive the luma and chroma signals from the green, blue and red or Y, Z and X primaries of the i-th attribute.

[0281] attribute_cicp_video_full_range_flag[i] indicates, for the i-th attribute, the black level and range of the extracted luma and chroma signals, E'' and E', or E'' and E' component signals (real-valued component signals).

[0282] known_attribute_label_flag[i] indicates whether known_attribute_label is signaled for the i-th attribute. If known_attribute_label_flag[i] has a value of 1, known_attribute_label_flag[i] indicates that known_attribute_label is signaled for the i-th attribute. If known_attribute_label_flag[i] has a value of 0, known_attribute_label_flag[i] indicates that attribute_label_four_bytes[i] is signaled for the i-th attribute.

[0283] The following is known_attribute_label_flag[i] signaled when the value of known_attribute_label_flag[i] is 1.

[0284] known_attribute_label[i] has a value between 0 and 2.

[0285] If the value of known_attribute_label_flag[i] is 0, the attribute is hue or color.

[0286] If the value of known_attribute_label_flag[i] is 1, the attribute is reflectance.

[0287] If the value of known_attribute_label_flag[i] is 2, the attribute is a frame index.

[0288] The following is attribute_label_four_bytes[i] signaled when the value of known_attribute_label_flag[i] is 0.

[0289] attribute_label_four_bytes[i] is a 4-byte code indicating the known attribute type. If the value of attribute_label_four_bytes[i] is 0, the attribute type is hue or color. If the value of attribute_label_four_bytes[i] is 1, the attribute type is reflectance.

[0290] Below is some information related to slicing.

[0291] split_slice_flag indicates whether the point cloud data is split into one or more slices. If the value of split_slice_flag is 1, split_slice_flag indicates whether the point cloud data (geometry bitstream, attribute bitstream) is split into one or more slices. If the value of split_slice_flag is 0, split_slice_flag indicates that the point cloud data (geometry bitstream, attribute bitstream) is contained in one slice. If the value of split_slice_flag is 1, the SPS includes split_type.

[0292] The split_type field indicates the method or type of splitting of point cloud data into one or more slices, and the layering method or type of the geometry bitstream and the attribute bitstream. For example, if the value of split_type is 0, split_type indicates that the point cloud data is split into one or more slices based on the LOD (e.g., the geometry sub-bitstream or the attribute sub-bitstream described in FIGS. 20 and 21). For example, as described in FIGS. 21 and 22, point cloud data split based on the maximum LOD (e.g., the geometry sub-bitstream 2113 and the attribute sub-bitstream 2123) are included in one slice. If the value of split_type is 1, split_type indicates that the point cloud data is split into one or more slices based on the level of the octree structure (e.g., the geometry sub-bitstream or the attribute sub-bitstream described in FIGS. 20 and 21). For example, one slice includes point cloud data sampled based on the octree level (e.g., point cloud data sampled based on the level of the octree structure that has been colored for matching attributes). That is, the geometry sub-bitstreams according to the embodiment correspond to octree levels.

[0293] If the value of split_type is 0, the SPS contains the following information:

[0294] Num_LOD indicates the number of LODs. As mentioned above, the point cloud data is divided into one or more slices depending on the number of LODs.

[0295] The full_res_flag indicates whether point cloud data corresponding to the entire LOD or a partial LOD is to be transmitted. When the value of the full_res_flag is 1, the full_res_flag indicates that the geometry and attributes corresponding to the entire LOD are to be transmitted. That is, when the value of the full_res_flag is 1, the full_res_flag indicates that information constituting the entire point cloud data is to be transmitted. When the value of the full_res_flag is 0, the full_res_flag indicates that the geometry and attributes corresponding to a partial LOD are to be transmitted. Therefore, based on this information, the receiving device provides point cloud content of various resolutions as described in FIGS. 18 and 19.

[0296] If the value of full_res_flag is 0, the SPS contains the following information:

[0297] full_geo_present_flag indicates whether the entire geometry is to be transmitted. If the value of full_geo_present_flag is 1, full_geo_present_flag indicates that the entire geometry is to be transmitted (regardless of whether the entire attribute is transmitted). Therefore, even if a receiving device receives attributes corresponding to a partial LOD, it can restore the entire geometry and restore the attributes corresponding to the entire LOD. If the value of full_geo_present_flag is 0, full_geo_present_flag indicates that a partial geometry is to be transmitted. The partial geometry corresponds to the same LOD as the transmitted partial attribute. If the partial geometry corresponds to a different LOD than the transmitted partial attribute, information about the LOD corresponding to the partial geometry and the LOD corresponding to the partial attribute is signaled separately.

[0298] split_info_present_in_slice_header_flag indicates whether additional information is sent in the slice header. If the value of split_info_present_in_slice_header_flag is 1, split_info_present_in_slice_header_flag indicates that additional information is sent in the slice header. If the value of split_info_present_in_slice_header_flag is 0, split_info_present_in_slice_header_flag indicates that additional information is not sent in the slice header.

[0299] sps_extension_present_flag indicates whether the sps_extension_data syntax structure exists within the SPS syntax structure. If the value of sps_extension_present_flag is 0, sps_extension_present_flag indicates that the sps_extension_data syntax structure does not exist. If the value of sps_extension_present_flag is 1, sps_extension_present_flag indicates that the SPS_extension_data syntax structure exists.

[0300] sps_extension_data_flag has any value. The presence and value of this information does not affect the decoding performance of the receiving device.

[0301] The syntax for the SPS according to the embodiment shown in FIG. 23 is not limited to the above examples, and may further include additional information (or fields, parameters, etc.) not shown.

[0302] FIG. 24 is an example of syntax for a geometry slice bitstream according to an embodiment.

[0303] FIG. 24 is an example of syntax for a geometry bitstream (or geometry slice bitstream) corresponding to one slice when point cloud data is divided into one or more slices.

[0304] 24 shows an example of a syntax for a geometry slice bitstream according to an embodiment. The geometry slice bitstream includes a geometry slice header (geometry_slice_header) and geometry slice data (geometry_slice_data).

[0305] The second syntax 2410 shown in Figure 24 is an example of a syntax for a geometry slice header according to an embodiment. The syntax for a geometry slice header includes the following information (or fields, parameters, etc.):

[0306] The gsh_geometry_parameter_set_id indicates the identifier or value of the active GPS.

[0307] The gsh_tile_id indicates the identifier or identifier value of the tile to which the geometry slice header refers. The value of gsh_tile_id in an embodiment ranges from 0 to any value.

[0308] The gsh_slice_id identifies the slice header referenced by other syntax. The value of gsh_slice_id ranges from 0 to any value.

[0309] The independent_decodable_flag indicates whether this slice can be independently decoded. If the value of the independent_decodable_flag is 1, the independent_decodable_flag indicates that this slice can be independently decoded. If the value of the independent_decodable_flag is 0, the independent_decodable_flag indicates that this slice cannot be independently decoded. Therefore, a receiving device (e.g., the receiving device 10004 in FIG. 1, the point cloud decoders in FIGS. 10 and 11, and the receiving device in FIG. 13) decodes this slice based on other slices.

[0310] 1 to 22, the higher the LOD level, the lower the point cloud data of the LOD level must be referenced. Therefore, when the value of split_type described in Fig. 23 is 0, the value of independent_decodable_flag of the slice corresponding to LOD0 is 1, and the values ​​of independent_decodable_flag of the remaining slices corresponding to LOD1 to LOD N are 0.

[0311] The following shows the relevant information when the value of independent_decodable_flag is 1:

[0312] gps_box_present_flag indicates whether a bounding box for this slice exists. If the value of gps_box_present_flag is 1, then gps_box_present_flag indicates that a bounding box exists for this slice. If the value of gps_box_present_flag is 1, then the syntax for the geometry slice header further includes the following information:

[0313] gps_gsh_box_log2_scale_present_flag indicates whether the original scale of the bounding box is the same as gsh_box_log2_scale. In the embodiment, gsh_box_log2_scale indicates the scale factor applied to each slice. In the embodiment, GPS_gsh_box_log2_scale indicates the scale factor applied commonly to all slices.

[0314] If the value of gps_gsh_box_log2_scale_present_flag is 0, it indicates that the value of the original scale is defined to be the same as the value of gsh_box_log2_scale. If the value of gps_gsh_box_log2_scale_present_flag is 1, it indicates that the value of the original scale is defined to be the same as the value of gps_gsh_box_log2_scale.

[0315] If the value of gps_gsh_box_log2_scale_present_flag is 1, the syntax 2410 for the geometry slice header includes gsh_box_log2_scale.

[0316] gsh_box_log2_scale indicates the scaling factor of the bounding box origin for each slice identified by gsh_slice_id.

[0317] gsh_box_origin_x indicates the x value of the origin of the bounding box scaled by the value indicated by gsh_box_log2_scale. gsh_box_origin_x is the same as the center x of the slice (slice_origin_x) and is smaller than the original scale.

[0318] gsh_box_origin_y indicates the y value of the origin of the bounding box scaled by the value indicated by gsh_box_log2_scale. gsh_box_origin_y is the same as the center y of the slice (slice_origin_y) and is smaller than the original scale.

[0319] Gsh_box_origin_z indicates the z value of the origin of the bounding box scaled by the value indicated by gsh_box_log2_scale. Gsh_box_origin_z is the same as the slice center z (slice_origin_z) and is smaller than the original scale.

[0320] If the value of gps_box_present_flag is 0, the center x, y and z values ​​of the slice are inferred to be 0.

[0321] The following shows the relevant information when the value of independent_decodable_flag is 0:

[0322] ref_geom_slice_id indicates the ID of the geometry slice bitstream to be referenced when decoding the slice. The value of ref_geom_slice_id is the same as the value of gsh_slice_id of the corresponding geometry slice bitstream.

[0323] split_info_present_in_slice_header_flag indicates whether additional information is sent in the geometry slice header. If the value of split_info_present_in_slice_header_flag is 1, split_info_present_in_slice_header_flag indicates that additional information is sent in the geometry slice header. If the value of split_info_present_in_slice_header_flag is 0, split_info_present_in_slice_header_flag indicates that additional information is not sent in the geometry slice header.

[0324] If the value of split_info_present_in_slice_header_flag is 1, the syntax 2410 for the geometry slice header includes layer_info.

[0325] The layer_info indicates the layer of the geometry data included in the slice (for example, the layer described in Figures 19 to 22). If the value of the split_type information described in Figure 23 is 0, the layer_info indicates the LOD level. If the value of the split_type information described in Figure 23 is 1, the layer_info indicates the octree level (or octree depth).

[0326] The syntax 2410 for a geometry slice header according to an embodiment further includes the following information:

[0327] gsh_log2_cur_nodesize indicates the node size of the octree structure used to represent the point cloud data included in the slice. When slices are divided by layering based on the octree structure, gsh_log2_cur_nodesize indicates the node size of the octree depth corresponding to each layer. In this embodiment, the node represents a 3D cube, so the node size is expressed as NxNxN, where N is an integer. For example, if the octree depth layer is Max, it is expressed as 1 (i.e., the octree node size is 1x1x1), if it is Max-1, it is expressed as 2, and if it is Max-2, it is expressed as 2x2. A value of 1 in gsh_log2_cur_nodesize indicates the same resolution as the original point cloud data. A value of 0 in full_res_flag, described in Figure 23, indicates that sampled point cloud data is represented.

[0328] Also, if a slice contains geometry data corresponding to one or more octree depths, gsh_log2_cur_nodesize indicates the node size for the depth closest to the leaf node.

[0329] gsh_log2_max_nodesize is the MaxNodeSize variable used in the decryption process.

[0330] The value of MaxNodeSize is expressed as follows:

[0331] MAXNODESIZE=2 ( gbh_log2_max_nodesize )

[0332] The depth of the octree is calculated as follows:

[0333] Octree depth layer=gsh_log2_max_nodesize / gsh_log2_cur_nodesize

[0334] Therefore, a receiving device according to an embodiment can check the depth of the octree structure of the point cloud data (or geometry) included in the current geometry slice based on gsh_log2_cur_nodesize and gsh_log2_max_nodesize.

[0335] gsh_points_number indicates the number of coded points in this slice.

[0336] The syntax 2410 for the geometry slice header according to the embodiment shown in FIG. 24 is not limited to the above examples, and further includes additional information (or fields, parameters, etc.) not shown.

[0337] FIG. 25 is an example of syntax for a feature slice bitstream according to an embodiment.

[0338] FIG. 25 is an example of syntax for a feature bitstream (or feature slice bitstream) corresponding to one slice when point cloud data is divided into one or more slices.

[0339] 25 shows an example of a syntax for an attribute slice bitstream according to an embodiment. The attribute slice bitstream includes an attribute slice header (attribute_slice_header) and attribute slice data (attribute_slice_data).

[0340] 25 is an example of a syntax for an attribute slice header according to an embodiment. The syntax for an attribute slice header includes the following information (or fields, parameters, etc.):

[0341] The ash_attr_parameter_set_id has the same value as the aps_attr_parameter_set_id of the active APS (eg, the aps_attr_parameter_set_id included in the syntax for the APS).

[0342] ash_attr_SPS_attr_idx identifies the attribute set contained in the active SPS.

[0343] The value of ash_attr_SPS_attr_idx ranges from 0 to the SPS_num_attribute_sets value contained within the active SPS.

[0344] ash_attr_geom_slice_id indicates the value of gsh_slice_id of the active geometry slice header (e.g., gsh_slice_id in syntax 2410 for the geometry slice header described in FIG. 24). As described in FIGS. 1 to 24, attribute decoding is performed based on geometry decoding. Therefore, ash_attr_geom_slice_id indicates the geometry slice referenced by the corresponding attribute slice.

[0345] aps_slice_qp_delta_present_flag indicates whether the attribute slice header includes a quantization parameter.

[0346] If the value of aps_slice_qp_delta_present_flag is 1, the syntax 2510 for the attribute slice header includes ash_qp_delta_luma and ash_qp_delta_chroma.

[0347] ash_qp_delta_luma indicates the luma delta ap of the initial slice qp in the active feature parameter set. If this information is not signaled, the value of ash_qp_delta_luma is inferred to be 0.

[0348] ash_qp_delta_chroma indicates the chroma delta ap of the initial slice qp in the active characteristic parameter set. If this information is not signaled, the value of ash_qp_delta_chroma is inferred to be 0.

[0349] The independent_decodable_flag indicates whether this slice can be independently decoded. If the value of the independent_decodable_flag is 1, the independent_decodable_flag indicates that this slice can be independently decoded. If the value of the independent_decodable_flag is 0, the independent_decodable_flag indicates that this slice cannot be independently decoded. Therefore, a receiving device (e.g., the receiving device 10004 in FIG. 1, the point cloud decoders in FIGS. 10 and 11, and the receiving device in FIG. 13) decodes this slice based on other slices.

[0350] The following shows the relevant information when the value of independent_decodable_flag is 0:

[0351] ref_attr_slice_id indicates the ID of another slice (or attribute slice bitstream) to be referenced when decoding a slice. The value of ref_attr_slice_id is the same as the value of ash_slice_id of the corresponding attribute slice bitstream. As described above, attribute decoding is based on geometry decoding. Therefore, a receiving device according to an embodiment can confirm the geometry slice required when decoding the slice identified by ref_attr_slice_id using ref_geom_slice_id (e.g., ref_geom_slice_id in the syntax 2410 for the geometry slice header described in FIG. 24). In addition, the syntax 2510 for the attribute slice header further includes information (e.g., ash_attr_geom_slice_id) for identifying the geometry slice required to decode the slice indicated by ref_attr_slice_id.

[0352] split_info_present_in_slice_header_flag indicates whether additional information is sent in the attribute slice header. If the value of split_info_present_in_slice_header_flag is 1, split_info_present_in_slice_header_flag indicates that additional information is sent in the attribute slice header. If the value of split_info_present_in_slice_header_flag is 0, split_info_present_in_slice_header_flag indicates that additional information is not sent in the attribute slice header.

[0353] If the value of split_info_present_in_slice_header_flag is 1, the syntax 2510 for the attribute slice header includes layer_info.

[0354] The layer_info indicates the layer of the attribute data included in the slice (for example, the layer described in Figures 19 to 22). If the value of the split_type information described in Figure 23 is 0, the layer_info indicates the LOD level. If the value of the split_type information described in Figure 23 is 1, the layer_info indicates the octree level (or octree depth).

[0355] The syntax 2510 for the attribute slice header according to the embodiment shown in FIG. 25 is not limited to the above examples and may further include additional information (or fields, parameters, etc.) not shown.

[0356] FIG. 26 shows the structure of signaling information according to an embodiment.

[0357] As illustrated in Figures 23 to 25, a bitstream according to an embodiment includes an SPS, one or more geometry slices, and one or more feature slices. A point cloud data processing device according to an embodiment (e.g., the transmitting device 10000 illustrated in Figure 1 or the transmitting device illustrated in Figure 12) transmits the entire layered geometry (e.g., geometry corresponding to LOD 0 to LOD N (the highest level N, e.g., 2)) and layered features (e.g., features corresponding to LOD 0 to LOD N (the highest level N, e.g., 2)) as shown in example 1910 of Figure 19. Figure 26 shows an example of the structure of signaling information when the entire geometry and the entire feature having a three-level LOD-based layer structure are transmitted.

[0358] Example 2600 in FIG. 26 shows the syntax for an SPS. An SPS according to this embodiment includes the information (or parameters) included in the syntax for an SPS described in FIG. 23. Therefore, a description of the information included in the syntax for an SPS will be omitted. As shown in FIG. 26, an SPS includes an SPS-id. Since the geometry and attributes according to this embodiment have a three-level LOD-based layer structure, the split_slice flag is 1, the split_type is 0, and the value of num_LOD is 3. Furthermore, since the geometry and attributes according to this embodiment correspond to the full LOD, the value of full_res_flag is 1.

[0359] As described in FIGS. 19 to 25, one slice includes one geometry sub-bitstream or a geometry sub-bitstream corresponding to one layer.

[0360] As described above, the geometry has a three-level LOD-based layer structure, so that the geometry according to the embodiment includes a first geometry slice corresponding to the lowest first level (or first layer), a second geometry slice corresponding to the second level (or second layer), and a third geometry slice corresponding to the highest third level (or third layer).

[0361] Each geometry slice includes a geometry slice header and geometry slice data (or geometry sub-bitstreams as described in Figures 19 to 22). The first geometry slice includes a first geometry sub-bitstream (e.g., first geometry sub-bitstream 2111), the second geometry slice includes a second geometry sub-bitstream (e.g., second geometry sub-bitstream 2112), and the third geometry slice includes a third geometry sub-bitstream (e.g., third geometry sub-bitstream 2113).

[0362] 26 shows syntax for a first geometry slice header 2611, syntax for a second geometry slice header 2612, and syntax for a third geometry slice header 2613. The syntax for the first geometry slice header 2611, syntax for the second geometry slice header 2612, and syntax for the third geometry slice header 2613 are included in each geometry slice, and in this example 2610, for convenience of explanation, the syntax for the geometry slice headers is shown together.

[0363] The syntax for each geometry slice header according to the embodiment includes the information (or parameters) included in the syntax 2410 for the geometry slice header described in Fig. 24. Therefore, a description of the information included in the syntax for the geometry slice header according to the embodiment will be omitted.

[0364] The first geometry slice is the first geometry slice. Therefore, the value of gsh_slice_id included in the syntax 2611 for the first geometry slice header according to the embodiment is 0. As described above, the first geometry slice corresponds to the lowest level (or layer). Therefore, a receiving device according to the embodiment (e.g., the receiving device 10004 of FIG. 1, the point cloud decoders of FIGS. 10 and 11, and the receiving device of FIG. 13) can decode the data included in the first geometry slice and provide point cloud content without other geometry slices. Therefore, the value of independent_decodable_flag is 1. The lower the LOD level, the larger the node size of the octree depth. The syntax 2611 for the first geometry slice header includes gsh_log2_cur_nodesize having a value of 8 and gsh_log2_max_nodesize having a value of 128. The second geometry slice is the second geometry slice. Therefore, the value of gsh_slice_id included in the syntax 2612 for the second geometry slice header according to the embodiment is 1. As described above, the second geometry slice corresponds to the second level (or layer). Therefore, in order for a receiving device according to an embodiment to decode data included in the second geometry slice, a geometry slice of a lower level than the second geometry slice is required. Therefore, the value of independent_decodable_flag is 0. The syntax 2612 for the second geometry slice header according to an embodiment further includes ref_geom_slice_id and layer_info. In an embodiment, the value of ref_geom_slice_id is the same as the value of gsh_slice_id of the first geometry slice, which is 0. In an embodiment, layer_info indicates the layer of the second geometry slice. Therefore, the value of layer_info is 1.

[0365] The higher the LOD level, the smaller the node size of the octree depth. The syntax 2612 for the second geometry slice header includes gsh_log2_cur_nodesize with a value of 4 and gsh_log2_max_nodesize with a value of 128.

[0366] The third geometry slice is the third geometry slice. Therefore, the value of gsh_slice_id included in the syntax 2613 for the third geometry slice header according to the embodiment is 2. As described above, the third geometry slice corresponds to the third level (or layer). Therefore, in order for the receiving device according to the embodiment to decode the data included in the third geometry slice, a geometry slice of a lower level than the third geometry slice is required. Therefore, the value of independent_decodable_flag is . The syntax 2613 for the third geometry slice header according to the embodiment further includes ref_geom_slice_id and layer_info. The value of ref_geom_slice_id according to the embodiment is the same as the value of gsh_slice_id of the second geometry slice, which is 1. The layer_info according to the embodiment indicates the layer of the third geometry slice. Therefore, the value of layer_info is 2.

[0367] The higher the LOD level, the smaller the node size of the octree depth. The syntax 2613 for the third geometry slice header includes gsh_log2_cur_nodesize with a value of 1 and gsh_log2_max_nodesize with a value of 128.

[0368] As described in Figures 19 to 25, one slice includes one feature sub-bitstream or a feature sub-bitstream corresponding to one layer.

[0369] As described above, the feature has a three-level LOD-based layer structure, so that the feature according to the embodiment includes a first feature slice corresponding to the lowest first level (or first layer), a second feature slice corresponding to the second level (or second layer), and a third feature slice corresponding to the highest third level (or third layer).

[0370] Each attribute slice includes an attribute slice header and attribute slice data (or attribute sub-bitstreams as described in Figures 19 to 22). The first geometry slice includes a first attribute sub-bitstream (e.g., first attribute sub-bitstream 2121), the second attribute slice includes a second attribute sub-bitstream (e.g., second attribute sub-bitstream 2122), and the third attribute slice includes a third attribute sub-bitstream (e.g., third geometry sub-bitstream 2123).

[0371] 26 shows syntax for a first attribute slice header 2621, syntax for a second attribute slice header 2622, and syntax for a third attribute slice header 2623. The syntax for the first attribute slice header 2621, syntax for the second attribute slice header 2622, and syntax for the third attribute slice header 2623 are included in each attribute slice, and in this example 2620, for convenience of explanation, the syntax for the attribute slice headers is shown together.

[0372] The syntax for each attribute slice header according to the embodiment includes the information (or parameters) included in the syntax 2510 for the geometry slice header described in Figure 25. Therefore, a description of the information included in the syntax for the attribute slice header according to the embodiment will be omitted.

[0373] The first attribute slice is the first attribute slice. Therefore, the value of ash_slice_id included in the syntax 2621 for the first attribute slice header according to this embodiment is 0. As described above, the first attribute slice corresponds to the lowest level (or layer). Therefore, a receiving device according to this embodiment can decode the data included in the first attribute slice and provide point cloud content without other attribute slices. Therefore, the value of independent_decodable_flag is 1.

[0374] The second attribute slice is the second attribute slice. Therefore, the value of ash_slice_id included in the syntax 2622 for the second attribute slice header according to the embodiment is 1. As described above, the second attribute slice corresponds to the second level (or layer). Therefore, in order for the receiving device according to the embodiment to decode the data included in the second attribute slice, an attribute slice of a lower level than the second attribute slice is required. Therefore, the value of independent_decodable_flag is 0. The syntax 2622 for the second attribute slice header according to the embodiment further includes ref_attr_slice_id and layer_info. The value of ref_attr_slice_id according to the embodiment is the same as the value of ash_slice_id of the first attribute slice, which is 0. The layer_info according to the embodiment indicates the layer of the second attribute slice. Therefore, the value of layer_info is 1.

[0375] The third attribute slice is the third attribute slice. Therefore, the value of ash_slice_id included in the syntax 2623 for the third attribute slice header according to the embodiment is 2. As described above, the third attribute slice corresponds to the third level (or layer). Therefore, in order for a receiving device according to the embodiment to decode the data included in the third attribute slice, an attribute slice of a level lower than the third attribute slice is required. Therefore, the value of independent_decodable_flag is 0. The syntax 2623 for the third attribute slice header according to the embodiment further includes ref_attr_slice_id and layer_info. The value of ref_attr_slice_id according to the embodiment is the same as the value of ash_slice_id of the second attribute slice, which is 1. The layer_info according to the embodiment indicates the layer of the third attribute slice. Therefore, the value of layer_info is 2.

[0376] The syntax for an attribute slice header according to an embodiment further includes information about the associated geometry slice.

[0377] Therefore, a receiving device can receive a bitstream, obtain signaling information from the SPS, geometry slice header, and attribute slice header described in Figures 24 to 26, and decode the geometry and attributes corresponding to a particular layer (or level) to perform scalable representation.

[0378] FIG. 27 shows the structure of signaling information according to an embodiment.

[0379] FIG. 27 shows an example of the structure of the signaling information described in FIG.

[0380] Example 2700 shown in the upper part of Figure 27 is an example of the structure of signaling information when partial geometries and partial attributes corresponding to one or more layers are transmitted for geometries and attributes having a three-level LOD-based layer structure. Example 2700 shown in the upper part of Figure 27 shows the structure of signaling information when transmitting first and second geometry slices (e.g., the first and second geometry slices described in Figure 26) and first and second attribute slices (e.g., the first and second attribute slices described in Figure 26) corresponding to the first and second layers.

[0381] The example 2700 of FIG. 27 includes a syntax 2710 for an SPS. The information included in the syntax 2710 for an SPS is the same as that described in FIG. 26, except for the value of full_res_flag. Because the geometry and attributes according to this embodiment correspond to only a partial LOD, the value of full_res_flag is 0 and the value of full_geo_present_flag is 0. The syntax 2720 for a first geometry slice header and the syntax 2721 for a second geometry slice header included in the example 2700 of FIG. 27 are the same as the syntax 2611 for a first geometry slice header and the syntax 2612 for a second geometry slice header described in FIG. 26, and therefore detailed description thereof will be omitted. The syntax 2730 for the first attribute slice header and the syntax 2731 for the second attribute slice header included in the example 2700 of Figure 27 are the same as the syntax 2621 for the first attribute slice header and the syntax 2622 for the second attribute slice header described in Figure 26, so detailed explanations will be omitted.

[0382] Example 2740 shown at the bottom of Figure 27 is an example of the structure of signaling information when partial geometries and partial attributes corresponding to one or more layers are transmitted for geometry and attributes having a two-level LOD-based layer structure. Example 2740 shown at the bottom of Figure 27 shows the structure of signaling information when transmitting a first geometry slice (e.g., the first geometry slice described in Figure 26) and a first attribute slice (e.g., the first attribute slice described in Figure 26) corresponding to a first layer.

[0383] Example 2740 of FIG. 27 includes syntax 2750 for an SPS. Syntax 2750 for an SPS includes the same information as syntax 2710 for an SPS described above, but the LOD level is different. Therefore, the value of num_LOD is 2. Syntax 2760 for a first geometry slice header included in example 2740 of FIG. 27 is the same as syntax 2611 for a first geometry slice header described in FIG. 26, so a detailed description thereof will be omitted. Syntax 2770 for a first attribute slice header included in example 2740 of FIG. 27 is the same as syntax 2621 for a first attribute slice header described in FIG. 26, so a detailed description thereof will be omitted.

[0384] Therefore, a receiving device according to an embodiment (e.g., receiving device 10004 of Figure 1, point cloud decoder of Figures 10 and 11, receiving device of Figure 13) can receive a bitstream, obtain signaling information from the SPS, geometry slice header and attribute slice header described in Figures 24 to 27, and decode the geometry and attributes corresponding to a particular layer (or level) to perform scalable representation.

[0385] FIG. 28 shows the structure of signaling information according to an embodiment.

[0386] Figure 28 is an example of the structure of signaling information described in Figures 26 and 27. Example 2800 shown in Figure 28 is an example of the structure of signaling information when first to third geometry slices corresponding to first to third layers (e.g., first to third geometry slices described in Figure 26) and first and second attribute slices corresponding to first and second layers (e.g., first and second attribute slices described in Figure 26) are transmitted for geometries and attributes having a three-level LOD-based layer structure.

[0387] The example 2800 in Figure 28 includes syntax 2810 for SPS. The information included in the syntax 2810 for SPS is the same as that described in Figure 26, but since the geometry corresponds to the entire LOD and the characteristics correspond to only a partial LOD, the value of full_res_flag is 0 and the value of full_geo_present_flag is 1.

[0388] An example 2820 of the syntax for the first geometry slice header through the syntax for the third geometry slice header included in example 2800 of Fig. 28 is the same as the syntax 2611 for the first geometry slice header through the syntax 2623 for the third geometry slice header described in Fig. 26, so a detailed description thereof will be omitted. An example 2830 of the syntax for the first attribute slice header and the syntax for the second attribute slice header included in example 2800 of Fig. 28 is the same as the syntax 2621 for the first attribute slice header and the syntax 2622 for the second attribute slice header described in Fig. 26, so a detailed description thereof will be omitted.

[0389] Therefore, a receiving device according to an embodiment (e.g., receiving device 10004 of Figure 1, point cloud decoder of Figures 10 and 11, receiving device of Figure 13) can receive a bitstream, obtain signaling information from the SPS, geometry slice header and attribute slice header described in Figures 24 to 27, and decode the geometry and attributes corresponding to a particular layer (or level) to perform scalable representation.

[0390] FIG. 29 illustrates an example of a point cloud data processing device according to an embodiment.

[0391] The point cloud data processing device 2900 according to the embodiment shown in Figure 29 is an example of the receiving device 10004 of Figure 1, the point cloud decoder of Figures 10 and 11, and the receiving device of Figure 13. Therefore, the point cloud data processing device 2900 performs operations corresponding to the reverse process of the operations of the point cloud data processing device 2000 described in Figure 20. The point cloud data processing device 2900 performs operations that are the same as or similar to the operations of the receiving device described in Figures 1 to 28. Although not shown in Figure 29, the point cloud data processing device 2900 further includes one or more elements for performing the decoding operations described in Figures 1 to 28.

[0392] The point cloud data processing device 2900 includes a receiver 2910, a de-mux 2920, a metadata parser 2930, a sub-bitstream classifier 2940, a geometry decoder 2950, ​​an attribute decoder 2960, and a renderer 2970.

[0393] The receiver 2910 according to the embodiment receives data output by a transmitting device (for example, the transmitting device 10000 described in FIG. 1 or the transmitting device described in FIG. 12, or the point cloud data processing device 2000 described in FIG. 20). The data received by the receiver 2910 corresponds to the bit stream described in FIG. 1 and the data output by the transmitter 2060 described in FIG. 20.

[0394] The demultiplexer 2920 according to the embodiment demultiplexes the received data, i.e., the received data is demultiplexed into a geometry bitstream or geometry sub-bitstream, a characteristic bitstream or characteristic sub-bitstream, and metadata (or signaling information, parameters).

[0395] The metadata parser 2930 according to the embodiment acquires the metadata output from the demultiplexer 2920. The metadata according to the embodiment includes the SPS described in FIGS. 23 to 28, etc. Therefore, the point cloud data processing device 2900 according to the embodiment can obtain information related to geometry and feature layering based on the metadata. The metadata parser 2930 also transmits information required for geometry decoding and / or feature decoding to the respective decoders. The sub-bitstream classifier 2930 according to the embodiment classifies the geometry sub-bitstreams and feature sub-bitstreams required for decoding based on information included in the headers (e.g., the geometry slice headers or feature slice headers described in FIGS. 24 and 25) of the geometry bitstream or geometry sub-bitstream (e.g., the geometry sub-bitstreams described in FIGS. 21 and 22) and the feature bitstream or feature sub-bitstream (e.g., the feature sub-bitstreams described in FIGS. 21 and 22) and information output from the metadata parser 2930. The sub-bitstream classifier 2940 also selects geometry and feature layers. The geometry decoder 2950 according to the embodiment receives one or more geometry sub-bitstreams corresponding to one or more layers and decodes the geometry. The operation of the geometry decoder 2950 according to the embodiment is the same as or similar to the operation of the arithmetic decoder 11000, the octree synthesis unit 11001, the surface approximation synthesis unit 11002, the geometry reconstruction unit 11003, and the coordinate system inverse transformation unit 11004 described in FIG. 11 , and therefore a detailed description thereof will be omitted.

[0396] The feature decoder 2960 according to the embodiment receives one or more feature sub-bitstreams corresponding to one or more layers based on the decoding by the geometry decoder 2950 and performs feature decoding. The operation of the feature decoder 2960 according to the embodiment is the same as or similar to the operation of the arithmetic decoder 11005, the inverse quantization unit 11006, the RAHT transform unit 11007, the LOD generation unit 11008, the inverse lift unit 11009 and / or the inverse hue transform unit 11010 described in Fig. 11, so a detailed description will be omitted. The renderer 2970 according to the embodiment can receive geometry and features and convert them into a format for final output.

[0397] FIG. 30 illustrates the operation of point cloud decoding according to an embodiment.

[0398] FIG. 30 illustrates the operation of a point cloud data processing device (eg, point cloud data processing device 2900 or sub-bitstream classifier 2940 described in FIG. 29) that receives a bitstream and classifies or selects it into one or more substreams.

[0399] The example 3000 shown at the top of Figure 30 shows a process in which a point cloud data processing device according to an embodiment processes geometry sub-bitstreams (e.g., a first geometry sub-bitstream 2111 including a first geometry corresponding to LOD0, LOD1, and LOD2 described in Figure 21, a second geometry sub-bitstream 2112 including a second geometry R1 corresponding to LOD1 and LOD2, and a third geometry sub-bitstream 2113 including a third geometry R2 corresponding only to LOD2) having an LOD-based layer structure and feature sub-bitstreams (e.g., a first feature sub-bitstream 2121 including a first feature corresponding to LOD0, LOD1, and LOD2 described in Figure 21, a second feature sub-bitstream 2122 including a second feature R1 corresponding to LOD1 and LOD2, and a third feature sub-bitstream 2123 including a third feature R2 corresponding only to LOD2) using the same layer. A point cloud data processing device according to the embodiment selects and decodes geometry sub-bitstreams and feature sub-bitstreams according to layers corresponding to LOD levels up to 1. As shown on the right side of example 3000, in order to process point cloud data corresponding to LOD level 1, point cloud data corresponding to LOD level 0 must also be processed. Therefore, the point cloud processing device selects 3001 geometry sub-bitstreams and feature sub-bitstreams corresponding to LOD levels 0 and 1 from the received bitstream and performs geometry decoding and feature decoding, respectively. R1 shown in example 3000 indicates geometry and features included only in LOD level 1. The point cloud processing device does not select 3002 geometry sub-bitstreams and feature bitstreams corresponding to LOD levels higher than LOD level 1 (e.g., LOD level 2) from the received bitstream. R2 shown in example 3000 indicates geometry and features included only in LOD level 2.

[0400] Example 3000 shown at the bottom of Figure 30 illustrates a process in which a point cloud data processing device according to an embodiment processes geometry sub-bitstreams and feature sub-bitstreams having an LOD-based layer structure according to different layers. The point cloud data processing device according to an embodiment selects and decodes geometry sub-bitstreams according to layers corresponding to up to LOD level 2, and feature sub-bitstreams according to layers corresponding to up to LOD level 1. As shown on the right side of example 3010, to decode geometry corresponding to LOD level 2, geometries corresponding to LOD levels 0 and 1 are required. Therefore, the point cloud processing device selects 3011 geometry sub-bitstreams corresponding to LOD levels 0, 1, and 2 and feature sub-bitstreams of LOD levels 0 and 1 from the received bitstream, and performs geometry decoding and feature decoding, respectively. The point cloud processing device does not select 3012 feature bitstreams corresponding to LOD levels greater than LOD level 1 (e.g., LOD level 2) from the received bitstream.

[0401] A transmitting device according to an embodiment (for example, the transmitting device 10000 described in FIG. 1 or the transmitting device described in FIG. 12, or the point cloud data processing device 2000 described in FIG. 20) can generate and transmit a bitstream in which some geometry and some features or some features have been removed, as shown in examples 3000 and 3010 in FIG. 30, depending on the performance of the receiving device.

[0402] FIG. 31 illustrates a point cloud data transmission and reception configuration according to an embodiment.

[0403] FIG. 31 shows a point cloud data transmission and reception configuration for scalable decoding and representation.

[0404] The example 3100 shown in the upper part of FIG. 31 is an example of a configuration for transmitting and receiving a partial PCC bitstream. A transmitting device according to an embodiment (e.g., the transmitting device 10000 described in FIG. 1, the transmitting device described in FIG. 12, or the point cloud data processing device 2000 described in FIG. 20) performs scalable encoding on source geometry and source attribute, and generates a partial PCC bitstream by selecting a bitstream (or a sub-bitstream described in FIG. 21 and FIG. 22) according to the layer described in FIG. 18 to FIG. 29. Thus, the transmitting device transmits only necessary data, enabling efficient information transmission in terms of bandwidth. A receiving device according to an embodiment (e.g., the receiving device 10004 in FIG. 1, the point cloud decoder in FIG. 10 and FIG. 11, the receiving device in FIG. 13, or the point cloud data processing device 2900 described in FIG. 29) receives and decodes the partial PCC bitstream and outputs partial geometry and partial attribute 3102. The configuration shown in example 3100 does not require further data processing (e.g., decoding and transcoding) in the transmitting device, thereby reducing the probability of delays occurring during the data processing. In addition, the transmitting device also transmits signaling information related to the layer (e.g., the signaling information described in Figures 23 to 28), so that the receiving device can efficiently decode the partial geometry and partial characteristics based on the signaling information.

[0405] An example 3110 shown at the bottom of Figure 31 is an example of a configuration for transmitting and receiving a complete PCC bitstream. The transmitting device according to the embodiment performs scalable encoding and stores the encoded geometry and attributes. According to the embodiment, the transmitting device layers the geometry and attributes in a slice unit and generates a bitstream (e.g., an entire PCC bitstream) including the geometry and attributes in a slice unit 3111. Since the transmitting device also transmits information related to the layer of the geometry and attributes (e.g., the signaling information described in Figures 23 to 28), the receiving device can obtain information about the layer and slice. Therefore, the receiving device receives the bitstream and selects a geometry sub-bitstream and / or an attribute sub-bitstream in a slice unit before decoding to decode the entire geometry and attributes, or decodes the geometry and attributes corresponding to some layers and output partial geometry and partial attributes 3112. According to the configuration shown in this example 3110, the receiving device selects sub-bitstreams according to the density of the point cloud data to be represented according to the performance or the field. In addition, the receiving device selects layers before decoding to perform more efficient decoding and can support decoders with various performance capabilities.

[0406] FIG. 32 is an example flowchart of a method for processing point cloud data according to an embodiment.

[0407] Flowchart 3200 in Fig. 32 shows a point cloud data processing method of a point cloud data processing device (for example, the transmitting device described in Figs. 1, 11, 14 and 15 and the point cloud data processing device 2000 described in Fig. 20). The point cloud data processing device according to the embodiment performs operations that are the same as or similar to the encoding operations described in Figs. 1 to 31.

[0408] The point cloud data processing device according to the embodiment encodes 3210 the point cloud data including geometry information and attribute information. The geometry information according to the embodiment is information indicating the positions of the points of the point cloud data. The attribute information according to the embodiment is information indicating the attributes of the points of the point cloud data.

[0409] A point cloud data processing device according to an embodiment encodes geometry information and outputs a geometry bitstream. Also, the point cloud data processing device encodes feature information and outputs a feature bitstream. The point cloud data processing device according to an embodiment performs an operation that is the same as or similar to the geometry information encoding operation described with reference to FIGS. 1 to 31. Also, the point cloud data processing device performs an operation that is the same as or similar to the feature information encoding operation described with reference to FIGS. 1 to 31. A point cloud data processing device according to an embodiment (e.g., the sub-bitstream generator 2030 described with reference to FIG. 20) generates one or more geometry sub-bitstreams (e.g., the first geometry sub-bitstream 2111, the second geometry sub-bitstream 2112, and the third geometry sub-bitstream 2113 described with reference to FIG. 21) corresponding to one or more layers from the geometry bitstream. A point cloud data processing device according to an embodiment (e.g., the sub-bitstream generator 2030 described in FIG. 20) generates one or more feature sub-bitstreams (e.g., the first feature sub-bitstream 2121, the second feature sub-bitstream 2122, and the third feature sub-bitstream 2123 described in FIG. 21) corresponding to one or more layers of a feature bitstream. As described in FIGS. 20 to 28, the geometry sub-bitstream is included in the geometry slice, and the feature sub-bitstream is included in the feature slice. A bitstream according to an embodiment includes one or more geometry slices and one or more feature slices. A detailed description of the bitstreams is as described in FIG. 23.

[0410] A bitstream according to the embodiment transmits signaling information (e.g., the signaling information described in FIGS. 23 to 28). The bitstream includes slice-related information (e.g., the slice-related information described in FIG. 23), and the slice-related information includes first information (e.g., split_slice_flag described in FIG. 23) indicating whether the geometry bitstream and the attribute bitstream are divided into one or more slices. If the first information indicates that the geometry bitstream and the attribute bitstream are divided into one or more slices, the slice-related information includes second information (e.g., split_type described in FIG. 23) indicating whether the layers of the geometry sub-bitstream and the attribute sub-bitstream are based on a Level of Detail (LOD) or a depth of an octree structure. The signaling information according to the embodiment is the same as that described in FIGS. 23 to 28, and therefore a detailed description thereof will be omitted.

[0411] FIG. 33 is an example of a flowchart of a method for processing point cloud data according to an embodiment.

[0412] Flowchart 3300 in Fig. 33 shows a point cloud data processing method of a point cloud data processing device (for example, the point cloud data receiving device or point cloud data decoder described in Figs. 1, 13, 14, 16 to 25, or the point cloud data processing device 2900 described in Fig. 29). The point cloud data processing device according to the embodiment performs operations that are the same as or similar to the decoding operations described in Figs. 1 to 31.

[0413] A point cloud data processing device according to an embodiment receives a bitstream including point cloud data 3310. Geometry information according to an embodiment is information indicating the positions of points in the point cloud data. Attribute information according to an embodiment is information indicating attributes of points in the point cloud data. The structure of the bitstream according to an embodiment is the same as that described with reference to FIGS. 23 to 28, and therefore detailed description thereof will be omitted. As described with reference to FIGS. 20 to 28, geometry sub-bitstreams are included in geometry slices, and attribute sub-bitstreams are included in attribute slices. A bitstream according to an embodiment includes a geometry slice including a geometry sub-bitstream and an attribute sub-bitstream. One or more geometry sub-bitstreams (e.g., the first geometry sub-bitstream 2111, the second geometry sub-bitstream 2112, and the third geometry sub-bitstream 2113 described with reference to FIG. 21) and one or more attribute sub-bitstreams (e.g., the first attribute sub-bitstream 2121, the second attribute sub-bitstream 2122, and the third attribute sub-bitstream 2123 described with reference to FIG. 21) correspond to one or more layers. A bitstream according to the embodiment transmits signaling information (e.g., the signaling information described in FIGS. 23 to 28). The bitstream includes slice-related information (e.g., the slice-related information described in FIG. 23), and the slice-related information includes first information (e.g., split_slice_flag described in FIG. 23) indicating whether the geometry bitstream and the attribute bitstream are split into one or more slices. If the first information indicates that the geometry bitstream and the attribute bitstream are split into one or more slices, the slice-related information includes second information (e.g., split_type described in FIG. 23) indicating whether the layers of the geometry sub-bitstream and the attribute sub-bitstream are based on a Level of Detail (LOD) or a depth of an octree structure. The signaling information according to the embodiment is the same as that described in FIGS. 23 to 28, and therefore a detailed description thereof will be omitted.

[0414] The point cloud data processing device according to the embodiment decodes the point cloud data 3320. The point cloud data processing device according to the embodiment decodes a geometry sub-bitstream corresponding to an arbitrary layer based on information related to the slice, and decodes a characteristic sub-bitstream corresponding to an arbitrary layer, to perform scalable representation.

[0415] The components of the point cloud data processing apparatus according to the embodiments described in FIGS. 1 to 33 may be implemented as hardware, software, firmware, or a combination thereof, including one or more processors coupled to a memory. The components of the device according to the embodiments may be implemented as a single chip, for example, a single hardware circuit. The components of the point cloud data processing apparatus according to the embodiments may be implemented as separate chips. Any of the components of the point cloud data processing apparatus according to the embodiments may be configured with one or more processors capable of executing one or more programs, and the one or more programs may include instructions for causing or executing any one or more operations / methods of the point cloud data processing apparatus described in FIGS. 1 to 33.

[0416] For convenience of explanation, the figures have been described separately. However, it is possible to combine the embodiments described in the figures to realize a new embodiment. Furthermore, designing a computer-readable recording medium on which a program for executing the previously described embodiments is recorded, as needed by those skilled in the art, is also within the scope of the embodiments. As described above, the apparatus and method according to the embodiments are not limited to the configurations and methods of the described embodiments. The embodiments may be configured by selectively combining all or part of each embodiment, allowing for various modifications. While preferred embodiments have been shown and described, the embodiments are not limited to the specific embodiments described above. Various modifications may be made by those skilled in the art to which the invention pertains without departing from the spirit and scope of the embodiments claimed in the claims. Such modifications should not be interpreted as being separate from the technical ideas and perspectives of the embodiments.

[0417] The descriptions of the apparatus and method according to the embodiments may be applied to complement each other. For example, the point cloud data transmitting method according to the embodiments is performed by the point cloud data transmitting apparatus according to the embodiments or components included in the point cloud data transmitting apparatus. Also, the point cloud data receiving method according to the embodiments is performed by the point cloud data receiving apparatus according to the embodiments or components included in the point cloud data receiving apparatus.

[0418] Various components of the apparatus according to the embodiments may be implemented by hardware, software, firmware, or a combination thereof. Various components of the embodiments may be implemented by a single chip, e.g., a single hardware circuit. In some embodiments, components according to the embodiments may be implemented by individual chips. In some embodiments, any of the components of the apparatus according to the embodiments may be implemented by one or more processors capable of executing one or more programs, and the one or more programs include instructions for causing or causing any one or more of the operations / methods according to the embodiments to be performed. Executable instructions for performing the methods / operations of the apparatus according to the embodiments may be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or may be stored in a transient CRM or other computer program product configured to be executed by one or more processors. In addition, the term "memory" in the embodiments is used as a concept that encompasses not only volatile memory (e.g., RAM, etc.), but also non-volatile memory, flash memory, PROM, etc. It may also be implemented in the form of a carrier wave, such as transmission over the Internet. In addition, a processor-readable recording medium may be distributed among network-connected computer systems, and the processor-readable code may be stored and executed in a distributed manner.

[0419] In this specification, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Furthermore, "A / B / C" means "any of A, B, and / or C." Also, "A, B, C" means "any of A, B, and / or C." Furthermore, in this document, "or" is interpreted as "and / or." For example, "A or B" means 1) only "A," 2) only "B," or 3) "A and B." In other words, in this specification, "or" means "additionally or alternatively."

[0420] Terms such as "first," "second," and the like are used to describe various components of the embodiments. However, the interpretation of the various components according to the embodiments should not be limited by these terms. These terms are merely used to distinguish one component from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of such terms does not depart from the scope of the various embodiments. Although the first user input signal and the second user input signal are both user input signals, they do not refer to the same user input signal unless the context clearly indicates otherwise.

[0421] Terms used to describe the embodiments are used to describe particular embodiments and are not intended to limit the embodiments. As used in the description of the embodiments and in the claims, the singular includes the plural unless the context clearly dictates otherwise. The term "and / or" is used to include all possible combinations between terms. "Comprises" describes the presence of features, numbers, steps, elements, and / or components, but does not imply the absence of additional features, numbers, steps, elements, and / or components. Conditional expressions such as "if" and "when" used to describe the embodiments are only intended to be selective and not limiting. It is intended that when a specific condition is met, a related action is performed in response to a specific condition, or a related definition is interpreted.

[0422] The best mode for carrying out the invention will be described in detail below. [Industrial Applicability]

[0423] It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible within the spirit and scope of the present invention. Thus, the present invention covers the modifications and variations of the present invention provided within the scope of the appended claims and their equivalents.

Claims

1. 1. A method for decoding point cloud data, comprising: receiving a bitstream including point cloud data; the point cloud data includes geometry information of layers associated with levels of an octree and characteristic information of the layers; the geometry information indicates positions of points in the point cloud data; the characteristic information indicates characteristics of the points of the point cloud data; the geometry information is conveyed based on geometry sub-slice; the feature information is conveyed based on feature sub-slice; decoding the geometry information of the geometry sub-slice; and decoding the quality information of the quality portion slice; The bitstream comprises: a slice identifier identifying the geometry sub-slice or the attribute sub-slice; layer information for indicating the layer for the geometry sub-slice or the attribute sub-slice; size information for indicating a size of a subgroup associated with the layer.

2. the geometry bitstream including the geometry information includes a first geometry sub-bitstream including a first depth and a second geometry sub-bitstream including a second depth; 2. The method of claim 1, wherein the feature bitstream containing the feature information includes a first feature sub-bitstream containing a first LOD (Level of Detail) and a second feature sub-bitstream containing residual information between a second LOD and the first LOD.

3. 1. An apparatus for decoding point cloud data, comprising:

1. A receiver configured to receive a bitstream comprising point cloud data, the bitstream comprising: the point cloud data includes geometry information of layers associated with levels of an octree and characteristic information of the layers; the geometry information indicates positions of points in the point cloud data; the characteristic information indicates characteristics of the points of the point cloud data; the geometry information is conveyed based on geometry sub-slice; a receiver, wherein the quality information is conveyed based on quality sub-slice; a decoder configured to decode the geometry information of the geometry sub-slice and to decode the feature information of the feature sub-slice; The bitstream comprises: a slice identifier identifying the geometry sub-slice or the attribute sub-slice; layer information for indicating the layer for the geometry sub-slice or the attribute sub-slice; size information for indicating a size of a subgroup associated with the layer.

4. the geometry bitstream including the geometry information includes a first geometry sub-bitstream including a first depth and a second geometry sub-bitstream including a second depth; The device of claim 3, wherein the feature bitstream containing the feature information includes a first feature sub-bitstream containing a first LOD (Level of Detail) and a second feature sub-bitstream containing residual information between a second LOD and the first LOD.

5. 1. A method for encoding point cloud data, comprising: encoding geometry information of the geometry sub-slice for a layer associated with a level of the octree; encoding quality information of a quality sub-slice for the layer; the geometry information indicates positions of points in the point cloud data; the characteristic information indicates characteristics of the points of the point cloud data; the geometry information is conveyed based on the geometry sub-slice; the feature information is conveyed based on the feature portion slice; transmitting a bitstream containing the encoded point cloud data; The bitstream comprises: a slice identifier identifying the geometry sub-slice or the attribute sub-slice; layer information for indicating the layer for the geometry sub-slice or the attribute sub-slice; size information for indicating a size of a subgroup associated with the layer.

6. the geometry bitstream including the geometry information includes a first geometry sub-bitstream including a first depth and a second geometry sub-bitstream including a second depth; The method of claim 5, wherein the feature bitstream containing the feature information includes a first feature sub-bitstream containing a first LOD (Level of Detail) and a second feature sub-bitstream containing residual information between a second LOD and the first LOD.

7. 1. An apparatus for encoding point cloud data, comprising:

1. An encoder configured to encode geometry information of geometry sub-slice for a layer associated with a level of an octree and to encode attribute information of attribute sub-slice for said layer, the geometry information indicates positions of points in the point cloud data; the characteristic information indicates characteristics of the points of the point cloud data; the geometry information is conveyed based on the geometry sub-slice; an encoder, wherein the quality information is conveyed based on the quality sub-slice; a transmitter configured to transmit a bitstream including the encoded point cloud data; The bitstream comprises: a slice identifier identifying the geometry sub-slice or the attribute sub-slice; layer information for indicating the layer for the geometry sub-slice or the attribute sub-slice; size information for indicating a size of a subgroup associated with the layer.

Citation Information

Patent Citations

  • Method and apparatus for near-lossless compression and decompression of 3D meshes and point clouds

    US20160086353A1

  • Scalable point cloud compression with transform, and corresponding decompression

    US20170347122A1

  • Point cloud compression

    WO2019055963A1

  • Information processing device and method

    WO2019078000A1

  • Image processing device and method

    WO2019198521A1