Point cloud data transmission apparatus and method, point cloud data reception apparatus and method

CN116349229BActive Publication Date: 2026-08-28LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180068605.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-10-06
Publication Date
2026-08-28
Estimated Expiration
2041-10-06

AI Technical Summary

Benefits of technology

[0009]根据实施方式的装置和方法可以高效率地处理点云数据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116349229B_ABST
    Figure CN116349229B_ABST
Patent Text Reader

Abstract

A point cloud data transmission method according to an embodiment can include encoding point cloud data and transmitting a bitstream including the point cloud data. Also, a point cloud data reception method according to an embodiment can include receiving a bitstream including point cloud data and decoding the point cloud data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments relate to methods and apparatus for processing point cloud content. Background Technology

[0002] Point cloud content is content represented by point clouds, which are collections of points belonging to a coordinate system representing three-dimensional space. Point cloud content can represent media configured in three dimensions and is used to provide various services such as virtual reality (VR), augmented reality (AR), mixed reality (MR), and autonomous driving services. However, tens of thousands to hundreds of thousands of points are needed to represent point cloud content. Therefore, methods for efficiently processing large amounts of point data are required. Summary of the Invention

[0003] Technical issues

[0004] The embodiments provide apparatus and methods for efficiently processing point cloud data. The embodiments also provide point cloud data processing methods and apparatuses to address latency and encoding / decoding complexity.

[0005] The technical scope of the implementation is not limited to the technical objectives mentioned above, and can be extended to other technical objectives that can be inferred by those skilled in the art based on all the contents disclosed herein.

[0006] Technical solution

[0007] To achieve these objectives and other advantages, and in accordance with the purposes of this disclosure, as implemented and broadly described herein, a method for transmitting point cloud data may include encoding the point cloud data and transmitting a bit stream containing the point cloud data. A method for receiving point cloud data may include receiving a bit stream containing the point cloud data and decoding the point cloud data.

[0008] Beneficial effects

[0009] The apparatus and method according to the embodiments can process point cloud data efficiently.

[0010] The apparatus and method according to the embodiments can provide high-quality point cloud services.

[0011] The apparatus and method according to the embodiments can provide point cloud content for providing general services such as VR services and autonomous driving services. Attached Figure Description

[0012] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated in and constitute a part of this application. They illustrate embodiments of the disclosure and, together with the description, serve to illustrate the principles of the disclosure. For a better understanding of the various embodiments described below, reference should be made to the description of the following embodiments in conjunction with the accompanying drawings. The same reference numerals will be used throughout the drawings to refer to the same or similar parts. In the drawings:

[0013] Figure 1 An exemplary point cloud content providing system according to an embodiment is shown;

[0014] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an implementation method;

[0015] Figure 3 An exemplary process for capturing point cloud video according to an embodiment is illustrated;

[0016] Figure 4 An exemplary point cloud encoder according to an implementation method is illustrated;

[0017] Figure 5 An example of a voxel according to an embodiment is shown;

[0018] Figure 6 An example of an octree and occupancy code according to an implementation is shown;

[0019] Figure 7 An example of a neighboring node pattern according to an implementation method is shown;

[0020] Figure 8 An example of point configuration in each LOD according to the implementation method is illustrated;

[0021] Figure 9 An example of point configuration in each LOD according to the implementation method is illustrated;

[0022] Figure 10 An example of a point cloud decoder according to an implementation method is shown;

[0023] Figure 11 An example of a point cloud decoder according to an implementation method is shown;

[0024] Figure 12 An example of a transmitting device according to an embodiment is shown;

[0025] Figure 13 An example of a receiving device according to an embodiment is shown;

[0026] Figure 14 An exemplary structure operable in conjunction with a point cloud data transmission / reception method / apparatus according to an embodiment is illustrated;

[0027] Figure 15 The process of encoding, transmitting, and decoding point cloud data according to the implementation method is illustrated;

[0028] Figure 16 An example of a layer-based point cloud data configuration according to an implementation method is shown;

[0029] Figure 17 The geometric bitstream structure and attribute bitstream structure according to the implementation method are illustrated;

[0030] Figure 18 A bitstream configuration according to an implementation method is illustrated;

[0031] Figure 19 An example of a bitstream arrangement method according to an implementation method is shown;

[0032] Figure 20 An example of a bitstream arrangement method according to an implementation method is shown;

[0033] Figure 21 A method for selecting geometric data and attribute data according to an implementation method is illustrated;

[0034] Figure 22 An example of a bitstream selection method according to an implementation method is shown;

[0035] Figure 23 An example is illustrated of a method for configuring slices containing point cloud data according to an implementation method;

[0036] Figure 24 A bitstream configuration according to an implementation method is illustrated;

[0037] Figure 25 The syntax for the sequence parameter set and geometric parameter set according to the implementation method is shown;

[0038] Figure 26 The syntax of the attribute parameter set according to the implementation method is shown;

[0039] Figure 27 The syntax of the geometric data cell header according to the implementation method is shown;

[0040] Figure 28 The syntax of the attribute data cell header according to the implementation method is shown;

[0041] Figure 29 The structure of a point cloud data transmission apparatus according to an embodiment is illustrated;

[0042] Figure 30 The structure of a point cloud data receiving device according to an embodiment is illustrated;

[0043] Figure 31 This is a flowchart illustrating a point cloud data receiving device according to an embodiment;

[0044] Figure 32 The transmission / reception of point cloud data according to the implementation method is illustrated;

[0045] Figure 33 Examples of a single-slice-based geometric tree structure and a segmented-slice-based geometric tree structure according to an embodiment are shown;

[0046] Figure 34 The hierarchical structure of the geometric coding tree and the aligned hierarchical structure of the attribute coding tree according to the embodiments are illustrated.

[0047] Figure 35 The hierarchical structure of the geometric tree and the independent hierarchical structure of the attribute coding tree according to the implementation method are illustrated;

[0048] Figure 36 The syntax for the parameter set according to the implementation method is shown;

[0049] Figure 37 The geometric data cell header according to an embodiment is shown;

[0050] Figure 38 A method for transmitting point cloud data according to an embodiment is illustrated; and

[0051] Figure 39 A method for receiving point cloud data according to an embodiment is illustrated. Detailed Implementation

[0052] Now, reference will be made in detail to preferred embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The following detailed description, given with reference to the accompanying drawings, is intended to explain exemplary embodiments of the present disclosure and not to show the only embodiments that can be implemented according to the present disclosure. The following detailed description includes specific details to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details.

[0053] While most of the terms used in this disclosure are selected from commonly used terms in the art, some terms have been arbitrarily chosen by the applicant, and their meanings will be explained in detail in the following description as needed. Therefore, this disclosure should be understood based on the literal meaning of the terms rather than their simple names or connotations.

[0054] Figure 1 An exemplary point cloud content delivery system according to an implementation is shown.

[0055] Figure 1The point cloud content providing system illustrated herein may include a transmitting device 10000 and a receiving device 10004. The transmitting device 10000 and the receiving device 10004 are capable of wired or wireless communication to transmit and receive point cloud data.

[0056] The point cloud data transmission device 10000 according to an embodiment can acquire and process point cloud video (or point cloud content) and transmit the point cloud video (or point cloud content). According to an embodiment, the transmission device 10000 may include a fixed station, a base transceiver system (BTS), a network, an artificial intelligence (AI) device and / or system, a robot, an AR / VR / XR device, and / or a server. According to an embodiment, the transmission device 10000 may include devices configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)), robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers.

[0057] The transmitting device 10000 according to the embodiment includes a point cloud video acquirer 10001, a point cloud video encoder 10002 and / or a transmitter (or communication module) 10003.

[0058] The point cloud video acquirer 10001 according to an embodiment acquires point cloud video through processing procedures such as capture, synthesis, or generation. Point cloud video is point cloud content represented by a point cloud as a set of points in 3D space, and may be referred to as point cloud video data. The point cloud video according to an embodiment may include one or more frames. A frame represents a still image / picture. Therefore, point cloud video may include point cloud images / frames / pictures, and may be referred to as point cloud images, frames, or pictures.

[0059] The point cloud video encoder 10002 according to an embodiment encodes the acquired point cloud video data. The point cloud video encoder 10002 can encode the point cloud video data based on point cloud compression coding. The point cloud compression coding according to an embodiment may include geometry-based point cloud compression (G-PCC) coding and / or video-based point cloud compression (V-PCC) coding or next-generation coding. The point cloud compression coding according to an embodiment is not limited to the above embodiments. The point cloud video encoder 10002 can output a bitstream containing the encoded point cloud video data. The bitstream may contain not only the encoded point cloud video data but also signaling information related to the encoding of the point cloud video data.

[0060] According to an embodiment, transmitter 10003 transmits a bitstream containing encoded point cloud video data. The bitstream, according to an embodiment, is encapsulated in a file or segment (e.g., a streaming segment) and transmitted over various networks such as broadcast networks and / or broadband networks. Although not shown in the figures, transmitting device 10000 may include an encapsulator (or encapsulation module) configured to perform encapsulation operations. According to an embodiment, the encapsulator may be included in transmitter 10003. According to an embodiment, the file or segment can be transmitted over a network to receiving device 10004 or stored in a digital storage medium (e.g., USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.). Transmitter 10003 according to an embodiment is capable of wired / wireless communication with receiving device 10004 (or receiver 10005) via networks such as 4G, 5G, and 6G. Additionally, the transmitter can perform necessary data processing operations depending on the network system (e.g., a 4G, 5G, or 6G communication network system). Transmitting device 10000 can transmit encapsulated data on demand.

[0061] The receiving device 10004 according to an embodiment includes a receiver 10005, a point cloud video decoder 10006, and / or a renderer 10007. According to an embodiment, the receiving device 10004 may include devices, robots, vehicles, AR / VR / XR devices, portable devices, home appliances, Internet of Things (IoT) devices, and AI devices / servers configured to communicate with base stations and / or other wireless devices using radio access technologies (e.g., 5G New RAT (NR), Long Term Evolution (LTE)).

[0062] According to an embodiment, receiver 10005 receives a bitstream containing point cloud video data or a file / segment encapsulating the bitstream from a network or storage medium. Receiver 10005 can perform necessary data processing according to the network system (e.g., 4G, 5G, 6G, etc. communication network systems). According to an embodiment, receiver 10005 can decapsulate the received file / segment and output the bitstream. According to an embodiment, receiver 10005 may include a decapsulator (or decapsulator module) configured to perform a decapsulation operation. The decapsulator may be implemented as a separate element (or component) from receiver 10005.

[0063] The point cloud video decoder 10006 decodes a bitstream containing point cloud video data. The point cloud video decoder 10006 can decode the point cloud video data according to a method used to encode the point cloud video data (e.g., in the inverse process of the operation of the point cloud video encoder 10002). Therefore, the point cloud video decoder 10006 can decode the point cloud video data by performing point cloud decompression encoding, which is the inverse process of point cloud compression. Point cloud decompression encoding includes G-PCC encoding.

[0064] Renderer 10007 renders the decoded point cloud video data. Renderer 10007 can output point cloud content by rendering not only the point cloud video data but also the audio data. According to one embodiment, renderer 10007 may include a display configured to display the point cloud content. According to another embodiment, the display may be implemented as a separate device or component, rather than being included in renderer 10007.

[0065] The arrows indicated by the dashed lines in the diagram represent the transmission paths of the feedback information acquired by the receiving device 10004. The feedback information reflects the interactivity of the user consuming the point cloud content and includes information about the user (e.g., head orientation information, viewport information, etc.). Specifically, when the point cloud content is the content of a service requiring user interaction (e.g., autonomous driving service, etc.), the feedback information can be provided to the content sender (e.g., the sending device 10000) and / or the service provider. Depending on the implementation, the feedback information may be used in both the receiving device 10004 and the sending device 10000, or it may not be provided.

[0066] The head orientation information according to the embodiment is information about the user's head position, orientation, angle, movement, etc. The receiving device 10004 according to the embodiment can calculate viewport information based on the head orientation information. The viewport information can be information related to the area of ​​the point cloud video that the user is viewing. The viewpoint is the point through which the user is viewing the point cloud video, and can refer to the center point of the viewport area. That is, the viewport is the area centered on the viewpoint, and the size and shape of this area can be determined by the field of view (FOV). Therefore, in addition to the head orientation information, the receiving device 10004 can also extract viewport information based on the vertical or horizontal FOV supported by the device. Furthermore, the receiving device 10004 performs gaze analysis, etc., to examine the way the user consumes the point cloud, the area the user gazes at in the point cloud video, the gaze duration, etc. According to the embodiment, the receiving device 10004 can send feedback information including the gaze analysis results to the transmitting device 10000. The feedback information according to the embodiment can be obtained during rendering and / or display processing. According to the embodiment, the feedback information can be obtained by one or more sensors included in the receiving device 10004. According to the embodiment, the feedback information can be obtained by the renderer 10007 or a separate external component (or device, part, etc.). Figure 1The dashed lines in the diagram represent the processing of feedback information received by renderer 10007. The point cloud content providing system can process (encode / decode) point cloud data based on the feedback information. Therefore, point cloud video decoder 10006 can perform decoding operations based on the feedback information. Receiving device 10004 can send feedback information to transmitting device 10000. Transmitting device 10000 (or point cloud video encoder 10002) can perform encoding operations based on the feedback information. Therefore, the point cloud content providing system can efficiently process necessary data (e.g., point cloud data corresponding to the user's head position) based on feedback information instead of processing (encoding / decoding) the entire point cloud data, and provide point cloud content to the user.

[0067] According to the implementation, the transmitting device 10000 may be referred to as an encoder, transmitting device, transmitter, etc., and the receiving device 10004 may be referred to as a decoder, receiving device, receiver, etc.

[0068] (Through a series of processes including acquisition / encoding / sending / decoding / rendering) according to the implementation method Figure 1 The point cloud data processed in the point cloud content provision system can be referred to as point cloud content data or point cloud video data. According to the implementation, point cloud content data can be used as a concept encompassing metadata or signaling information related to the point cloud data.

[0069] Figure 1 The components of the point cloud content providing system illustrated herein can be implemented by hardware, software, processors, and / or combinations thereof.

[0070] Figure 2 This is a block diagram illustrating the operation of providing point cloud content according to an implementation method.

[0071] Figure 2 The block diagram shows Figure 1 The operation of the point cloud content providing system described herein. As mentioned above, the point cloud content providing system can process point cloud data based on point cloud compression encoding (e.g., G-PCC).

[0072] A point cloud content providing system according to an embodiment (e.g., point cloud sending device 10000 or point cloud video acquirer 10001) can acquire point cloud video (20000). The point cloud video is represented by a point cloud belonging to a coordinate system used to represent 3D space. The point cloud video according to an embodiment may include a Ply (polygon file format or Stanford triangle format) file. When the point cloud video has one or more frames, the acquired point cloud video may include one or more Ply files. The Ply file contains point cloud data such as point geometry and / or attributes. Geometry includes the position of the points. The position of each point may be represented by parameters (e.g., values ​​of the X, Y, and Z axes) representing a three-dimensional coordinate system (e.g., a coordinate system consisting of X, Y, and Z axes). Attributes include the attributes of the points (e.g., information about the texture, color (YCbCr or RGB), reflectivity r, transparency, etc., of each point). A point has one or more attributes. For example, a point may have an attribute as color or two attributes as color and reflectivity. According to the implementation, geometry can be referred to as location, geometric information, geometric data, etc., and attributes can be referred to as attributes, attribute information, attribute data, etc. A point cloud content providing system (e.g., a point cloud sending device 10000 or a point cloud video acquirer 10001) can obtain point cloud data from information related to the acquisition and processing of point cloud video (e.g., depth information, color information, etc.).

[0073] A point cloud content providing system (e.g., a transmitting device 10000 or a point cloud video encoder 10002) according to an embodiment can encode point cloud data (20001). The point cloud content providing system can encode point cloud data based on point cloud compression encoding. As described above, point cloud data can include the geometry and attributes of points. Therefore, the point cloud content providing system can perform geometry encoding to encode the geometry and output a geometry bitstream. The point cloud content providing system can perform attribute encoding to encode the attributes and output an attribute bitstream. According to an embodiment, the point cloud content providing system can perform attribute encoding based on geometry encoding. The geometry bitstream and attribute bitstream according to an embodiment can be multiplexed and output as a single bitstream. The bitstream according to an embodiment may also contain signaling information related to geometry encoding and attribute encoding.

[0074] A point cloud content providing system according to an embodiment (e.g., transmitting device 10000 or transmitter 10003) can transmit encoded point cloud data (20002). For example... Figure 1As illustrated, the encoded point cloud data can be represented by a geometric bitstream and an attribute bitstream. Additionally, the encoded point cloud data can be transmitted as a bitstream along with signaling information related to the encoding of the point cloud data (e.g., signaling information related to geometric and attribute encoding). The point cloud content providing system can encapsulate the bitstream carrying the encoded point cloud data and transmit the bitstream as a file or segment.

[0075] The point cloud content providing system (e.g., receiving device 10004 or receiver 10005) according to the embodiment can receive a bitstream containing encoded point cloud data. Additionally, the point cloud content providing system (e.g., receiving device 10004 or receiver 10005) can demultiplex the bitstream.

[0076] A point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode encoded point cloud data (e.g., geometric bitstream, attribute bitstream) transmitted in a bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode point cloud video data based on signaling information related to the encoding of point cloud video data contained in the bitstream. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can decode the geometric bitstream to reconstruct the location (geometry) of points. The point cloud content providing system can reconstruct the attributes of points by decoding the attribute bitstream based on the reconstructed geometry. The point cloud content providing system (e.g., receiving device 10004 or point cloud video decoder 10005) can reconstruct point cloud video based on the location of the reconstructed geometry and the decoded attributes.

[0077] A point cloud content providing system (e.g., receiving device 10004 or renderer 10007) according to an embodiment can render decoded point cloud data (20004). The point cloud content providing system (e.g., receiving device 10004 or renderer 10007) can use various rendering methods to render the geometry and attributes decoded through the decoding process. Points in the point cloud content can be rendered as vertices with a certain thickness, cubes with a specific minimum size centered at the corresponding vertex position, or circles centered at the corresponding vertex position. All or part of the rendered point cloud content is provided to a user through a display (e.g., a VR / AR display, a conventional display, etc.).

[0078] The point cloud content providing system (e.g., receiving device 10004) according to the embodiment can obtain feedback information (20005). The point cloud content providing system can encode and / or decode point cloud data based on the feedback information. The feedback information and operation and reference of the point cloud content providing system according to the embodiment... Figure 1The feedback information and operation described are the same, so a detailed description of them is omitted.

[0079] Figure 3 An exemplary process for capturing point cloud video according to an implementation method is illustrated.

[0080] Figure 3 Examples of references are provided. Figures 1 to 2 The described point cloud content provides an example of point cloud video capture processing for the system.

[0081] Point cloud content includes point cloud videos (images and / or videos) representing objects and / or environments located in various 3D spaces (e.g., 3D spaces representing real environments, 3D spaces representing virtual environments, etc.). Therefore, the point cloud content providing system according to embodiments can use one or more cameras (e.g., infrared cameras capable of acquiring depth information, RGB cameras capable of extracting color information corresponding to the depth information, etc.), projectors (e.g., infrared pattern projectors for acquiring depth information), LiDRA, etc., to capture point cloud videos. The point cloud content providing system according to embodiments can extract the shape of the geometry composed of points in 3D space from the depth information and extract the attributes of each point from the color information to obtain point cloud data. Images and / or videos according to embodiments can be captured based on at least one of inward-oriented and outward-oriented techniques.

[0082] Figure 3 The left side illustrates inward-facing technology. Inward-facing technology refers to the technique of capturing images of a central object using one or more cameras (or camera sensors) positioned around it. Inward-facing technology can be used to generate point cloud content that provides users with 360-degree images of key objects (e.g., VR / AR content that provides users with 360-degree images of objects such as characters, players, objects, or actors).

[0083] Figure 3 The right side illustrates outward-facing techniques. Outward-facing techniques refer to techniques that capture images of the environment of a central object, rather than the central object itself, using one or more cameras (or camera sensors) positioned around it. Outward-facing techniques can be used to generate point cloud content that provides the surrounding environment from a user's perspective (e.g., content representing the external environment that can be provided to users of autonomous vehicles).

[0084] As shown in the figure, point cloud content can be generated based on the capture operations of one or more cameras. In this case, the coordinate system is different in each camera; therefore, the point cloud content providing system can calibrate one or more cameras to set the global coordinate system before the capture operation. Alternatively, the point cloud content providing system can generate point cloud content by compositing arbitrary images and / or videos with images and / or videos captured using the aforementioned capture techniques. The point cloud content providing system may not perform this step when generating point cloud content representing virtual space. Figure 3 The capture operations described herein. The point cloud content providing system according to an embodiment can perform post-processing on the captured images and / or videos. In other words, the point cloud content providing system can remove unwanted areas (e.g., background), identify spaces to which the captured images and / or videos are connected, and perform a space-hole filling operation when space holes exist.

[0085] A point cloud content delivery system can generate point cloud content by performing coordinate transformations on points in a point cloud video obtained from each camera. The system can perform coordinate transformations on points based on the coordinates of each camera's location. Therefore, the system can generate content representing a wide range or point cloud content with high-density points.

[0086] Figure 4 An exemplary point cloud encoder according to an implementation method is illustrated.

[0087] Figure 4 It shows Figure 1 An example of a point cloud video encoder 10002. The point cloud encoder reconstructs and encodes point cloud data (e.g., point locations and / or attributes) to adjust the quality of the point cloud content (e.g., lossless, lossy, or near-lossless) based on network conditions or applications. When the total size of the point cloud content is large (e.g., 60Gbps of point cloud content for 30fps), the point cloud content providing system may not be able to stream the content in real time. Therefore, the point cloud content providing system can reconstruct the point cloud content based on a maximum target bitrate to provide the point cloud content according to network conditions, etc.

[0088] For reference Figures 1 to 2 The point cloud encoder described herein can perform geometric encoding and attribute encoding. Geometric encoding is performed before attribute encoding.

[0089] The point cloud encoder according to the implementation includes a coordinate transformer (transform coordinates) 40000, a quantizer (quantize and remove points (voxarization)) 40001, an octree analyzer (analyze octrees) 40002, a surface approximation analyzer (analyze surface approximations) 40003, an arithmetic encoder (arithmetic coding) 40004, a geometry reconstructor (reconstruct geometry) 40005, a color transformer (transform colors) 40006, an attribute transformer (transform attributes) 40007, a RAHT transformer (RAHT) 40008, an LOD generator (generate LODs) 40009, a lift transformer (lift) 40010, a coefficient quantizer (quantize coefficients) 40011, and / or an arithmetic encoder (arithmetic coding) 40012.

[0090] Coordinate transformer 40000, quantizer 40001, octree analyzer 40002, surface approximation analyzer 40003, arithmetic encoder 40004, and geometric reconstructor 40005 can perform geometric encoding. Geometric encoding according to the implementation may include octree geometric encoding, direct encoding, trisoup geometric encoding, and entropy encoding. Direct encoding and trisoup geometric encoding are applied selectively or in combination. Geometric encoding is not limited to the examples described above.

[0091] As shown in the figure, the coordinate transformer 40000 according to the embodiment receives the position and transforms it into coordinates. For example, the position can be transformed into position information in three-dimensional space (e.g., three-dimensional space represented by the XYZ coordinate system). The position information in three-dimensional space according to the embodiment can be referred to as geometric information.

[0092] The quantizer 40001 according to the embodiment quantizes geometry. For example, the quantizer 40001 can quantize points based on the minimum position value of all points (e.g., the minimum value on each of the X, Y, and Z axes). The quantizer 40001 performs the following quantization operation: multiplying the difference between the position value of each point and the minimum position value by a preset quantization scaling value, and then finding the nearest integer value by rounding the value obtained by multiplication. Thus, one or more points can have the same quantized position (or position value). The quantizer 40001 according to the embodiment performs voxelization based on the quantized position to reconstruct the quantized points. As in the case of a pixel, which is the smallest unit containing 2D image / video information, the points of the point cloud content (or 3D point cloud video) according to the embodiment can be included in one or more voxels. The term voxel, a compound word of volume and pixel, refers to the 3D cubic space generated when 3D space is divided into units (unit = 1.0) based on axes representing 3D space (e.g., X-axis, Y-axis, and Z-axis). The quantizer 40001 can match groups of points in 3D space to voxels. According to one implementation, a voxel may include only one point. According to another implementation, a voxel may include one or more points. To represent a voxel as a point, the center position of the voxel can be set based on the positions of the one or more points included in the voxel. In this case, attributes of all positions included in a voxel can be combined and assigned to that voxel.

[0093] The octree analyzer 40002 according to the implementation performs octree geometric encoding (or octree encoding) to represent voxels in an octree structure. The octree structure represents points based on the octree structure and voxel matching.

[0094] The surface approximation analyzer 40003 according to the embodiment can analyze and approximate octrees. The octree analysis and approximation according to the embodiment analyzes regions containing multiple points to efficiently provide octree and voxelization processing.

[0095] The arithmetic encoder 40004 according to the embodiment performs entropy encoding on an octree and / or an approximate octree. For example, the encoding scheme includes arithmetic encoding. As a result of the encoding, a geometric bitstream is generated.

[0096] The attribute encoding is performed by a color transformer 40006, an attribute transformer 40007, a RAHT transformer 40008, an LOD generator 40009, a boosting transformer 40010, a coefficient quantizer 40011, and / or an arithmetic encoder 40012. As described above, a point can have one or more attributes. The attribute encoding according to the embodiment is also applied to the attributes a point possesses. However, when an attribute (e.g., color) comprises one or more elements, the attribute encoding is applied independently to each element. The attribute encoding according to the embodiment includes color transformation encoding, attribute transformation encoding, region adaptive hierarchical transformation (RAHT) encoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) encoding, and interpolation-based hierarchical nearest neighbor prediction encoding with an update / boosting step (boosting transformation). Depending on the point cloud content, the RAHT encoding, prediction transformation encoding, and boosting transformation encoding described above can be used selectively, or a combination of one or more encoding schemes can be used. The attribute encoding according to the embodiment is not limited to the examples described above.

[0097] The color converter 40006 according to the embodiment performs color transformation encoding on the color values ​​(or textures) included in the transformation attributes. For example, the color converter 40006 can transform the format of color information (e.g., from RGB to YCbCr). The operation of the color converter 40006 according to the embodiment can be optionally applied according to the color values ​​included in the attributes.

[0098] The geometry reconstructor 40005, according to the implementation method, reconstructs (decompresses) octrees and / or approximate octrees. The geometry reconstructor 40005 reconstructs the octree / voxel based on the distribution of analysis points. The reconstructed octree / voxel can be referred to as the reconstructed geometry (recovered geometry).

[0099] The attribute transformer 40007 according to the embodiment performs attribute transformation to transform attributes based on the location and / or reconstructed geometry that has not undergone geometric encoding. As described above, since attributes depend on geometry, the attribute transformer 40007 can transform attributes based on reconstructed geometric information. For example, based on the position values ​​of points included in a voxel, the attribute transformer 40007 can transform the attributes of points at that location. As described above, when the position of the voxel center is set based on the positions of one or more points included in the voxel, the attribute transformer 40007 transforms the attributes of said one or more points. When performing trigonometric Thomson geometric encoding, the attribute transformer 40007 can transform attributes based on the trigonometric Thomson geometric encoding.

[0100] The attribute transformer 40007 performs attribute transformation by calculating the average of the attributes or attribute values ​​(e.g., color or reflectivity of each point) of neighboring points within a specific position / radius from the center (or position value) of each voxel. The attribute transformer 40007 can apply weights based on the distance from the center to each point when calculating the average. Therefore, each voxel has a position and a calculated attribute (or attribute value).

[0101] The attribute transformer 40007 can search for nearest neighbors within a specific location / radius of the center of each voxel based on a KD-tree or Morton code. A KD-tree is a binary search tree and supports a data structure that allows points to be managed based on location so that nearest neighbor search (NNS) can be performed quickly. Morton codes are generated by representing the coordinates (e.g., (x, y, z)) of the 3D location of all points as bit values ​​and mixing those bits. For example, when the coordinates representing the location of a point are (5, 9, 1), the bit values ​​for the coordinates are (0101, 1001, 0001). The bit values ​​are mixed according to the bit indices in the order of z, y, and x to produce 010001000111. This value is represented as the decimal number 1095. That is, the Morton code value for the point with coordinates (5, 9, 1) is 1095. The attribute transformer 40007 can sort the points based on their Morton code values ​​and perform NNS using a depth-first traversal process. After an attribute transformation operation, if an NNS is required in another transformation process used for attribute encoding, use a KD tree or Morton code.

[0102] As shown in the figure, the transformed attributes are input to the RAHT transformer 40008 and / or the LOD generator 40009.

[0103] According to the implementation, the RAHT transformer 40008 performs RAHT encoding for predicting attribute information based on the reconstructed geometric information. For example, the RAHT transformer 40008 can predict the attribute information of higher-level nodes in an octree based on the attribute information associated with lower-level nodes in the octree.

[0104] The LOD generator 40009 according to the embodiment generates a Level of Detail (LOD) to perform predictive transform coding. The LOD according to the embodiment represents the level of detail of the point cloud content. A decreasing LOD value indicates a decrease in the level of detail of the point cloud content. An increasing LOD value indicates an increase in the level of detail of the point cloud content. Points can be classified according to their LOD.

[0105] The lift transformer 40010 according to the embodiment performs lift transform coding to transform the attributes of the point cloud based on weights. As described above, lift transform coding may optionally be applied.

[0106] According to the implementation method, the coefficient quantizer 40011 quantizes the attribute after attribute encoding based on the coefficient.

[0107] According to the implementation method, the arithmetic encoder 40012 encodes the quantized attributes based on arithmetic encoding.

[0108] Although not shown in the figure, Figure 4 The elements of the point cloud encoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits, which are configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform the above-described... Figure 4 At least one of the operations and / or functions of the elements of the point cloud encoder. Additionally, one or more processors can operate on or execute a set of software programs and / or instructions to perform... Figure 4 The operation and / or function of the elements of the point cloud encoder. One or more memories according to the embodiments may include high-speed random access memory, or include non-volatile memory (e.g., one or more disk storage devices, flash storage devices or other non-volatile solid-state storage devices).

[0109] Figure 5 An example of a voxel according to an embodiment is shown.

[0110] Figure 5 This illustrates voxels in 3D space represented by a coordinate system consisting of three axes: the X-axis, Y-axis, and Z-axis. (See reference...) Figure 4 As described, a point cloud encoder (e.g., quantizer 40001) can perform voxelization. A voxel refers to the 3D cubic space generated when the 3D space is divided into cells (unit = 1.0) based on axes representing the 3D space (e.g., X-axis, Y-axis, and Z-axis). Figure 5 An example of voxels generated via an octree structure is shown, in which a bounding box aligned to the cubic axis, defined by two poles (0, 0, 0) and (2d, 2d, 2d), is recursively subdivided. A voxel comprises at least one point. The spatial coordinates of a voxel can be estimated based on its positional relationship to a group of voxels. As mentioned above, voxels possess properties similar to pixels in a 2D image / video (such as color or reflectivity). Details and references of voxels are provided. Figure 4 The details described are the same, so the description of it is omitted.

[0111] Figure 6 An example of an octree and occupancy code according to an implementation is shown.

[0112] For reference Figures 1 to 4The described point cloud content delivery system (point cloud video encoder 10002) or point cloud encoder (e.g., octree analyzer 40002) performs octree geometric encoding (or octree encoding) based on an octree structure to efficiently manage the regions and / or locations of voxels.

[0113] Figure 6 The upper part shows the octree structure. The 3D space of the point cloud content according to the embodiment is represented by the axes of the coordinate system (e.g., the X, Y, and Z axes). The octree structure is created by recursively subdividing bounding boxes aligned to the cubic axes defined by two poles (0, 0, 0) and (2d, 2d, 2d). Here, 2d can be set as the value of the minimum bounding box that constitutes all points around the point cloud content (or point cloud video). Here, d represents the depth of the octree. The value of d is determined by the following formula. In the following formula, (x int n ,y int n ,z int n ) indicates the position (or position value) of the quantization point.

[0114]

[0115] like Figure 6 As shown in the upper center, the entire 3D space can be divided into eight spaces according to partitions. Each partitioned space is represented by a cube with six faces. (See diagram below.) Figure 6 As shown in the upper right corner, each of the eight spaces is further divided based on the axes of the coordinate system (e.g., the X, Y, and Z axes). Thus, each space is divided into eight smaller spaces. These smaller spaces are also represented by cubes with six faces. This partitioning scheme is applied until the leaf nodes of the octree become voxels.

[0116] Figure 6 The lower part shows the octree occupancy code. The occupancy code of the octree is generated to indicate whether each of the eight partitioned spaces resulting from partitioning a space contains at least one point. Therefore, a single occupancy code is represented by eight child nodes. Each child node represents the occupancy of the partitioned space, and the child node has a 1-bit value. Therefore, the occupancy code is represented as an 8-bit code. That is, when the space corresponding to the child node contains at least one point, the node is assigned a value of 1. When the space corresponding to the child node does not contain a point (the space is empty), the node is assigned a value of 0. Since... Figure 6The occupancy code shown is 00100001, indicating that the spaces corresponding to the third and eighth child nodes out of eight each contain at least one point. As shown, each of the third and eighth child nodes has eight child nodes, and each child node is represented by an 8-bit occupancy code. The figure shows the occupancy code for the third child node as 10000111, and the occupancy code for the eighth child node as 01001111. A point cloud encoder (e.g., an arithmetic encoder 40004) according to an embodiment can perform entropy coding on the occupancy code. To improve compression efficiency, the point cloud encoder can perform intra / inter-frame coding on the occupancy code. A receiving device (e.g., receiving device 10004 or point cloud video decoder 10006) according to an embodiment reconstructs the octree based on the occupancy code.

[0117] A point cloud encoder according to an implementation method (e.g., Figure 4 A point cloud encoder or octree analyzer (40002) can perform voxelization and octree encoding to store the locations of points. However, points are not always uniformly distributed in 3D space, so there will be specific regions with fewer points. Therefore, performing voxelization over the entire 3D space is inefficient. For example, when a specific region contains fewer points, it is not necessary to perform voxelization in that specific region.

[0118] Therefore, for the specific region mentioned above (or nodes other than the leaf nodes of the octree), the point cloud encoder according to the embodiment can skip voxelization and perform direct encoding to directly encode the positions of points included in the specific region. The coordinates of the directly encoded points according to the embodiment are called the Direct Encoding Mode (DCM). The point cloud encoder according to the embodiment can also perform trigonometric encoding based on the surface model to reconstruct the positions of points in the specific region (or node) based on voxels. Trigonometric encoding is a geometric encoding that represents an object as a series of triangular meshes. Therefore, the point cloud decoder can generate a point cloud from the mesh surface. Trigonometric encoding and direct encoding according to the embodiment can be selectively performed. In addition, trigonometric encoding and direct encoding according to the embodiment can be performed in combination with octree geometric encoding (or octree coding).

[0119] To perform direct encoding, the option to apply direct encoding using direct mode should be enabled. The node to be directly encoded is not a leaf node, and there should be fewer than a threshold number of points within that node. Furthermore, the total number of points to be directly encoded should not exceed a preset threshold. When the above conditions are met, the point cloud encoder (or arithmetic encoder 40004) according to the implementation can perform entropy encoding on the point positions (or position values).

[0120] A point cloud encoder according to an embodiment (e.g., a surface approximation analyzer 40003) can determine a specific level of the octree (a level less than the depth d of the octree) and can perform trigonometric tangent coding using a surface model starting from that level to reconstruct the location of points in the region of a node based on voxels (trigonometric tangent mode). The point cloud encoder according to an embodiment can specify the level at which trigonometric tangent coding will be applied. For example, when the specific level is equal to the depth of the octree, the point cloud encoder does not operate in trigonometric tangent mode. In other words, the point cloud encoder according to an embodiment can operate in trigonometric tangent mode only when the specified level is less than the depth value of the octree. The 3D cubic region of a node at the specified level according to an embodiment is called a block. A block may include one or more voxels. A block or voxel may correspond to a brick. Geometry is represented by a surface within each block. A surface according to an embodiment may intersect each edge of a block at most once.

[0121] A block has 12 edges, therefore there are at least 12 intersections within a block. Each intersection is called a vertex (or apex point). A vertex is detected along an edge when there is at least one occupied voxel adjacent to that edge in all blocks sharing that edge. An occupied voxel, according to the implementation, refers to a voxel containing a point. The position of a vertex detected along an edge is the average position of the edges of all voxels adjacent to that edge in all blocks sharing that edge.

[0122] Once a vertex is detected, the point cloud encoder according to the embodiment can perform entropy encoding on the starting point (x, y, z) of the edge, the direction vector (Δx, Δy, Δz) of the edge, and the vertex position value (relative position value within the edge). When applying trigonometric Tang geometry encoding, the point cloud encoder according to the embodiment (e.g., geometry reconstructor 40005) can generate the restored geometry (reconstructed geometry) by performing trigonometric reconstruction, upsampling, and voxelization.

[0123] Vertices at the edges of a block define the surface traversing the block. The surface, according to the implementation, is a non-planar polygon. In the triangulation process, the surface represented by triangles is reconstructed based on the starting point of the edge, the direction vector of the edge, and the position values ​​of the vertices. The triangulation process is performed by: i) calculating the centroid value of each vertex, ii) subtracting the centroid value from each vertex value, and iii) estimating the sum of squares of the values ​​obtained through the subtraction.

[0124]

[0125] Estimate the minimum value of the sum and perform projection processing based on the axis with the minimum value. For example, when element x is minimum, each vertex is projected onto the x-axis relative to the center of the block and onto the (y,z) plane. When the value obtained by projecting onto the (y,z) plane is (ai,bi), the value of θ is estimated by atan2(bi,ai), and the vertices are sorted according to the value of θ. Table 1 below shows the vertex combinations for creating triangles based on the number of vertices. The vertices are sorted from 1 to n. Table 1 below shows that for four vertices, two triangles can be constructed based on the combinations of vertices. The first triangle can be composed of vertices 1, 2, and 3 from the sorted vertices, and the second triangle can be composed of vertices 3, 4, and 1 from the sorted vertices.

[0126] Table 2-1. Triangles formed by vertices sorted by 1, ..., n

[0127]

[0128] Upsampling is performed to add points along the edges of the triangle at the center and voxelization is then performed. The added points are generated based on the upsampling factor and the width of the block. The added points are called thinned vertices. A point cloud encoder according to an implementation can voxelize the thinned vertices. Additionally, the point cloud encoder can perform attribute encoding based on the voxelized locations (or location values).

[0129] Figure 7 An example of a neighbor node pattern according to an implementation method is shown.

[0130] To improve the compression efficiency of point cloud videos, the point cloud encoder according to the implementation method can perform entropy coding based on context-adaptive arithmetic coding.

[0131] For reference Figures 1 to 6 The described point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, point cloud encoder or...) Figure 4 The arithmetic encoder 40004 can immediately perform entropy coding on the occupancy code. Alternatively, the point cloud content providing system or point cloud encoder can perform entropy coding (intra-frame coding) based on the occupancy code of the current node and the occupancy of neighboring nodes, or entropy coding (inter-frame coding) based on the occupancy code of the previous frame. According to the embodiment, a frame represents a collection of simultaneously generated point cloud videos. The compression efficiency of the intra-frame coding / inter-frame coding according to the embodiment can depend on the number of referenced neighboring nodes. As the number of bits increases, the operation becomes more complex, but the coding can be biased to one side, thereby increasing the compression efficiency. For example, when given a 3-bit context, 2... 3 = There are 8 methods to perform the encoding. The division of space for encoding affects the complexity of the implementation. Therefore, an appropriate level of compression efficiency and complexity must be achieved.

[0132] Figure 7 This illustrates a process for obtaining occupancy patterns based on the occupancy of neighboring nodes. A point cloud encoder, according to an implementation, determines the occupancy of the neighboring nodes of each node in an octree and obtains the value of the neighboring node pattern. This neighboring node pattern is then used to infer the occupancy pattern of the node. Figure 7 The left side of the diagram shows the cube corresponding to the node (the cube in the middle) and six cubes sharing at least one face with the cube (neighboring nodes). The nodes shown in the diagram are nodes at the same depth. The numbers shown in the diagram represent the weights associated with the six nodes (1, 2, 4, 8, 16, and 32). Weights are assigned sequentially based on the position of the neighboring nodes.

[0133] Figure 7 The right side of the diagram shows the neighbor node pattern values. The neighbor node pattern value is the sum of the values ​​multiplied by the weights of occupied neighbor nodes (neighbor nodes with points). Therefore, the neighbor node pattern values ​​range from 0 to 63. When the neighbor node pattern value is 0, it indicates that none of the node's neighbors have a point (unoccupied node). When the neighbor node pattern value is 63, it indicates that all neighbor nodes are occupied nodes. As shown in the diagram, since the neighbor nodes assigned weights 1, 2, 4, and 8 are occupied nodes, the neighbor node pattern value is 15, which is the sum of 1, 2, 4, and 8. The point cloud encoder can perform encoding based on the neighbor node pattern values ​​(e.g., 64 encodings can be performed when the neighbor node pattern value is 63). According to implementations, the point cloud encoder can reduce encoding complexity by changing the neighbor node pattern values ​​(e.g., based on a table that changes 64 to 10 or 6).

[0134] Figure 8 An example of point configuration in each LOD according to the implementation method is shown.

[0135] For reference Figures 1 to 7 The description describes the reconstruction (decompression) of encoded geometry before performing attribute encoding. When direct encoding is applied, geometry reconstruction operations may include changing the placement of directly encoded points (e.g., placing directly encoded points in front of the point cloud data). When triangulation geometry encoding is applied, geometry reconstruction processing is performed through triangulation, upsampling, and voxelization. Since attributes depend on geometry, attribute encoding is performed based on the reconstructed geometry.

[0136] Point cloud encoders (e.g., LOD generator 40009) can classify (reorganize) points using LOD. This figure illustrates the point cloud content corresponding to LOD. The leftmost image in the figure represents the original point cloud content. The second image from the left shows the distribution of points in the lowest LOD, and the rightmost image shows the distribution of points in the highest LOD. That is, points are sparsely distributed in the lowest LOD and densely distributed in the highest LOD. In other words, as the LOD increases in the direction indicated by the arrow at the bottom of the figure, the space (or distance) between points narrows.

[0137] Figure 9 An example of point configuration for each LOD according to the implementation method is shown.

[0138] For reference Figures 1 to 8 The described point cloud content providing system or point cloud encoder (e.g., point cloud video encoder 10002, ...) Figure 4 A point cloud encoder or LOD generator (40009) can generate LODs. LODs are generated by reorganizing points into a set of refinement levels based on a set of LOD distance values ​​(or a set of Euclidean distances). The LOD generation process is performed not only by the point cloud encoder but also by the point cloud decoder.

[0139] Figure 9 The upper part shows examples of points (P0 to P9) of point cloud content distributed in 3D space. Figure 9 In this context, the original order represents the order of points P0 to P9 before LOD generation. Figure 9 In this context, LOD-based order represents the order in which points are generated according to their LOD values. Points are reorganized using LOD. Furthermore, higher LOD values ​​include points belonging to lower LOD values. For example... Figure 9 As shown, LOD0 contains P0, P5, P4, and P2. LOD1 contains the points of LOD0, P1, P6, and P3. LOD2 contains the points of LOD0, the points of LOD1, P9, P8, and P7.

[0140] For reference Figure 4 The point cloud encoder described herein, according to the embodiments, can selectively or in combination perform predictive transform coding, lifting transform coding, and RAHT transform coding.

[0141] The point cloud encoder according to the implementation can generate predictors for points by performing predictive transform coding to set the predictive attributes (or predictive attribute values) for each point. That is, N predictors can be generated for N points. The predictors according to the implementation can calculate weights (= 1 / distance) based on the LOD value of each point, indexed information related to neighboring points existing within a set distance for each LOD, and the distance to the neighboring points.

[0142] According to the implementation, the predicted attribute (or attribute value) is set as the average of the values ​​obtained by multiplying the attributes (or attribute values) of neighboring points (e.g., color, reflectivity, etc.) set in the predictor of each point by a weight (or weight value) calculated based on the distance to each neighboring point. The point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation can quantize and dequantize the residual (which may be referred to as residual attribute, residual attribute value, or attribute prediction residual) obtained by subtracting the predicted attribute (attribute value) from the attribute (attribute value) of each point. This quantization process is configured as shown in the table below.

[0143] Table. Pseudocode for Attribute Prediction Residual Quantization

[0144]

[0145]

[0146] Table. Pseudocode for inverse quantization of attribute prediction residuals

[0147]

[0148] When the predictor for each point has neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the residual values ​​after quantization and dequantization as described above. When the predictor for each point has no neighboring points, the point cloud encoder (e.g., arithmetic encoder 40012) according to the embodiment can perform entropy encoding on the attributes of the corresponding point without performing the above operations.

[0149] The point cloud encoder (e.g., lift transformer 40010) according to the embodiment can generate a predictor for each point, set the calculated LOD and register neighboring points in the predictor, and set weights according to the distance to the neighboring points to perform lift transform coding. The lift transform coding according to the embodiment is similar to the prediction transform coding described above, but the difference is that the weights are applied cumulatively to the attribute values. The process of applying weights cumulatively to attribute values ​​according to the embodiment is configured as follows.

[0150] 1) Create an array QuantizationWeight(QW) to store the weight value for each point. All elements of QW are initialized to 1.0. The QW value of the predictor index of the neighboring nodes registered in the predictor is multiplied by the predictor weight of the current point, and the resulting values ​​are summed.

[0151] 2) Improved prediction processing: Subtract the value obtained by multiplying the attribute value of the point by the weight from the existing attribute value to calculate the predicted attribute value.

[0152] 3) Create a temporary array called updateweight, and update and initialize the temporary array to zero.

[0153] 4) The weights calculated by multiplying the weights computed for all predictors by the weights stored in QW corresponding to the predictor indices are cumulatively added to the update weight array as the indices of neighboring nodes. The values ​​obtained by multiplying the attribute values ​​of the neighboring node indices by the calculated weights are cumulatively added to the update array.

[0154] 5) Improved update processing: Divide the attribute values ​​of the update array for all predictors by the weight values ​​of the update weight array of the predictor index, and add the existing attribute values ​​to the values ​​obtained by division.

[0155] 6) The predicted attribute is calculated by multiplying the attribute value updated by the boost update process for all predictors by the weight updated by the boost prediction process (stored in the QW). The predicted attribute value is quantized by a point cloud encoder (e.g., coefficient quantizer 40011) according to the implementation. In addition, the point cloud encoder (e.g., arithmetic encoder 40012) performs entropy encoding on the quantized attribute value.

[0156] A point cloud encoder according to an embodiment (e.g., RAHT transform 40008) can perform RAHT transform coding, where attributes associated with lower-level nodes in an octree are used to predict attributes of higher-level nodes. RAHT transform coding is an example of intra-frame attribute coding performed via a backward scan of an octree. The point cloud encoder according to an embodiment scans the entire region from voxels and repeats a merging process in each step, merging voxels into larger blocks, until the root node is reached. The merging process according to an embodiment is performed only on occupied nodes. The merging process is not performed on empty nodes. The merging process is performed in a higher mode directly above an empty node.

[0157] The following equation represents the RAHT transformation matrix. In this equation, This represents the average attribute value of the voxel at level l. It can be based on... and To calculate and The weight is and

[0158]

[0159] here, It is a low-pass value and is used in the next higher level of merge processing. This represents the high-pass coefficient. The high-pass coefficient in each step is quantized and undergoes entropy encoding (e.g., via an arithmetic encoder 40012). Weights are calculated as follows: pass and The root node is calculated as follows.

[0160]

[0161] Figure 10 An example of a point cloud decoder according to an implementation method is shown.

[0162] Figure 10 The point cloud decoder shown in the example is Figure 1 The example of the point cloud video decoder 10006 described in [the document], and can perform [operations] with [other functions]. Figure 1 The point cloud video decoder 10006 illustrated in the figure operates in the same or similar manner. As shown in the figure, the point cloud decoder can receive a geometry bitstream and an attribute bitstream contained in one or more bitstreams. The point cloud decoder includes a geometry decoder and an attribute decoder. The geometry decoder performs geometry decoding on the geometry bitstream and outputs the decoded geometry. The attribute decoder performs attribute decoding based on the decoded geometry and the attribute bitstream and outputs the decoded attributes. The decoded geometry and the decoded attributes are used to reconstruct the point cloud content (the decoded point cloud).

[0163] Figure 11 An example of a point cloud decoder according to an implementation method is shown.

[0164] Figure 11 The point cloud decoder shown in the example is Figure 10 The example shown is a point cloud decoder, which can be executed as... Figures 1 to 9 The example above illustrates the decoding operation, which is the inverse of the encoding operation of a point cloud encoder.

[0165] For reference Figure 1 and Figure 10 As described, the point cloud decoder can perform geometry decoding and attribute decoding. Geometry decoding is performed before attribute decoding.

[0166] The point cloud decoder according to the implementation includes an arithmetic decoder (arithmetic decoding) 11000, an octree synthesizer (synthesized octree) 11001, a surface approximation synthesizer (synthesized surface approximation) 11002, a geometry reconstructor (reconstructed geometry) 11003, an inverse coordinate transformer (inverse coordinate transformation) 11004, an arithmetic decoder (arithmetic decoding) 11005, an inverse quantizer (inverse quantization) 11006, a RAHT transformer 11007, an LOD generator (generated LOD) 11008, an inverse lifter (inverse lift) 11009, and / or an inverse color transformer (inverse color transformation) 11010.

[0167] Arithmetic decoder 11000, octree synthesizer 11001, surface approximation synthesizer 11002, geometric reconstructor 11003, and coordinate inverse transformer 11004 can perform geometric decoding. Geometric decoding according to embodiments may include direct encoding and trigonometric Thomson geometric decoding. Direct encoding and trigonometric Thomson geometric decoding are selectively applied. Geometric decoding is not limited to the examples described above and is provided for reference only. Figures 1 to 9 The inverse processing of the described geometric encoding is performed.

[0168] The arithmetic decoder 11000 according to the embodiment decodes the received geometric bitstream based on arithmetic coding. The operation of the arithmetic decoder 11000 corresponds to the inverse processing of the arithmetic encoder 40004.

[0169] The octree synthesizer 11001 according to the embodiment can generate an octree by obtaining occupancy codes from the decoded geometric bitstream (or geometric information obtained as a decoding result). See reference... Figures 1 to 9 Configure the occupancy code in detail.

[0170] When applying trisoup geometry encoding, the surface approximation synthesizer 11002 according to the implementation can synthesize the surface based on the decoded geometry and / or the generated octree.

[0171] The geometry reconstructor 11003 according to the embodiment can regenerate geometry based on surfaces and / or decoded geometry. See reference... Figures 1 to 9 As described, direct encoding and trigonometric Tangle geometry encoding are selectively applied. Therefore, geometry reconstructor 11003 directly imports and adds positional information about points where direct encoding has been applied. When trigonometric Tangle geometry encoding is applied, geometry reconstructor 11003 can reconstruct the geometry by performing reconstruction operations (e.g., triangulation, upsampling, and voxelization) of geometry reconstructor 40005. Details and References Figure 6 The details described are the same, so their description is omitted. The reconstructed geometry may include point cloud images or frames that do not contain attributes.

[0172] According to the implementation, the inverse coordinate transformer 11004 can obtain the position of a point based on the reconstructed geometric transformation coordinates.

[0173] Arithmetic decoder 11005, dequantizer 11006, RAHT transformer 11007, LOD generator 11008, inverse booster 11009, and / or inverse color transformer 11010 can perform reference... Figure 10The attribute decoding described herein includes Region Adaptive Hierarchical Transformation (RAHT) decoding, interpolation-based hierarchical nearest neighbor prediction (prediction transformation) decoding, and interpolation-based hierarchical nearest neighbor prediction decoding with an update / lifting step (lifting transformation). The three decoding schemes described above may be used selectively, or a combination of one or more decoding schemes may be used. The attribute decoding according to the embodiments is not limited to the examples described above.

[0174] According to the embodiment, the arithmetic decoder 11005 decodes the attribute bitstream by arithmetic encoding.

[0175] The dequantizer 11006 according to the implementation method dequantizes the information about the decoded attribute bitstream or attribute obtained as a decoding result, and outputs the dequantized attribute (or attribute value). Dequantization can be selectively applied based on the attribute encoding of the point cloud encoder.

[0176] According to the implementation, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can process the reconstructed geometry and dequantized attributes. As described above, the RAHT transformer 11007, LOD generator 11008, and / or inverse lifter 11009 can selectively perform decoding operations corresponding to the encoding of the point cloud encoder.

[0177] The inverse color transformer 11010 according to the embodiment performs inverse transformation encoding to inversely transform the color values ​​(or textures) included in the decoded attributes. The operation of the inverse color transformer 11010 can be selectively performed based on the operation of the color transformer 40006 of the point cloud video encoder.

[0178] Although not shown in the figure, Figure 11 The elements of the point cloud decoder can be implemented by hardware, software, firmware, or a combination thereof including one or more processors or integrated circuits, which are configured to communicate with one or more memories included in the point cloud providing device. The one or more processors can perform the above-described... Figure 11 The point cloud decoder's components include at least one or more of their operations and / or functions. Additionally, one or more processors can operate on or execute software programs and / or sets of instructions to perform... Figure 11 The operation and / or functions of the components of the point cloud decoder.

[0179] Figure 12 An example of a transmitting device according to an embodiment is shown.

[0180] Figure 12 The transmitting device shown is Figure 1 The transmitting device 10000 (or Figure 4Example of a point cloud encoder. Figure 12 The transmitting device illustrated in the example can perform and reference Figures 1 to 9 The described point cloud encoder operation and method are one or more of the same or similar operations and methods. The transmitting apparatus according to the embodiment may include a data input unit 12000, a quantization processor 12001, a voxelization processor 12002, an octree occupancy code generator 12003, a surface model processor 12004, an intra / inter-frame coding processor 12005, an arithmetic encoder 12006, a metadata processor 12007, a color transformation processor 12008, an attribute transformation processor 12009, a prediction / boosting / RAHT transformation processor 12010, an arithmetic encoder 12011, and / or a transmitting processor 12012.

[0181] According to the embodiment, the data input unit 12000 receives or acquires point cloud data. The data input unit 12000 can perform operations and / or acquisition methods similar to those of the point cloud video acquirer 10001 (or refer to...). Figure 2 The described acquisition process (20000) is the same as or similar to the operation and / or acquisition method.

[0182] The data input unit 12000, quantization processor 12001, voxelization processor 12002, octree occupancy code generator 12003, surface model processor 12004, intra / inter-frame coding processor 12005, and arithmetic encoder 12006 perform geometric coding. Geometric coding according to the embodiment and reference... Figures 1 to 9 The geometric codes described are the same or similar, so a detailed description of them is omitted.

[0183] The quantization processor 12001 according to the embodiment quantizes geometry (e.g., point position values). The operation of the quantization processor 12001 and / or quantization with reference... Figure 4 The operation and / or quantization of the described quantizer 40001 are the same as or similar. Details and references Figures 1 to 9 The details described are the same.

[0184] The voxelization processor 12002 according to the embodiment performs voxelization on the quantized position values ​​of points. The voxelization processor 12002 can perform operations similar to those described above. Figure 4 The operation and / or voxelization process of the described quantizer 40001 are the same as or similar to the operation and / or processing. Details and references Figures 1 to 9 The details described are the same.

[0185] The octree occupancy code generator 12003 according to the embodiment performs octree encoding based on the voxelized positions of points in the octree structure. The octree occupancy code generator 12003 can generate occupancy codes. The octree occupancy code generator 12003 can perform and reference...Figure 4 and Figure 6 The operations and / or methods described are the same as or similar to those of the point cloud video encoder (or octree analyzer 40002). Details and references Figures 1 to 9 The details described are the same.

[0186] According to the implementation, the surface model processor 12004 can perform trigonometric geometry encoding based on a surface model to reconstruct the positions of points in a specific region (or node) based on voxels. The surface model processor 12004 can perform operations related to reference... Figure 4 The operation and / or methods described are the same as or similar to those of the point cloud video encoder (e.g., surface approximation analyzer 40003). Details and references Figures 1 to 9 The details described are the same.

[0187] The intra / inter-frame coding processor 12005 according to the embodiment can perform intra / inter-frame coding on point cloud data. The intra / inter-frame coding processor 12005 can perform operations similar to those described above. Figure 7 The described intra / inter-frame coding is the same or similar. Details and references Figure 7 The details described are the same. According to the implementation, the intra / inter-frame coding processor 12005 may be included in the arithmetic encoder 12006.

[0188] The arithmetic encoder 12006 according to the embodiment performs entropy encoding on an octree and / or an approximate octree of point cloud data. For example, the encoding scheme includes arithmetic encoding. The arithmetic encoder 12006 performs the same or similar operations and / or methods as the arithmetic encoder 40004.

[0189] The metadata processor 12007 according to the embodiment processes metadata (e.g., set values) about point cloud data and provides it to necessary processing procedures such as geometric encoding and / or attribute encoding. Additionally, the metadata processor 12007 according to the embodiment can generate and / or process signaling information related to geometric encoding and / or attribute encoding. The signaling information according to the embodiment can be encoded separately from geometric encoding and / or attribute encoding. The signaling information according to the embodiment can be interleaved.

[0190] Color transformation processor 12008, attribute transformation processor 12009, prediction / boosting / RAHT transformation processor 12010, and arithmetic encoder 12011 perform attribute encoding. Attribute encoding and reference according to the implementation method. Figures 1 to 9 The attribute codes described are the same or similar, so detailed descriptions of them are omitted.

[0191] According to the embodiment, the color transformation processor 12008 performs color transformation encoding to transform color values ​​included in the attributes. The color transformation processor 12008 can perform color transformation encoding based on reconstructed geometry. Reconstructed geometry and reference... Figures 1 to 9 The description is the same. Additionally, it performs the same as the reference. Figure 4 The operation and / or methods of the described color converter 40006 are the same as or similar to those described. Detailed description of it is omitted.

[0192] The attribute transformation processor 12009 according to the implementation performs attribute transformation to transform attributes based on the reconstructed geometry and / or locations where geometric encoding has not been performed. The attribute transformation processor 12009 performs and references... Figure 4 The operation and / or method of the described attribute transformer 40007 are the same as or similar to those operations and / or methods. Detailed descriptions thereof are omitted. The prediction / boosting / RAHT transformation processor 12010 according to the embodiment can encode the transformed attribute by any one or a combination of RAHT encoding, prediction transformation encoding, and boosting transformation encoding. The prediction / boosting / RAHT transformation processor 12010 performs and references... Figure 4 The RAHT transformer 40008, LOD generator 40009, and lift transformer 40010 described herein operate in at least one of the same or similar manner. Furthermore, the predictive transform coding, lift transform coding, and RAHT transform coding are similar to those described in the reference... Figures 1 to 9 The descriptions are the same, so detailed descriptions of them are omitted.

[0193] The arithmetic encoder 12011 according to the embodiment can encode the encoded attributes based on arithmetic encoding. The arithmetic encoder 12011 performs the same or similar operations and / or methods as the arithmetic encoder 40012.

[0194] According to an embodiment, the transmitting processor 12012 can transmit each bitstream containing encoded geometric and / or encoded attribute and metadata information, or transmit a bitstream configured with encoded geometric and / or encoded attribute and metadata information. When the encoded geometric and / or encoded attribute and metadata information according to an embodiment is configured as a bitstream, the bitstream may include one or more sub-bitstreams. The bitstream according to an embodiment may include signaling information, including a sequence parameter set (SPS) for sequence-level signaling, a geometric parameter set (GPS) for signaling for geometric information encoding, an attribute parameter set (APS) for signaling for attribute information encoding, and a tile parameter set (TPS) for tile-level signaling, and tile data. The tile data may include information about one or more tiles. A tile according to an embodiment may include a geometric bitstream Geom0.0 and one or more attribute bitstreams Attr0 0 and Attr1 0 .

[0195] A slice is a series of syntax elements that represent all or part of an encoded point cloud frame.

[0196] The TPS according to the embodiment may include information about each of one or more tiles (e.g., height / size information and coordinate information about the bounding box). The geometric bitstream may include a header and a payload. The header of the geometric bitstream according to the embodiment may include a parameter set identifier (geom_parameter_set_id), a tile identifier (geom_tile_id), and a slice identifier (geom_slice_id) included in the GPS, as well as information about the data contained in the payload. As described above, the metadata processor 12007 according to the embodiment may generate and / or process signaling information and send it to the transmit processor 12012. According to the embodiment, the elements for performing geometry encoding and the elements for performing attribute encoding may share data / information with each other, as indicated by the dashed lines. The transmit processor 12012 according to the embodiment may perform the same or similar operations and / or transmission methods as the transmitter 10003. Details and References Figure 1 and Figure 2 The details described are the same, so the description of it is omitted.

[0197] Figure 13 An example of a receiving device according to an embodiment is shown.

[0198] Figure 13 The receiving device illustrated in the example is Figure 1 The receiving device 10004 (or Figure 10 and Figure 11 Example of a point cloud decoder. Figure 13 The receiving device illustrated in the example can perform the same operation as the reference. Figures 1 to 11 The operations and methods described in the point cloud decoder are one or more of the same or similar operations and methods.

[0199] The receiving apparatus according to the embodiment includes a receiver 13000, a receiving processor 13001, an arithmetic decoder 13002, an octree reconstruction processor based on occupancy code 13003, a surface model processor (triangulation, upsampling, voxelization) 13004, an inverse quantization processor 13005, a metadata parser 13006, an arithmetic decoder 13007, an inverse quantization processor 13008, a prediction / boost / RAHT inverse transform processor 13009, a color inverse transform processor 13010, and / or a renderer 13011. Each element for decoding according to the embodiment can perform the inverse processing of the operation of the corresponding element for encoding according to the embodiment.

[0200] Receiver 13000, according to an embodiment, receives point cloud data. Receiver 13000 can perform operations related to... Figure 1 The operation and / or receiving method of the receiver 10005 are the same as or similar to those of the receiver. Detailed description of it is omitted.

[0201] According to the embodiment, the receiving processor 13001 can acquire geometric bitstreams and / or attribute bitstreams from received data. The receiving processor 13001 may be included in the receiver 13000.

[0202] The arithmetic decoder 13002, the octet-based octree reconstruction processor 13003, the surface model processor 13004, and the dequantization processor 13005 can perform geometric decoding. Geometric decoding and reference according to the implementation method. Figures 1 to 10 The described geometric decodings are the same or similar, so a detailed description of them is omitted.

[0203] The arithmetic decoder 13002 according to the embodiment can decode geometric bitstreams based on arithmetic coding. The arithmetic decoder 13002 performs the same or similar operations and / or coding as the arithmetic decoder 11000.

[0204] According to the embodiment, the octree reconstruction processor 13003 based on occupancy codes can reconstruct an octree by obtaining occupancy codes from the decoded geometric bitstream (or geometric information obtained as a decoding result). The octree reconstruction processor 13003 performs operations and / or methods identical or similar to those of the octree synthesizer 11001 and / or the octree generation method. When applying trigonometric Tang geometry encoding, the surface model processor 13004 according to the embodiment can perform trigonometric Tang geometry decoding and related geometric reconstruction (e.g., triangulation, upsampling, voxelization) based on surface model methods. The surface model processor 13004 performs operations identical or similar to those of the surface approximation synthesizer 11002 and / or the geometry reconstructor 11003.

[0205] The dequantization processor 13005 according to the embodiment can dequantize the decoded geometry.

[0206] The metadata parser 13006 according to the implementation can parse metadata contained in received point cloud data, such as set values. The metadata parser 13006 can transmit metadata for geometry decoding and / or attribute decoding. Metadata and reference Figure 12 The metadata described is the same, so a detailed description of it is omitted.

[0207] The arithmetic decoder 13007, the inverse quantization processor 13008, the prediction / boost / RAHT inverse transform processor 13009, and the color inverse transform processor 13010 perform attribute decoding. Attribute decoding and reference Figures 1 to 10 The properties described are decoded in the same or similar ways, so detailed descriptions of them are omitted.

[0208] The arithmetic decoder 13007 according to the embodiment can decode the attribute bitstream via arithmetic coding. The arithmetic decoder 13007 can decode the attribute bitstream based on the reconstructed geometry. The arithmetic decoder 13007 performs the same or similar operations and / or coding as the arithmetic decoder 11005.

[0209] The dequantization processor 13008 according to the embodiment can dequantize the decoded attribute bitstream. The dequantization processor 13008 performs the same or similar operations and / or methods as the dequantizer 11006 and / or dequantization method.

[0210] The prediction / boosting / RAHT inverse transform processor 13009 according to the embodiment can process the reconstructed geometry and the dequantized attributes. The prediction / boosting / RAHT inverse transform processor 13009 performs one or more operations and / or decodings that are the same as or similar to those of the RAHT transformer 11007, the LOD generator 11008, and / or the inverse booster 11009. The color inverse transform processor 13010 according to the embodiment performs inverse transform encoding to inversely transform the color values ​​(or textures) included in the decoded attributes. The color inverse transform processor 13010 performs operations and / or inverse transform encodings that are the same as or similar to those of the inverse color transformer 11010. The renderer 13011 according to the embodiment can render point cloud data.

[0211] Figure 14 An exemplary structure of a combined point cloud data transmission / reception method / apparatus according to an embodiment is illustrated.

[0212] Figure 14The structure represents a configuration in which at least one of server 1460, robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, home appliance 1450, and / or head-mounted display (HMD) 1470 is connected to cloud network 1400. Robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, or home appliance 1450 are referred to as devices. Additionally, XR device 1430 may correspond to a point cloud data (PCC) device according to an embodiment, or may be operatively connected to a PCC device.

[0213] Cloud network 1400 can refer to a network that forms part of or exists within a cloud computing infrastructure. Here, cloud network 1400 can be configured using a 3G network, a 4G or Long Term Evolution (LTE) network, or a 5G network.

[0214] Server 1460 can be connected via cloud network 1400 to at least one of robot 1410, autonomous vehicle 1420, XR device 1430, smartphone 1440, home appliance 14500 / or HMD 1470, and can assist at least a portion of the processing of connected devices 1410 to 1470.

[0215] HMD 1470 represents one type of implementation of an XR device and / or PCC device according to an embodiment. An HMD-type device according to an embodiment includes a communication unit, a control unit, a memory, an I / O unit, a sensor unit, and a power supply unit.

[0216] Various embodiments of the apparatus 1410 to 1450 applying the above-described technology will be described below. According to the above embodiments, Figure 14 The devices 1410 to 1450 illustrated herein can be operatively connected to / coupled to point cloud data transmitting and receiving devices.

[0217] <PCC+XR>

[0218] The XR / PCC device 1430 can employ PCC technology and / or XR (AR+VR) technology, and can be implemented as an HMD, a head-up display (HUD) installed in a vehicle, a television, a mobile phone, a smartphone, a computer, a wearable device, a home appliance, a digital signage, a vehicle, a stationary robot, or a mobile robot.

[0219] The XR / PCC device 1430 can analyze 3D point cloud data or image data obtained through various sensors or from external devices, and generate position data and attribute data about 3D points. Accordingly, the XR / PCC device 1430 can obtain information about surrounding spaces or real objects, and render and output XR objects. For example, the XR / PCC device 1430 can match an XR object including auxiliary information about the recognized object with the recognized object, and output the matched XR object.

[0220] <PCC+XR+Mobile Phone>

[0221] The XR / PCC device 1430 may be implemented as the mobile phone 1440 by applying PCC technology.

[0222] The mobile phone 1440 can decode and display point cloud content based on PCC technology.

[0223] <PCC + autonomous driving + XR>

[0224] The autonomous driving vehicle 1420 may be implemented as a mobile robot, a vehicle, an unmanned aerial vehicle, etc. by applying PCC technology and XR technology.

[0225] The autonomous driving vehicle 1420 applying XR / PCC technology may refer to an autonomous driving vehicle provided with a device for providing XR images or an autonomous driving vehicle that serves as a control / interaction target in XR images. Specifically, the autonomous driving vehicle 1420 that serves as a control / interaction target in XR images may be distinguished from the XR device 1430 and may be operatively connected to the XR device 1430.

[0226] The autonomous driving vehicle 1420 provided with a device for providing XR / PCC images can acquire sensor information from sensors including a camera, and output a generated XR / PCC image based on the acquired sensor information. For example, the autonomous driving vehicle 1420 may have a HUD and output XR / PCC images thereto, thereby providing an occupant with an XR / PCC object corresponding to a real object or an object existing on a screen.

[0227] When an XR / PCC object is output to the HUD, at least a portion of the XR / PCC object can be output to overlap with a real object that the occupant's eyes are directed towards. On the other hand, when an XR / PCC object is output on a display arranged inside the autonomous driving vehicle, at least a portion of the XR / PCC object can be output to overlap with an object on the screen. For example, the autonomous driving vehicle 1220 can output XR / PCC objects corresponding to objects such as roads, other vehicles, traffic lights, traffic signs, two-wheeled vehicles, pedestrians, and buildings.

[0228] Virtual reality (VR), augmented reality (AR), mixed reality (MR), and / or point cloud compression (PCC) technologies according to the implementation methods are applicable to various devices.

[0229] In other words, VR technology is a technology that only provides CG images of real-world objects, backgrounds, etc. AR technology, on the other hand, refers to the technology of displaying virtually created CG images on top of images of real objects. MR technology is similar to AR technology in that the virtual objects to be displayed are mixed and combined with the real world. However, MR technology differs from AR technology in that AR technology explicitly distinguishes between real objects and virtual objects created as CG images, using virtual objects as supplementary objects to real objects, while MR technology treats virtual objects as objects with the same characteristics as real objects. More specifically, an example of MR technology application is holographic services.

[0230] Recently, VR, AR, and MR technologies have sometimes been referred to as Extended Reality (XR) technologies without being clearly distinguished from each other. Therefore, embodiments of this disclosure are applicable to any of VR, AR, MR, and XR technologies. Encoding / decoding based on PCC, V-PCC, and G-PCC technologies are suitable for such technologies.

[0231] The PCC method / apparatus according to the implementation method can be applied to vehicles that provide autonomous driving services.

[0232] Vehicles providing autonomous driving services connect to the PCC device for wired / wireless communication.

[0233] When the point cloud data (PCC) transmitting / receiving device according to the embodiment is connected to a vehicle for wired / wireless communication, the device can receive / process content data related to AR / VR / PCC services that can be provided along with autonomous driving services and transmit it to the vehicle. When the PCC transmitting / receiving device is installed in the vehicle, it can receive / process content data related to AR / VR / PCC services based on user input signals input through a user interface device and provide it to the user. The vehicle or user interface device according to the embodiment can receive user input signals. User input signals according to the embodiment may include signals indicating autonomous driving services.

[0234] The point cloud data transmission method / apparatus according to the embodiments can be interpreted as referring to... Figure 1 The transmitting device 10000 in Figure 1 The point cloud video encoder 10002 in Figure 1 Transmitter 10003 in Figure 2 The code snippet shows the steps for obtaining 20000 / encoding 20001 / sending 20002. Figure 4encoder in Figure 12 The transmitting device in Figure 14 The device in Figure 18 encoder, Figure 30 Terms related to sending methods, etc.

[0235] The point cloud data receiving method / apparatus according to the embodiments can be interpreted as referring to... Figure 1 The receiving device 10004 Figure 1 Receiver 10005 Figure 1 Point cloud video decoder 10006 Figure 2 Sending 20002 / Decoding 20003 / Rendering 20004 Figure 10 and Figure 11 decoder Figure 13 The receiving device Figure 14 The device Figure 19 decoder Figure 31 Terminology related to receiving methods, etc.

[0236] The method / apparatus for sending or receiving point cloud data according to the embodiments can be simply referred to as a method / apparatus.

[0237] According to the implementation method, the geometric data, geometric information, and location information constituting point cloud data will be interpreted as having the same meaning. Similarly, the attribute data, attribute information, and attribute information constituting point cloud data will be interpreted as having the same meaning.

[0238] Considering scalable transmission, the method / apparatus according to the implementation can process point cloud data based on the point cloud data structure according to the implementation.

[0239] Regarding the method / apparatus according to the embodiments, this document describes a method for efficiently supporting selective decoding of a portion of data when selective decoding is required due to receiver performance or transmission speed during the transmission and reception of point cloud data. The proposed method proposes a way to select desired information or eliminate unnecessary information in bitstream units by dividing geometric and attribute data transmitted in data units into semantic units such as geometric octrees and levels of detail (LoD).

[0240] Techniques for configuring data structures composed of point clouds, according to embodiments, are described. Specifically, a method disclosed in the embodiments for packaging and processing relevant signaling information to efficiently transmit layer-configured PCC data will be described, and a method for applying it to scalable PCC-based services is proposed.

[0241] Reference Figure 4 and Figure 11The point cloud data transmitting / receiving apparatus (or simply encoder / decoder) shown in the embodiment comprises point cloud data consisting of the position (e.g., XYZ coordinates) and attributes (e.g., color, reflectivity, intensity, gray level, opacity, etc.) of each data point. In point cloud compression (PCC), octree-based compression is performed to efficiently compress the non-uniform distribution characteristics in three-dimensional space, and attribute information is compressed accordingly. Figure 4 neutralization Figure 11 The PCC transmitting / receiving apparatus illustrated herein can handle operations according to the implementation method through each component device.

[0242] Figure 15 The process of encoding, transmitting, and decoding point cloud data according to the implementation method is illustrated.

[0243] The point cloud encoder 15000 is a transmitting device that implements the transmitting method according to the embodiment, and can scalably encode and transmit point cloud data.

[0244] The point cloud decoder 15010 is a receiving device that implements the receiving method according to the embodiment, and can scalably decode point cloud data.

[0245] The source data received by encoder 15000 may include geometric data and / or attribute data.

[0246] The encoder 15000 scalably encodes point cloud data but does not immediately generate a partial PCC bitstream. Instead, when it receives complete geometry and attribute data, it stores the data in memory connected to the encoder. The encoder can then perform transcoding for partial encoding and generate and send a partial PCC bitstream. The decoder 15010 can receive and decode the partial PCC bitstream to reconstruct partial geometry and / or partial attributes.

[0247] Upon receiving complete geometry and complete attributes, encoder 15000 can store the data in a memory connected to the encoder and transcode the point cloud data using a low quantization parameter (QP) to generate and transmit a complete PCC bitstream. Decoder 15010 can receive and decode the complete PCC bitstream to reconstruct the complete geometry and / or complete attributes. Decoder 15010 can select partial geometry and / or partial attributes from the complete PCC bitstream via data selection.

[0248] The method / apparatus according to the embodiment compresses and transmits point cloud data by dividing the location information of data points and feature information such as color / brightness / reflectivity into geometric information and attribute information. In this case, an octree structure with layers can be configured according to the level of detail, or PCC data can be configured according to the level of detail (LOD). Scalable point cloud data encoding and representation can then be performed based on the configured structure or data. In this case, due to the performance of the receiver or the transmission rate, only a portion of the point cloud data can be decoded or represented.

[0249] In this process, the method / apparatus according to the implementation can remove unnecessary data in advance. In other words, when only a portion of the scalable PCC bitstream needs to be transmitted (i.e., only some layers are decoded in scalable decoding), there is no way to select and transmit only the necessary portion. Therefore, 1) the necessary portion needs to be re-encoded after decoding (15020), or 2) the receiver must selectively apply the operation after the entire data has been transmitted (15030). However, in case 1), delays may occur due to the time spent on decoding and re-encoding (15020). In case 2), bandwidth efficiency may decrease due to the transmission of unnecessary data. Additionally, when using fixed bandwidth, it may be necessary to reduce data quality during transmission (15030).

[0250] Therefore, the method / apparatus according to the implementation can define a sliced ​​structure for point cloud data and signal the scalable layer and slice structure for scalable transmission.

[0251] In implementation, to ensure efficient bitstream transmission and decoding, the bitstream can be divided into specific units for processing.

[0252] According to the implementation, the unit may be referred to as a LOD, layer, slice, etc. LOD is the same term as LOD in attribute data encoding, but can refer to a data unit used for a hierarchical structure of bitstreams. This can be a concept corresponding to a bundle of one or two or more depths of a hierarchical structure based on point cloud data (e.g., an octree or the depth (hierarchy) of multiple trees). Similarly, a layer is a unit set up to generate a sub-bitstream and is a concept corresponding to a bundle of one or two or more depths, and can correspond to one LOD or two or more LODs. Additionally, a slice is a unit used to constitute a sub-bitstream unit and can correspond to one depth, a portion of one depth, or two or more depths. Furthermore, it can correspond to one LOD, a portion of one LOD, or two or more LODs. According to the implementation, LODs, layers, and slices can correspond to each other, or one of LODs, layers, and slices can be included in another. Furthermore, the unit according to the implementation may include LODs, layers, slices, layer groups, or subgroups, and can be referred to as complementary to each other.

[0253] Figure 16 An example of a layer-based point cloud data configuration according to an implementation method is shown.

[0254] The transmission method / apparatus according to the implementation method can be configured as follows: Figure 16 The layer-based point cloud data shown is used for encoding and decoding point cloud data.

[0255] Depending on the application domain, point cloud data can be layered with structures such as SNR, spatial resolution, color, temporal frequency, and bit depth. Layers can also be constructed in the direction of increasing data density based on octree or LoD structures.

[0256] Figure 17 Geometric bitstream structures and attribute bitstream structures according to the implementation method are illustrated.

[0257] The method / apparatus according to the implementation can be based on Figure 16 The layering shown in the diagram is used for configuration, encoding, and decoding. Figure 17 The geometric bitstream and attribute bitstream are shown in the figure.

[0258] According to the embodiments, the bitstream acquired by the transmitting device / encoder through point cloud compression can be divided into geometric data bitstream and attribute data bitstream according to the data type and then transmitted.

[0259] Each bitstream according to the implementation method can be composed of slices. Regardless of the layer information or LoD information, the geometric data bitstream and the attribute data bitstream can each be configured as a slice and transmitted. In this case, when only a portion of the layer or LoD will be used, the following should be performed: 1) decode the bitstream, 2) select only the desired portion and remove unnecessary portions, and 3) re-encode only based on the necessary information.

[0260] Figure 18 A bitstream configuration according to an implementation method is illustrated.

[0261] The transmission method / apparatus according to the implementation method can generate, for example... Figure 18 The bitstream shown, and the receiving method / apparatus according to the embodiment can decode it as follows: Figure 18 The bitstream shown contains point cloud data.

[0262] Bitstream configuration according to the implementation method

[0263] In implementations, to avoid unnecessary intermediate processes, the bitstream can be divided into layers (or LODs) and sent.

[0264] For example, in the case of LoD-based PCC technology, lower LoDs are included in higher LoDs. Information included in the current LoD but not in previous LoDs—that is, information newly added to each LoD—can be referred to as R (Rest). Figure 18 As shown, the initial LoD information and the information R newly added to each LoD can be divided into independent units and sent.

[0265] The transmission method / apparatus according to the implementation can encode geometric data and generate a geometric bitstream. The geometric bitstream can be configured for each LOD or layer. The geometric bitstream can include a header (geometric header) for each LOD or layer. The header can include reference information for the next LOD or layer. The current LOD (layer) can also include information R (geometric data) that was not included in previous LODs (layers).

[0266] The receiving method / apparatus according to the embodiments can encode attribute data and generate an attribute bitstream. The attribute bitstream can be configured for each LOD or layer, and the attribute bitstream can include a header (attribute header) for each LOD or layer. The header can include reference information for the next LOD or layer. The current LOD (layer) can also include information R (attribute data) that was not included in previous LODs (layers).

[0267] The receiving method / apparatus according to the embodiments can receive bit streams consisting of LODs or layers, and efficiently decode only the necessary data without complex intermediate processes.

[0268] Figure 19 An example of a bitstream arrangement method according to an implementation method is shown.

[0269] The method / apparatus according to the embodiments can be as follows: Figure 19 The layout shown Figure 18 The bitstream.

[0270] Bit stream arrangement method according to the implementation method.

[0271] The transmission method / apparatus according to the implementation method can transmit bit streams as follows: Figure 19 The diagram shows the serial transmission of geometry and attributes. In this case, the entire geometry information (geometric data) can be transmitted first, depending on the data type, followed by the attribute information (attribute data). In this scenario, the geometry information can be quickly reconstructed based on the transmitted bitstream information.

[0272] exist Figure 19 For example, the layer LOD containing geometric data can be set first in the bitstream, and the layer LOD containing attribute data can be set after the geometric layer. Since the attribute data depends on the geometric data, the geometric layer can be set first. Furthermore, the position can be changed depending on the implementation. References can be made between geometric headers and between attribute headers and geometric headers.

[0273] Figure 20 An example of a bitstream arrangement method according to an implementation method is shown.

[0274] like Figure 19 Same, Figure 20 This is an example of a bitstream arrangement according to an implementation method.

[0275] Bitstreams comprising both geometric and attribute data in the same layer can be grouped and transmitted. In this case, decoding execution time can be reduced by using compression techniques capable of parallel decoding of geometry and attributes. Information that needs to be processed first (where geometry should precede attributes by a small LoD) can be placed first.

[0276] The first layer 2000 includes geometric and attribute data corresponding to the minimum LOD 0 (layer 0) along with each header, and the second layer 2010 includes LOD 0 (layer 0) and includes geometric and attribute data as R1 information for points in a newer, more detailed layer 1 (LOD 1) that are not in LOD 0 (layer 0). Similarly, a third layer 2020 may follow.

[0277] The transmitting / receiving method / apparatus according to the embodiments can efficiently select the desired layer (or LoD) in the application domain at the bitstream level when transmitting and receiving bitstreams. In the embodiments ( Figure 19 When grouping and transmitting bitstreams according to the geometry of the bitstream arrangement method, there may be empty sections in the middle after selecting the bitstream level. In this case, it may be necessary to rearrange the bitstream. When geometry and attributes are grouped and transmitted according to the layer ( Figure 20 Unnecessary information can be selectively removed based on the application area, as follows.

[0278] Figure 21 A method for selecting geometric data and attribute data according to an implementation method is illustrated.

[0279] Bitstream selection according to the implementation method

[0280] When it is necessary to select a bitstream as described above, the method / apparatus according to the implementation can be as follows: Figure 21 The data selection at the bitstream level is shown as follows: 1) symmetric selection of geometry and attributes; 2) asymmetric selection of geometry and attributes; or 3) a combination of the above two methods.

[0281] 1) Symmetrical selection of geometry and attributes

[0282] Figure 21 The example illustrates the case where only LoD1 (LOD 0+R1) 21000 is selected and sent or decoded. In this case, when sending and decoding are performed, the information corresponding to R2 (the new part of LOD 2) 21010 corresponding to the higher layer is removed.

[0283] Figure 22 An example of a bitstream selection method according to an implementation method is given.

[0284] 2) Asymmetric selection of geometry and attributes

[0285] According to the method / apparatus of the implementation, geometry and attributes can be transmitted asymmetrically. Only the attributes of the upper level (attribute R2 22000) can be removed, and all geometry (level 0 (root level) to level 7 (leaf level) in the triangle octree structure) can be selected and transmitted / decoded (22010).

[0286] Reference Figure 16 When point cloud data is represented in an octree structure and divided into LODs (or layers) according to levels, scalable encoding / decoding (scalability) can be supported.

[0287] The scalability features, depending on the implementation, may include slice-level scalability and / or octree-level scalability.

[0288] According to the implementation, the LoD (Level of Detail) can be used as a unit to represent a collection of one or more octree layers. Alternatively, it can mean a bundle of octree layers that will be configured as slices.

[0289] In attribute encoding / decoding, the LOD according to the implementation can be extended in a broader sense and used as a unit for detailed data partitioning.

[0290] That is, spatial scalability can be provided for each octree layer through actual octree layers (or scalable attribute layers). However, when configuring scalability in slices before bitstream parsing, the choice can be made in LoD depending on the implementation.

[0291] In an octree structure, LOD0 can correspond to the root level to level 4, LOD1 can correspond to the root level to level 5, and LOD2 can correspond to the root level to level 7, i.e., the leaf level.

[0292] That is, such as Figure 16 As shown, when scalability is utilized in a slice, as in the case of scalable transmission, the provided scalable steps can correspond to the three steps LoD0, LoD1, and LoD2, and the scalable steps provided by the octree structure in the decoding operation can correspond to the eight steps from root to leaf.

[0293] According to the implementation method, for example, in Figure 16 In the process, when LoD0 to LoD2 are configured as the corresponding slices, the transcoder of the receiver or transmitter ( Figure 15 The transcoder 15040 can be selected as 1) LoD0 only, 2) LoD0 and LoD1, or 3) LoD0, LoD1 and LoD2 for scalable processing.

[0294] Example 1: When only LoD0 is selected, the maximum octree level can be 4, and a scalable layer can be selected from octree levels 0 to 4 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0295] Example 2: When LoD0 and LoD1 are selected, layer 5 can be added. Therefore, the maximum octree level can be 5, and a scalable layer can be selected from octree layers 0 to 5 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0296] According to the implementation method, the octree depth, octree level, and octree grade can be units in which data is divided in detail.

[0297] Example 3: When LoD0, LoD1, and LoD2 are selected, layers 6 and 7 can be added. Therefore, the maximum octree level can be 7, and a scalable layer can be selected from octree layers 0 to 7 during decoding. In this case, the receiver can treat the node size obtainable through the maximum octree depth as a leaf node and can send the node size via signaling information.

[0298] Figure 23 An example is provided of a method for configuring slices containing point cloud data according to an implementation method.

[0299] Slicing configuration according to the implementation method

[0300] According to the implementation method / apparatus / encoder, the G-PCC bitstream can be configured by segmenting the bitstream according to a slice structure. The data unit used for detailed data representation can be a slice.

[0301] For example, one or more octree layers can be matched with a slice.

[0302] According to the implementation method / apparatus (e.g., encoder), the bit stream can be configured based on slice 2301 by scanning the nodes (points) contained in the octree in the direction of scan sequence 2300.

[0303] Figure 23 (a): A slice can contain some nodes of an octree layer.

[0304] An octree layer, for example, level 0 to level 4, can form a slice 2002.

[0305] For example, data from level 5 of an octree can form slices 2003, 2004, and 2005.

[0306] In an octree, for example, data from level 6 can form each slice.

[0307] Figure 23 (b) and Figure 23 (c) When multiple octree layers match a slice, only a portion of the nodes from each layer may be included. In this way, when multiple slices constitute a geometry / attribute frame, the information required to configure the layers can be transmitted to the receiver. This information may include information about the layers contained in each slice and information about the nodes contained in each layer.

[0308] Figure 23 (b): Some data from octree levels, such as levels 0 to 3 and 4, can be configured in a slice.

[0309] For example, partial data at level 4 and partial data at level 5 of an octree can be configured as a slice.

[0310] An octree layer, for example, can have parts of data at level 5 and parts of data at level 6 configured as a slice.

[0311] For example, a portion of the data in an octree level 6 can be configured as a slice.

[0312] Figure 23 (c): Data in an octree layer, such as level 0 to level 4, can be configured in a slice.

[0313] Some data from each of the octree levels 5, 6, and 7 can be configured in a slice.

[0314] The encoder and the corresponding device according to the embodiment can encode point cloud data and generate and transmit a bit stream containing encoded data and parameter information related to the point cloud data.

[0315] Furthermore, when generating the bitstream, the implementation can be based on the bitstream structure (e.g., see [link to implementation details]). Figures 17 to 23 (etc.) to generate a bitstream. Accordingly, the receiving device, decoder, and corresponding device according to the embodiment can receive and parse the bitstream configured to suit a selective partial data decoding structure, and partially decode the point cloud data to efficiently provide data (see...). Figure 15 ).

[0316] Scalable transmission according to the implementation method

[0317] The point cloud data transmission method / apparatus according to the embodiments can transmit bit streams containing point cloud data in a scalable manner, and the point cloud data reception method / apparatus according to the embodiments can receive and decode bit streams in a scalable manner.

[0318] When Figures 17 to 23 When the bitstream illustrated in the embodiment is used for scalable transmission, information for selecting the slice required by the receiver can be sent to the receiver. Scalable transmission may mean transmitting or decoding only a portion of the bitstream instead of decoding the entire bitstream, which can result in low-resolution point cloud data.

[0319] When sending point cloud data to scalable applications for octree-based geometry bitstreams, the data may need to be configured for each octree layer from the root node to the leaf node. Figure 16 The range of the bitstream is limited to the information of a specific octree layer.

[0320] For this purpose, the target octree layer should not depend on information about lower octree layers. This can be a constraint applied jointly to geometric encoding and attribute encoding.

[0321] Additionally, in scalable transmission, a scalable structure for selecting scalable layers by the transmitter / receiver must be transmitted. Considering the octree structure according to the implementation, all octree layers can support scalable transmission. Alternatively, scalable transmission can be allowed only for specific octree layers and subsequent layers. When some octree layers are included, it can be indicated that scalable layers including slices are included. Thus, it can be determined whether slices are necessary / unnecessary at the bitstream level. Figure 23 In example (a), a scalable layer can be configured in the yellow-marked section starting from the root node without supporting scalable transmission, and subsequent octree layers can be configured to match the scalable layer in a one-to-one correspondence. Typically, scalability can be supported for the portions corresponding to leaf nodes. When such... Figure 23 As shown in (c), when a slice includes multiple octree layers, a scalable layer can be defined as being configured for the layer.

[0322] In this context, scalable transmission and scalable decoding can be used separately depending on the purpose. Scalable transmission can be used on the transmit / receive side for the purpose of selecting information up to a specific layer without involving the decoder. Scalable decoding is used to select a specific layer during encoding. That is, scalable transmission can support the selection of necessary information in a compressed state (at the bitstream level) without involving the decoder, so that this information can be determined or transmitted by the receiver. On the other hand, scalable decoding can support encoding / decoding only the required portion of data during the encoding / decoding process, and therefore can be used as a scalable representation in this case.

[0323] In this context, the layer configuration for scalable transmission can differ from the layer configuration for scalable decoding. For example, in scalable transmission, three low-level octree layers including leaf nodes can constitute a single layer. However, when all layer information is included in scalable decoding, scalable decoding can be performed for each of leaf node layer n, leaf node layer n-1, and leaf node layer n-2.

[0324] The following sections will describe the slice structure used for the above layer configuration and the signaling method for scalable transmission.

[0325] Figure 24 A bitstream configuration according to an implementation method is illustrated.

[0326] The method / apparatus according to the implementation method can generate, as follows: Figure 24 The bitstream shown is a bitstream that can contain encoded geometric and attribute data, as well as parameter information.

[0327] The following describes the syntax and semantics of parameter information.

[0328] According to the implementation method, information about the separated slices can be defined in the parameter set of the bitstream and the SEI message as follows.

[0329] A bitstream can include a Sequence Parameter Set (SPS), a Geometry Parameter Set (GPS), an Attribute Parameter Set (APS), a geometry slice header, and an attribute slice header. In this regard, depending on the application or system, the scope and method to be applied can be defined in corresponding or individual locations and used differently. That is, a signal can have different meanings depending on where it is transmitted. If a signal is defined in an SPS, it can be applied equivalently to the entire sequence. If a signal is defined in a GPS, this can indicate that the signal is used for location reconstruction. If a signal is defined in an APS, this can indicate that the signal is applied to attribute reconstruction. If a signal is defined in a TPS, this can indicate that the signal is applied only to points within a tile. If a signal is transmitted in a slice, this can indicate that the signal is applied only to the slice. Furthermore, the scope and method to be applied can be defined in corresponding or individual locations depending on the application or system and used differently. Additionally, when the syntax elements defined below apply to multiple point cloud data streams besides the current point cloud data stream, they can be carried in the parent parameter set.

[0330] The abbreviations used in this article are: SPS: Sequence Parameter Set; GPS: Geometric Parameter Set; APS: Attribute Parameter Set; TPS: Tile Parameter Set; Geom: Geometric Bitstream = Geometric Slice Header + Geometric Slice Data; Attr: Attribute Bitstream = Attribute Slice Header + Attribute Slice Data.

[0331] While the implementation method defines information independently of the encoding technique, information can be defined in conjunction with the encoding technique. To support different scalability across regions, information can be defined in the tile parameter set of the bitstream. Furthermore, when the syntax elements defined below apply not only to the current point cloud data stream but also to multiple point cloud data streams, they can be carried in the parent parameter set, etc.

[0332] Alternatively, network abstraction layer (NAL) units can be defined for the bitstream, and information such as layer_id for selecting layers can be transmitted. This allows for bitstream selection at the system level.

[0333] In the following text, the parameters according to the embodiment (which may be referred to as metadata, signaling information, etc.) can be generated in the processing of the transmitter according to the embodiment and sent to the receiver according to the embodiment for use in the reconstruction process.

[0334] For example, parameters can be generated by the metadata processor (or metadata generator) of the transmitting device according to the embodiment, which will be described later, and can be obtained by the metadata parser of the receiving device according to the embodiment.

[0335] The following text will refer to Figures 25 to 28 Describe the syntax / semantics of the parameters contained in the bitstream.

[0336] Figure 25 The syntax of the sequence parameter set and geometric parameter set according to the implementation is shown.

[0337] Figure 26 The syntax of the attribute parameter set according to the implementation method is shown.

[0338] Figure 27 The syntax of the geometric data cell header according to the implementation is shown.

[0339] Figure 28 The syntax of the attribute data cell header according to the implementation method is shown.

[0340] Below, description Figures 25 to 28 The semantics of the parameters included in the implementation method.

[0341] `scalable_transmission_enable_flag`: When equal to 1, it indicates that the bitstream is configured for scalable transmission. That is, since the bitstream consists of multiple slices, selection information can be provided at the bitstream level. Scalable layer configuration information can be sent to indicate that slice selection is available in the transmitter or receiver, and that geometry and / or attributes are compressed to enable partial decoding. When `scalable_transmission_enable_flag` is 1, the transcoder of the receiver or transmitter can be used to determine whether scalable transmission of geometry and / or attributes is allowed. The transcoder can be coupled to or included in the transmitting and receiving devices.

[0342] geom_scalable_transmission_enable_flag and attr_scalable_transmission_enable_flag: When equal to 1, they can indicate that geometry or attributes are compressed to enable scalable transmission.

[0343] For example, for geometry, this flag could indicate that the geometry consists of octree-based layers, or take into account scalable transmission of sliced ​​segments already performed (see...). Figure 23 ).

[0344] When geom_scalable_transmission_enable_flag or attr_scalable_transmission_enable_flag is 1, the receiver can know that scalable transmission is available for geometry or attributes.

[0345] For example, a geom_scalable_transmission_enable_flag of 1 can indicate the use of octree-based geometry encoding and disable QTBT, or encode the geometry in a form such as octree by performing encoding in the order of BT-QT-OT.

[0346] A value of 1 for attr_scalable_transmission_enable_flag can indicate whether to use pre-lifting encoding or scalable RAHT (e.g., Haar-based RAHT) by using scalable LOD generation.

[0347] `num_scalable_layers` can indicate the number of layers that support scalable transmission. Depending on the implementation, a layer can refer to a Level of Dimension (LOD).

[0348] The `scalable_layer_id` specifies the indicator of the layer that constitutes the scalable transmission. When the scalable layer consists of multiple slices, common information can be carried in the parameter set through `scalable_layer_id`, and different information can be carried in the data cell header according to the slice.

[0349] `num_octree_layers_in_scalable_layer` can indicate the number of octree layers included in or corresponding to the layers that make up the scalable layer. It can refer to the corresponding layer when no scalable layer is configured based on an octree.

[0350] tree_depth_start can indicate the starting octree depth (relative to the root) included in the layers that make up the scalable transmission or in the octree layers corresponding to that layer.

[0351] tree_depth_end can indicate the depth of the last octree (relative to the closest leaf) included in the layers that make up the scalable transmission or in the octree layers corresponding to that layer.

[0352] `node_size` can indicate the size of the nodes in the output point cloud data when reconstructing a scalable layer via scalable transmission. For example, `node_size` equal to 1 can indicate leaf nodes. Although the implementation assumes that the XYZ node size is constant, arbitrary node sizes can be indicated by signaling the size in each direction of the transformed coordinate system, such as (r(radius), phi, theta), or in the XYZ direction.

[0353] num_nodes can indicate the number of nodes included in the corresponding scalable layer.

[0354] `num_slices_in_scalable_layer` indicates the number of slices that belong to a scalable layer.

[0355] The slice_id specifies an indicator used to distinguish slices or data units, and can transmit the indicator of data units belonging to the scalable layer.

[0356] `aligned_slice_structure_enabled_flag`: When equal to 1, it indicates that the attribute-scalable layer structure and / or slice configuration matches the geometrically scalable layer structure and / or slice configuration. In this case, information about the attribute-scalable layer structure and / or slice configuration can be identified using information about the geometrically scalable layer structure and / or slice configuration. That is, the geometric layer / slice structure is the same as the attribute layer / slice structure.

[0357] `slice_id_offset` can indicate the offset used to obtain attribute slices or data cells based on geometry slice IDs. According to the implementation, when `aligned_slice_structure_enabled_flag` is 1, that is, when the attribute slice structure matches the geometry slice structure, the attribute slice ID can be obtained based on the geometry slice ID as follows.

[0358] Slice_id(attr)=slice_id(geom)slice_id_offset

[0359] In this case, the values ​​provided in the geometry parameter set can be used as variables for configuring the attribute slice structure: num_scalable_layers, scalable_layer_id, tree_depth_start, tree_depth_end, node_size, num_nodes, and num_slices_in_scalable_layer.

[0360] The corresponding_geom_scalable_layer can indicate the geometrically scalable layer corresponding to the attribute scalable layer structure.

[0361] num_tree_depth_in_data_unit can indicate the tree depth including nodes that belong to the data unit.

[0362] tree_depth can indicate the depth of the tree.

[0363] num_nodes can indicate the number of nodes belonging to tree_depth among the nodes belonging to the data unit.

[0364] The aligned_geom_data_unit_id can indicate the geometry data unit ID when the attribute data unit conforms to the scalable send layer / slice structure of the corresponding geometry data unit.

[0365] ref_slice_id can be used to refer to the slice that should precede the current slice used for decoding (see example...). Figures 18 to 20 ).

[0366] Figure 29 The structure of a point cloud data transmission apparatus according to an embodiment is illustrated.

[0367] Figure 29 The transmitting device according to the embodiment corresponds to Figure 1 The transmitting device 10000 in Figure 1 The point cloud video encoder 10002, Figure 1 Transmitter 10003 in Figure 2 The code snippet shows the steps for obtaining 20000 / encoding 20001 / sending 20002. Figure 4 encoder in Figure 12 The transmitting device in Figure 14 The device in Figure 18 encoder in Figure 30 Methods for sending data, etc. Figure 29 Each component can correspond to hardware, software, processor, and / or a combination thereof.

[0368] Operation of the encoder or transmitter according to the implementation method:

[0369] When point cloud data is input to the transmitting device, the encoder can encode position information (geometric data (e.g., XYZ coordinates, phi-theta coordinates, etc.)) and attribute information (attribute data (e.g., color, reflectivity, intensity, gray level, opacity, medium, material, gloss, etc.)) (geometric encoding and attribute encoding).

[0370] The compressed (encoded) data is divided into units for transmission. The data can be divided by a sub-bitstream generator into units suitable for selecting necessary information from the bitstream units based on hierarchical structure information, and then it can be packaged.

[0371] The hierarchical structure information according to the implementation method is an indication Figures 16 to 23 Information on bitstream configuration, arrangement, and selection, as well as slice configuration, and representation. Figures 24 to 28 The information shown is hierarchical structure information, which can be generated by the metadata generator. The sub-bitstream generator can divide the bitstream, generate hierarchical structure information indicating the division process, and send this information to the metadata generator. The metadata generator can receive information indicating geometric encoding and attribute encoding from the encoder and generate metadata (parameters).

[0372] The transmitting device according to the implementation can be reused and transmit the parameters and sub-bit streams of each layer.

[0373] Figure 30 The structure of a point cloud data receiving device according to an embodiment is illustrated.

[0374] Figure 30 The receiving device according to the embodiment corresponds to Figure 1 The receiving device 10004 Figure 1 Receiver 10005 Figure 1 Point cloud video decoder 10006 Figure 2 Sending 20002 / Decoding 20003 / Rendering 20004 Figure 10 and Figure 11 decoder Figure 13 The receiving device Figure 14 The device Figure 19 decoder Figure 31 The receiving method, etc. Figure 30 Each component can correspond to hardware, software, processor, and / or a combination thereof.

[0375] The operation of the decoder / receiver according to the implementation method:

[0376] When a bitstream is input to a receiving device, the receiver can process the bitstream separately for location information and attribute information (demultiplexing). In this case, a sub-bitstream classifier can send the sub-bitstream to the appropriate decoder based on the information in the bitstream header. Alternatively, the layer required by the receiver can be selected during this process. Geometric data and attribute data can be reconstructed from the classified bitstream by the geometric decoder and attribute decoder respectively according to the characteristics of the data, and then converted into the format for the final output of the renderer.

[0377] The sub-bitstream classifier can classify / select bitstreams based on metadata obtained by the metadata parser.

[0378] The geometry decoder and attribute decoder can decode geometric data and attribute data respectively based on the metadata obtained by the metadata parser.

[0379] Figure 30 The operation of each component of the receiving device can follow Figure 29 The operation of the corresponding component of the transmitting device or its reverse process.

[0380] Figure 31 This is a flowchart illustrating a point cloud data receiving device according to an embodiment.

[0381] Figure 31 More detailed examples Figure 30 The operation of the sub-bitstream classifier is shown in the figure.

[0382] The receiving device receives data slice by slice, and the metadata parser transmits parameter set information such as SPS, GPS, APS, and TPS. Based on the transmitted information, scalability can be determined. When the data is scalable, such as... Figure 31 The diagram illustrates the identification of the slice structure used for scalable transmission. First, the geometric slice structure can be identified based on information such as num_scalable_layers, scalable_layer_id, tree_depth_start, tree_depth_end, node_size, num_nodes, num_slices_in_scalable_layer, and slice_id carried in GPS.

[0383] When aligned_slice_structure_enabled_flag equals 1, the attribute slice structure can also be identified in the same way (e.g., geometry encoded based on an octree, attributes encoded based on scalable LoD or scalable RAHT, and geometry / attribute slices generated by the same slice segmentation have the same number of nodes for the same octree layer).

[0384] When the structures are the same, the range of geometry slice IDs is determined based on the target scalable layer, and the range of attribute slice IDs is determined by slice_id_offset. Geometry / attribute slices are then selected based on the determined ranges.

[0385] When `aligned_slice_structure_enabled_flag` = 0, the attribute slice structure can be identified based on information such as `num_scalable_layers`, `scalable_layer_id`, `tree_depth_start`, `tree_depth_end`, `node_size`, `num_nodes`, `num_slices_in_scalable_layer`, and `slice_id` transmitted via APS, and the range of necessary attribute slice IDs can be limited according to scalable operations. Based on this range, the desired slices can be selected by each slice ID before reconstruction. The geometry / attribute slices selected through the above process are sent as input to the decoder.

[0386] The decoding process based on the slice structure has been described above based on the receiver's scalable transmission or scalable selection. However, when `scalable_transmission_enabled_flag` equals 0, the operation of determining the range of geom / attr slice IDs can be skipped, and all slices can be selected, making them usable even in non-scalable operations. Even in this case, information about previous slices (e.g., slices belonging to higher layers or slices specified by `ref_slice_id`) can be used through slice structure information carried in parameter sets such as SPS, GPS, APS, or TPS.

[0387] Bitstreams can be received based on scalable transmission, and the scalable bitstream structure can be identified based on the parameter information included in the bitstream.

[0388] It is possible to estimate geometrically scalable layers.

[0389] Geometric slices can be identified based on geom_slice_id.

[0390] You can select a geometric slice based on slice_id.

[0391] The decoder can decode the selected geometric slices.

[0392] When the aligned_slice_structure_enabled_flag included in the bitstream is equal to 1, the attribute slice ID corresponding to the geometric slice can be checked. The attribute slice can be accessed based on slice_id_offset.

[0393] You can select an attribute slice based on slice_id.

[0394] The decoder can decode the selected attribute slice.

[0395] When `aligned_slice_structure_enabled_flag` is not equal to 1, the scalable layer of an attribute can be estimated. Attribute slices can be identified based on their attribute slice IDs.

[0396] You can select an attribute slice based on slice_id.

[0397] The transmitting device according to the embodiment has the following effects.

[0398] For point cloud data, the transmitting device can divide and transmit compressed data according to specific standards. When using layered encoding according to the implementation method, compressed data can be divided and transmitted according to layers. Therefore, the storage and transmission efficiency on the transmitting side can be improved.

[0399] Reference Figure 15 It can compress and provide the geometry and attributes of point cloud data. In PCC-based services, the amount of data or the compression rate can be adjusted based on receiver performance or transmission environment.

[0400] When point cloud data is configured in a slice, as receiver performance or transmission environment changes, 1) a bitstream suitable for each environment can be transcoded and stored separately and can be selected at transmission time, or 2) transcoding may be required before transmission. In this case, if the number of receiver environments to be supported increases or the transmission environment changes frequently, it may cause storage space-related problems or latency due to transcoding.

[0401] Figure 32 The transmission / reception of point cloud data according to the implementation method is illustrated.

[0402] To address the above problems, the method / apparatus according to the implementation method can be as follows: Figure 32 The example demonstrates the processing of point cloud data.

[0403] When compressed data is partitioned and transmitted according to layers according to the implementation method, only the necessary portions of the pre-compressed data can be selectively transmitted at the bitstream level without a separate transcoding process. Since each stream requires only one storage space, storage space can be operated efficiently. Furthermore, efficient transmission can be achieved in terms of bandwidth (bitstream selector) because only the necessary layers are selected before transmission.

[0404] The receiving method / apparatus according to the embodiments can provide the following effects.

[0405] According to the implementation method, compressed data can be partitioned and transmitted based on a single standard used for point cloud data. When using hierarchical encoding, compressed data can be partitioned and transmitted according to layers. In this case, the efficiency of the receiving side can be improved.

[0406] Figure 15 This illustrates the operations on both the transmitting and receiving sides when transmitting point cloud data consisting of layers. In this case, when information for reconstructing the entire PCC data is transmitted regardless of receiver performance, the receiver needs to reconstruct the point cloud data through decoding and then select only the data corresponding to the desired layer (data selection or subsampling). In this situation, since the transmitted bitstream has already been decoded, delays may occur in receivers aiming for low latency, or decoding may fail depending on receiver performance.

[0407] Therefore, when a bitstream is divided into slices and transmitted, the receiver can selectively send the bitstream to the decoder based on the density of the point cloud data to be represented, depending on the decoder's performance or the application domain. In this case, since the selection is performed before decoding, decoder efficiency can be improved, and decoders with various performance levels can be supported.

[0408] The method / apparatus according to the implementation can use layer groups and subgroups to transmit bit streams, and further perform slicing and segmentation.

[0409] Figure 33 Examples of a single-slice-based geometric tree structure and a segmented-slice-based geometric tree structure according to embodiments are shown.

[0410] like Figure 33 As illustrated, the method / apparatus according to the embodiment can be configured to transmit point cloud data slices.

[0411] Figure 33 The diagram illustrates the geometric tree structure contained within different slice structures. According to G-PCC technology, the entire encoded bitstream can be contained within a single slice. For multiple slices, each slice can contain sub-bitstreams. The order of the slices can be the same as the order of the sub-bitstreams. Bitstreams can be accumulated in breadth-first order of the geometric tree, and each slice can be matched with a set of tree layers (see [link to G-PCC]). Figure 33 (b)). Segmented slices can inherit the hierarchical structure of the G-PCC bitstream.

[0412] A slice may not affect previous slices, just as higher layers in a geometric tree do not affect lower layers.

[0413] The segmented slicing method described in the implementation is effective in terms of error robustness, efficient transmission, and support for the region of interest.

[0414] 1) Error recovery

[0415] Compared to single-slice structures, segmented slicing offers greater error recovery. When a slice contains the entire bitstream of a frame, data loss can affect the entire frame. On the other hand, when the bitstream is segmented into multiple slices, even if some other slices are lost, some slices unaffected by the loss can still be decoded.

[0416] 2) Scalable transmission

[0417] It can support multiple decoders with different capabilities. When the encoded data is in a single slice, the LOD of the encoded point cloud can be determined before encoding. Accordingly, multiple pre-coded bitstreams of point cloud data with different resolutions can be sent independently. This can be inefficient in terms of large bandwidth or storage space.

[0418] When generating a PCC bitstream and including it in segmented slices, a single bitstream can support different levels of decoders. From the decoder's perspective, the receiver can select the target layer and can send a portion of the selected bitstream to the decoder. Similarly, by using a single PCC bitstream without splitting the entire bitstream, a partial PCC bitstream can be efficiently generated on the transmitter side.

[0419] 3) Region-based spatial scalability

[0420] Regarding G-PCC requirements, region-based spatial scalability can be defined as follows: The compressed bitstream can be configured to have one or more layers. A specific region of interest can have a high density of additional layers, and these layers can be predicted from lower layers.

[0421] To support this requirement, different levels of detail representation of the region must be supported. For example, in VR / AR applications, distant objects can be represented with low precision, while nearby objects can be represented with high precision. Alternatively, the decoder can increase the resolution of the region of interest upon request. This can be achieved using scalable structures of G-PCC such as geometric octrees and scalable attribute coding schemes. The decoder should access the entire bitstream based on the current slice structure, which includes the entire geometry or attributes. This can lead to inefficiencies in bandwidth, memory, and decoder performance. On the other hand, when the bitstream is segmented into multiple slices and each slice includes a sub-bitstream according to a scalable layer, the decoder, according to the implementation, can select slices as needed before efficiently parsing the bitstream.

[0422] Figure 34 The hierarchical structure of the geometric coding tree and the aligned hierarchical structure of the attribute coding tree according to the embodiments are illustrated.

[0423] The method / apparatus according to the implementation method can be used as follows: Figure 34The hierarchical structure of the point cloud data shown is used to generate slice layer groups.

[0424] The method / apparatus according to the implementation can apply segmentation of the geometric and attribute bitstreams contained in different slices. Furthermore, in terms of tree depth, a encoded tree structure containing each slice and its geometric and attribute encodings from the partial tree information can be used.

[0425] Figure 34 (a) shows an example of the geometric tree structure and the proposed segment.

[0426] For example, eight levels can be configured in an octree, and five slices can be used to contain sub-bitstreams of one or more levels. A group represents a set of geometric tree levels. For example, group 1 includes levels 0 through 4, group 2 includes level 5, and group 3 includes levels 6 and 7. Additionally, a group can be divided into three subgroups. Parent-child pairs exist within each subgroup. Groups 3-1 through 3-3 are subgroups of group 3. When using scalable attribute encoding, the tree structure is the same as the geometric tree structure. The same octree-slice mapping can be used to create attribute slices (…). Figure 35 (b)

[0427] A layer group represents a bundle of layer structure units generated in G-PCC encoding, such as octree layers and LoD layers.

[0428] A subgroup can be represented by the set of neighboring nodes based on the location information of a layer group. Alternatively, the set of neighboring points can be based on the lowest layer in the layer group (which can be the layer closest to the root, and can be in the...). Figure 34 In the case of group 3 (layer 6), the configuration can be based on Morton code order, distance, or encoding order. Additionally, rules can be defined to ensure that nodes with parent-child relationships exist in the same subgroup.

[0429] When defining subgroups, boundaries can be formed in the middle of layers. Whether to maintain continuity at the boundaries can be indicated using `sps_entropy_continuation_enabled_flag`, `gsh_entropy_continuation_flag`, etc., and `ref_slice_id` can be provided. This allows the continuation of previous slices to be preserved.

[0430] Figure 35 The hierarchical structure of the geometric tree and the independent hierarchical structure of the attribute coding tree according to the implementation method are illustrated.

[0431] The method / apparatus according to the implementation can generate geometry-based slice layers and attribute-based slice layers, such as... Figure 35 As shown in the image.

[0432] The attribute coding layer can have a structure different from that of the geometric coding tree. (See reference...) Figure 28 (b) can define groups independently of the geometric tree structure.

[0433] To efficiently utilize the hierarchical structure of G-PCC, segments of slices can be provided that are paired with the geometry and attribute hierarchical structure.

[0434] For a geometric segment, each segment can contain encoded data from a layer group. Here, a layer group is defined as a set of consecutive tree layers, where the start depth and end depth of the tree layer can be a specific number within the tree depth, and the start depth is less than the end depth.

[0435] For attribute segments, each segment can contain encoded data from a group of layers. Here, depending on the attribute encoding scheme, a layer can be a tree depth or a Level of Detail (LOD).

[0436] The order of the encoded data in a slice can be the same as the order of the encoded data in a single slice.

[0437] The following can be provided as a set of parameters included in the bitstream.

[0438] In the set of geometric parameters, the hierarchical structure corresponding to the geometric tree layer needs to be described by, for example, the number of groups, group identifiers, the number of tree depths in a group, and the number of subgroups in a group.

[0439] In the attribute parameter set, indication information indicating whether the slice structure is aligned with the geometric slice structure is required. The number of groups, group identifiers, tree depths, and segments are defined to describe the layer group structure.

[0440] Define the following components in the slice header.

[0441] In the geometry slice header, you can define group identifiers, subgroup identifiers, etc., to identify the group and subgroup of each slice.

[0442] In the attribute slice header, when the attribute layer structure is not aligned with the geometry group, the group and subgroup of each slice must be identified.

[0443] Figure 36 The syntax of the parameter set according to the implementation method is shown.

[0444] Figure 36 The syntax can be with Figures 25 to 28 The parameter information is included together. Figure 24 In the bitstream.

[0445] num_layer_groups_minus1+1 specifies the number of layer groups, where a layer group represents a set of consecutive tree layers that are part of a geometry or attribute-encoded tree structure.

[0446] layer_group_id specifies the layer group identifier for the i-th layer group.

[0447] num_tree_depth_minus1+1 specifies the number of tree depths contained in the i-th level group.

[0448] num_subgroups_minus1+1 specifies the number of subgroups in the i-th level group.

[0449] A `aligned_layer_group_structure_flag` value of 1 indicates that the layer group and subgroup structure of the attribute slice is the same as the geometric layer group and subgroup structure. A `aligned_layer_group_structure_flag` value of 0 indicates that the layer group and subgroup structure of the attribute slice is not the same as the geometric layer group and subgroup structure.

[0450] geom_parameter_set_id specifies a set of geometric parameters that contains information about the hierarchical and subgroup structures aligned with the attribute hierarchical structure.

[0451] Figure 37 A geometric data cell header according to an embodiment is shown.

[0452] Figure 37 The head can be with Figures 25 to 28 The parameter information is included together. Figure 24 In the bitstream.

[0453] The subgroup_id specifies an indicator of a subgroup within the layer group indicated by the layer_group_id. The subgroup_id can range from 0 to num_subgroups_minus1.

[0454] layer_group_id and subgroup_id can be used to indicate the order of slices and can be used to sort slices according to bitstream order.

[0455] Reference ​ According to the implementation method / apparatus and encoder, point cloud data can be transmitted by dividing it into units for transmission. A bitstream generator can divide and package the data into units suitable for selecting necessary information within the bitstream units based on hierarchical structure information. ​ ).

[0456] Reference​ According to the implementation method / apparatus and decoder, geometric data and attribute data can be reconstructed based on the bitstream layer. ​ ).

[0457] In this scenario, the sub-bitstream classifier can transmit appropriate data to the decoder based on information in the bitstream header. Alternatively, the layer required by the receiver can be selected during this process.

[0458] Reference ​ ,based on ​ The sliced ​​layered bitstream can be used to select geometric slices and / or attribute slices by referring to the necessary parameter information, and then decoded and rendered.

[0459] based on ​ Implementation methods, such as ​ As shown, compressed data can be partitioned and sent according to layers, and only the necessary portions of the pre-compressed data can be selectively sent at the bitstream level without a separate transcoding process. In this case, only one storage space is required per stream, thus allowing for efficient storage space management. Furthermore, efficient transmission can be achieved in terms of bandwidth (bitstream selector) because only the necessary layers are selected before transmission.

[0460] Furthermore, the receiving method / apparatus according to the embodiment can receive the bitstream slice by slice, and the receiver can selectively send the bitstream to the decoder based on the density of the point cloud data to be represented according to the decoder performance or the application domain. In this case, since the selection is performed before decoding, the decoder efficiency can be improved, and decoders with various performance levels can be supported.

[0461] ​ A method for sending point cloud data according to an implementation method is illustrated.

[0462] S3800: The point cloud data transmission method according to the embodiment may include encoding the point cloud data.

[0463] The encoding according to the implementation method may include ​ The transmitting device 10000 and the point cloud video encoder 10002 are included. ​ The code 20001 in ​ encoder in ​ The transmitting device in ​ XR device 1430 in ​ encoder in ​ LOD-based hierarchical data configuration in ​ LOD-based geometry / attribute bitstream configuration in ​ Slice-based bitstream configuration in​ Generation of a bitstream containing parameters ​ The parameters generated, geometry / attribute encoder and sub-bitstream generator, metadata generator, and ​ Multiplexers in ​ Geometry / attribute encoding and bitstream selection in ​ Slice segmentation and slice grouping in the middle ​ and ​ The generation of parameters in the process.

[0464] S3801: The point cloud data transmission method according to the embodiment may further include transmitting a bit stream containing point cloud data.

[0465] The transmission according to the implementation method may include ​ The transmitting device 10000 and transmitter 10003 in the middle ​ Sending 20002 in ​ The transmission of the encoded bit stream in the middle, ​ Data transmission of the XR device 1430 in the middle, according to ​ Sending all or part of the encoded bit stream in the middle. ​ Transmission of geometry / attribute bitstreams based on LOD (layer) ​ Slice-based bitstream transmission ​ Sending a bitstream containing parameters ​ Sending parameters in ​ The transmitter in ​ Partial bitstream transmission in ​ Slicing, segmenting, and bitstream transmission in the context of bitstream transmission. ​ and ​ Sending parameters in the code.

[0466] ​ A method for receiving point cloud data according to an embodiment is illustrated.

[0467] S3900: The point cloud data receiving method according to the embodiment may include receiving a bit stream containing point cloud data.

[0468] Receiving according to the implementation method may include ​ The receiving device 10004 and receiver 10005 in the middle, according to ​ Sending and receiving in Figure 13 Receiving bitstreams in Figure 14 The XR device 1430 receives data. Figure 15 Receiving all or part of the bit stream in the middle Figures 17 to 22 Reception of LOD-based geometry / attribute bitstreams in [the context of] Figure 23Slice-based bitstream reception Figure 24 Reception of bitstreams containing parameters Figures 25 to 28 Receiving parameters in Figure 30 The receiver in Figure 31 Bit stream reception in Figure 32 Partial bitstream reception in Figures 33 to 35 The reception of sliced ​​segmented bitstreams and sliced ​​packetized bitstreams, and Figure 36 and Figure 37 The parameters are received in the process.

[0469] S3901: The point cloud data receiving method according to the embodiment may further include decoding the point cloud data.

[0470] Decoding according to the implementation method may include Figure 1 The receiving device 10004 and the point cloud video decoder 10006 in the middle Figure 2 Decoding 20003 in Figure 10 and Figure 11 decoder in Figure 13 The receiving device in Figure 14 XR device 1430 in Figure 15 Full / partial bitstream decoding and bitstream selection and decoding in [the context of the program]. Figure 15 Decoding of all or part of the bitstream in the middle, Figures 17 to 22 Decoding of LOD-based geometry / attribute bitstreams in [the context of] ... Figure 23 Slice-based bitstream decoding Figure 24 Decoding of a bitstream containing parameters Figures 25 to 28 Decoding parameters in Figure 30 The separator, sub-bitstream classifier, metadata parser, geometry / attribute decoder, and renderer in the middle. Figure 31 Geometry / attribute slice selection in Figure 32 Partial bitstream decoding in Figures 33 to 35 Decoding of sliced ​​segmented bitstreams and sliced ​​block bitstreams, and Figure 36 and Figure 37 Decoding the parameters in the code.

[0471] The transmission method according to the embodiment (implemented by the transmission device) may include encoding point cloud data and transmitting a bit stream containing the point cloud data.

[0472] The transmission method (apparatus) according to the implementation can generate the following bit stream to have a structure based on distinguishing units such as layers / LODs / slices.

[0473] Specifically, the bitstream can contain geometric data and attribute data of the point cloud data. The bitstream for geometric data can contain levels of detail (LODs), where LOD 1 can contain the geometric data and additional geometric data contained in LOD 0, and the bitstream for attribute data can include LODs, where LOD 1 can include the attribute data and additional attribute data contained in LOD 0 (see [link to documentation]). Figure 17 and Figure 18 (LOD-based bitstream configuration method in [the context]).

[0474] The number of LODs can be, for example, 2. Depending on the level of detail in the data, the bitstream can have even more classification units.

[0475] Additionally, the bitstream used for geometry data may precede the bitstream used for attribute data. Alternatively, the geometry data corresponding to LOD 0 and the attribute data corresponding to LOD 0 may precede the geometry data corresponding to LOD 1 and the attribute data corresponding to LOD 1 in the bitstream (see [link to bitstream]). Figure 19 and Figure 20 (Bitstream arrangement method in the text).

[0476] Additionally, the bitstream can contain geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, and attribute data corresponding to LOD 1. Geometric data and attribute data corresponding to LOD 2 can be excluded from the bitstream.

[0477] Additionally, the bitstream can contain geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, attribute data corresponding to LOD 1, and geometric data corresponding to LOD 2. Attribute data corresponding to LOD 2 can be excluded from the bitstream (see [link to documentation]). Figure 21 and Figure 22 (Geometric-attribute symmetry / asymmetry in the text).

[0478] Additionally, bitstreams can contain point cloud data individually within LOD-based layers and slices of point cloud data based on layers (see [link]). Figure 23 (Slice configuration in the middle).

[0479] Additionally, the bitstream may include segmented slices containing split data (see...). Figure 33 (segmented slice structure in the text).

[0480] Additionally, the bitstream for geometric data may contain slices comprising groups of geometric data for one or more layers, and the bitstream for attribute data may contain slices comprising groups of attribute data for one or more layers.

[0481] Furthermore, the layer structure of the bitstream used for geometric data can be the same as or different from the layer structure of the bitstream used for attribute data (see [link]). Figure 34 and Figure 35 (layer group slices in the middle).

[0482] A point cloud data receiving device corresponding to a transmitting device that implements the transmitting method according to the embodiment can be configured to implement the receiving method according to the embodiment as follows.

[0483] The receiving apparatus may include a receiver configured to receive a bitstream containing point cloud data and a decoder configured to decode the point cloud data (see [link to documentation]). Figure 1 ).

[0484] The bitstream classifier used for the decoder (see...) Figure 30 The processed bitstream can contain geometric and attribute data of the point cloud data. The bitstream for geometric data can contain levels of detail (LODs), where LOD 1 can include the geometric data included in LOD 0 and additional geometric data. The bitstream for attribute data can include LODs, where LOD 1 can include the attribute data included in LOD 0 and additional attribute data.

[0485] Additionally, the bitstream for geometry data may precede the bitstream for attribute data. Alternatively, the geometry data corresponding to LOD 0 and the attribute data corresponding to LOD 0 may precede the geometry data corresponding to LOD 1 and the attribute data corresponding to LOD 1 in the bitstream.

[0486] Therefore, as Figure 32 As shown, compressed data is partitioned and sent according to layers, and only the necessary portions of the pre-compressed data can be selectively sent at the bitstream level without a separate transcoding process. In this case, only one storage space is required per stream, thus storage space can be operated efficiently. Furthermore, efficient transmission can be achieved in terms of bandwidth (bitstream selector) because only the necessary layers are selected before transmission.

[0487] Furthermore, the bitstream can be divided into slices for transmission / reception, and the receiver can selectively decode the bitstream based on the density of the point cloud data to be represented, depending on the decoder's performance or the application domain. In this case, since selection is performed before decoding, decoder efficiency can be improved, and decoders with various performance levels can be supported.

[0488] The implementation has been described from the perspective of methods and / or apparatus, and the description of methods and apparatus can be applied to complement each other.

[0489] Although the accompanying drawings have been described separately for simplicity, new embodiments can be designed by incorporating the embodiments illustrated in the corresponding figures. The design of a computer-readable recording medium on which a program for performing the above embodiments is recorded, as required by those skilled in the art, also falls within the scope of the appended claims and their equivalents. The apparatus and method according to the embodiments are not limited to the configuration and methods of the above embodiments. Various modifications can be made to the embodiments by selectively combining all or some of the embodiments. Although preferred embodiments have been described with reference to the accompanying drawings, those skilled in the art will appreciate that various modifications and variations can be made to the embodiments without departing from the spirit or scope of this disclosure as described in the appended claims. Such modifications will not be understood as independent of the technical concept or viewpoint of the embodiments.

[0490] Various elements of the apparatus according to the embodiments can be implemented by hardware, software, firmware, or a combination thereof. Various elements in the embodiments can be implemented by a single chip (e.g., a single hardware circuit). According to the embodiments, the components according to the embodiments can be implemented as separate chips. According to the embodiments, at least one or more components of the apparatus according to the embodiments can include one or more processors capable of executing one or more programs. The one or more programs can execute any one or more of the operations / methods according to the embodiments, or include instructions for executing them. Executable instructions for executing the methods / operations of the apparatus according to the embodiments can be stored in a non-transitory CRM or other computer program product configured to be executed by one or more processors, or can be stored in a transient CRM or other computer program product configured to be executed by one or more processors. Furthermore, the memory according to the embodiments can be used as a concept that covers not only volatile memory (e.g., RAM) but also non-volatile memory, flash memory, and PROM. Additionally, it can be implemented in a carrier form such as transmission via the Internet. Furthermore, the processor-readable recording medium can be distributed across a network-connected computer system, allowing processor-readable code to be stored and executed in a distributed manner.

[0491] In this specification, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Similarly, “A, B” can mean “A and / or B”. Additionally, “A / B / C” can mean “at least one of A, B, and / or C”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Additionally, in this specification, the term “or” should be interpreted as indicating “and / or”. For example, the expression “A or B” can mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term “or” as used in this document should be interpreted as indicating “additionally or alternatively”.

[0492] Terms such as "first" and "second" can be used to describe various elements of the embodiments. However, the various components according to the embodiments should not be limited to the above terms. These terms are only used to distinguish one element from another. For example, a first user input signal can be referred to as a second user input signal. Similarly, a second user input signal can be referred to as a first user input signal. The use of these terms should be interpreted without departing from the scope of the various embodiments. Both a first user input signal and a second user input signal are user input signals, but they do not mean the same user input signal unless the context clearly indicates otherwise.

[0493] The terminology used to describe embodiments is for the purpose of describing particular embodiments and is not intended to limit the embodiments. As used in the description of embodiments and in the claims, the singular forms “a,” “an,” and “the” include plural indicators unless the context clearly specifies otherwise. The word “and / or” is used to include all possible combinations of terms. Terms such as “comprising” or “having” are intended to indicate the presence of figures, numbers, steps, elements, and / or components and should be understood not to exclude the possibility of the additional presence of figures, numbers, steps, elements, and / or components. As used herein, conditional expressions such as “if” and “when” are not limited to optional cases but are intended to be interpreted as performing a related operation or interpreting a related definition based on a specific condition when that condition is met.

[0494] Operations according to the embodiments described in this specification can be performed by a transmitting / receiving device including a memory and / or a processor according to the embodiments. The memory may store programs for processing / controlling the operations according to the embodiments, and the processor may control the various operations described in this specification. The processor may be referred to as a controller, etc. In the embodiments, operations may be performed by firmware, software, and / or combinations thereof. Firmware, software, and / or combinations thereof may be stored in a processor or memory.

[0495] The operations according to the above embodiments can be performed by the transmitting and / or receiving devices according to the embodiments. The transmitting / receiving device includes a transmitter / receiver configured to transmit and receive media data, a memory configured to store instructions (program code, algorithms, flowcharts, and / or data) of the process according to the embodiments, and a processor configured to control the operation of the transmitting / receiving device.

[0496] The processor may be referred to as a controller, etc., and may correspond to, for example, hardware, software, and / or a combination thereof. The operations according to the above embodiments can be performed by the processor. Alternatively, the processor may be implemented as an encoder / decoder for the operations of the above embodiments.

[0497] Open mode

[0498] As described above, the relevant content has been described in the best mode for implementing the method.

[0499] Industrial applicability

[0500] As described above, the implementation methods can be applied in whole or in part to point cloud data transmission / reception devices and systems.

[0501] Those skilled in the art will understand that various changes or modifications can be made to the implementation methods within the scope of the implementation methods.

[0502] Therefore, the embodiments are intended to cover modifications and variations of this disclosure, provided that they fall within the scope of the appended claims and their equivalents.

Claims

1. A method for encoding point cloud data, the method comprising the following steps: The geometric data of the point cloud data is encoded based on the first local slice, wherein the first local slice is mapped to a subgroup in the layer group of the layer group structure. Attribute data for the attributes of the point cloud data is encoded based on a second local slice, wherein the second local slice is mapped to the subgroup within the layer group of the layer group structure; and Send a bitstream including the point cloud data. The bitstream further includes information representing the number of layer groups, information representing the number of tree depths within the layer groups, information for identifying the layer groups, and information for identifying the subgroups. The layer group structure includes multiple layer groups associated with the information used to represent the number of layer groups. Each of the plurality of layer groups includes a plurality of tree depths associated with the number of the aforementioned information for the tree depth in the layer group, and The first local slice and the second local slice are identified by the same pair of information for identifying the layer group and information for identifying the subgroup.

2. The method according to claim 1, wherein, The bitstream contains the geometric data and attribute data of the point cloud data. The bitstream used for the geometric data includes Levels of Detail (LODs) including LOD 0 and LOD 1. LOD 1 contains the geometric data and additional geometric data contained in LOD 0. The bitstream used for the attribute data includes a Level of Detail (LOD), and LOD 1 includes the attribute data and additional attribute data included in LOD 0.

3. The method according to claim 2, wherein, In the bitstream, the bitstream used for the geometric data precedes the bitstream used for the attribute data, or In the bitstream, the geometric data and attribute data corresponding to LOD 0 are placed before the geometric data and attribute data corresponding to LOD 1.

4. The method according to claim 2, wherein, The bitstream contains geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, and attribute data corresponding to LOD 1. The bitstream does not contain geometric data or attribute data corresponding to LOD 2. The bitstream contains geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, attribute data corresponding to LOD 1, and geometric data corresponding to LOD 2, but does not contain attribute data corresponding to LOD 2.

5. The method according to claim 1, wherein, The bitstream includes point cloud data classified into LOD-based layers, and includes slices containing point cloud data based on the layers.

6. The method according to claim 1, wherein, The bitstream includes segmented slices containing segmented point cloud data.

7. The method according to claim 6, wherein, The bitstream for geometric data includes slices comprising groups of said geometric data for one or more layers. The bitstream for attribute data includes slices comprising groups of attribute data for the one or more layers. The layer structure of the bitstream used for the geometric data and the layer structure of the bitstream used for the attribute data may be the same as or different from each other.

8. An apparatus for encoding point cloud data, the apparatus comprising: An encoder configured to encode geometric data of location for point cloud data based on a first local slice, wherein the first local slice is mapped to a subgroup in a layer group structure, and to encode attribute data of attributes for the point cloud data based on a second local slice, wherein the second local slice is mapped to the subgroup in the layer group structure; and A transmitter configured to transmit a bitstream comprising the point cloud data. The bitstream further includes information representing the number of layer groups, information representing the number of tree depths within the layer groups, information for identifying the layer groups, and information for identifying the subgroups. The layer group structure includes multiple layer groups associated with the information used to represent the number of layer groups. Each of the plurality of layer groups includes a plurality of tree depths associated with the number of the aforementioned information for the tree depth in the layer group, and The first local slice and the second local slice are identified by the same pair of information for identifying the layer group and information for identifying the subgroup.

9. The apparatus according to claim 8, wherein, The bitstream contains the geometric data and attribute data of the point cloud data through the bitstream generator of the encoder. The bitstream used for the geometric data includes Levels of Detail (LODs) including LOD 0 and LOD 1. LOD 1 contains the geometric data and additional geometric data contained in LOD 0. The bitstream used for the attribute data includes a Level of Detail (LOD), and LOD 1 includes the attribute data and additional attribute data included in LOD 0.

10. The apparatus according to claim 9, wherein, In the bitstream, the bitstream used for the geometric data precedes the bitstream used for the attribute data, or In the bitstream, the geometric data and attribute data corresponding to LOD 0 are placed before the geometric data and attribute data corresponding to LOD 1.

11. The apparatus according to claim 9, wherein, The bitstream contains geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, and attribute data corresponding to LOD 1. The bitstream does not contain geometric data or attribute data corresponding to LOD 2. The bitstream contains geometric data corresponding to LOD 0, attribute data corresponding to LOD 0, geometric data corresponding to LOD 1, attribute data corresponding to LOD 1, and geometric data corresponding to LOD 2, but does not contain attribute data corresponding to LOD 2.

12. The apparatus according to claim 8, wherein, The bitstream contains point cloud data classified into LOD-based layers and includes slices containing point cloud data based on the layers, through the bitstream generator of the encoder.

13. The apparatus according to claim 8, wherein, The bitstream is generated by the bitstream generator of the encoder and includes segmented slices of point cloud data.

14. The apparatus according to claim 13, wherein, The bitstream for geometric data includes slices comprising groups of said geometric data for one or more layers. The bitstream for attribute data includes slices comprising groups of attribute data for the one or more layers. The layer structure of the bitstream used for the geometric data and the layer structure of the bitstream used for the attribute data may be the same as or different from each other.

15. A method for decoding point cloud data, the method comprising the following steps: Receive a bitstream containing point cloud data; The geometric data for the location of the point cloud data is decoded based on a first local slice, wherein the first local slice is mapped to the geometry of a subgroup within a layer group structure; and The attribute data of the point cloud data is encoded based on the second local slice, wherein the second local slice is mapped to the attributes of the subgroup in the layer group of the layer group structure. The bitstream includes information representing the number of layer groups, information representing the number of tree depths within the layer groups, information for identifying the layer groups, and information for identifying the subgroups. The layer group structure includes multiple layer groups associated with the information used to represent the number of layer groups. Each of the plurality of layer groups includes a plurality of tree depths associated with the number of the aforementioned information for the tree depth in the layer group, and The first local slice and the second local slice are identified by the same pair of information for identifying the layer group and information for identifying the subgroup.

16. The method according to claim 15, wherein, The bitstream contains the geometric data and attribute data of the point cloud data. The bitstream used for the geometric data includes Levels of Detail (LODs) including LOD 0 and LOD 1. LOD 1 contains the geometric data and additional geometric data contained in LOD 0. The bitstream used for the attribute data includes a Level of Detail (LOD), and LOD 1 includes the attribute data and additional attribute data included in LOD 0.

17. The method according to claim 16, wherein, In the bitstream, the bitstream used for the geometric data precedes the bitstream used for the attribute data, or In the bitstream, the geometric data and attribute data corresponding to LOD 0 are placed before the geometric data and attribute data corresponding to LOD 1.

18. An apparatus for receiving point cloud data, the apparatus comprising: A receiver configured to receive a bitstream containing point cloud data; as well as A decoder is configured to decode geometric data representing the location of the point cloud data based on a first local slice, wherein the first local slice is mapped to the geometry of a subgroup within a layer group of a layer group structure; and to encode attribute data representing the attributes of the point cloud data based on a second local slice, wherein the second local slice is mapped to the attributes of the subgroup within the layer group of the layer group structure. The bitstream includes information representing the number of layer groups, information representing the number of tree depths within the layer groups, information for identifying the layer groups, and information for identifying the subgroups. The layer group structure includes multiple layer groups associated with the information used to represent the number of layer groups. Each of the plurality of layer groups includes a plurality of tree depths associated with the number of the aforementioned information for the tree depth in the layer group, and The first local slice and the second local slice are identified by the same pair of information for identifying the layer group and information for identifying the subgroup.

19. The apparatus according to claim 18, wherein, The bitstream contains the geometric data and attribute data of the point cloud data, and the bitstream is processed by a bitstream classifier for the decoder. The bitstream used for the geometric data includes Levels of Detail (LODs) including LOD 0 and LOD 1. LOD 1 contains the geometric data and additional geometric data contained in LOD 0. The bitstream used for the attribute data includes a Level of Detail (LOD), and LOD 1 includes the attribute data and additional attribute data included in LOD 0.

20. The apparatus according to claim 19, wherein, In the bitstream, the bitstream used for the geometric data precedes the bitstream used for the attribute data, or In the bitstream, the geometric data and attribute data corresponding to LOD 0 are placed before the geometric data and attribute data corresponding to LOD 1.